houayang

houayang

I am trying to test oban job retries locally, my perform job returns {:error, %{}, and I see the job is in a retryable state, but it never runs and attempts is still 1.

oban: “~> 2.15.0”
oban_pro: “~> 0.14”
oban_web: “~> 2.6”

Showing Posts 1 to 10

sorentwo

sorentwo

Oban Core Team

That’s not much info to help you debug. Is this in the development environment? Can you verify that the queue is running with Oban.check_queue(queue: "name_of_queue")?

houayang

houayang OP

Yes this is my local test environment. The queue is running, tries to run the job once and when it errors, just stays in a retryable state.

Oban.check_queue(Manifold.Oban, queue: :delivery_destination_queue) %{ name: "Manifold.Oban", node: "M1-Max", running: [], queue: "delivery_destination_queue", started_at: ~U[2024-08-02 19:54:57.208154Z], updated_at: ~U[2024-08-02 19:55:20.222256Z], global_limit: nil, local_limit: 1, paused: false, rate_limit: nil, retry_attempts: 5, retry_backoff: 1000 }

sorentwo

sorentwo

Oban Core Team

Is there a custom backoff on the job? When is the scheduled_at timestamp for it to retry?

houayang

houayang OP

no custom backoff

inserted_at: 2024-08-05 15:39:11.287173
scheduled_at: 2024-08-05 15:39:28.327155
attempted_at: 2024-08-05 15:39:11.309171
completed_at: NULL
attempted_by: "{M1-Max,0726da66-d47b-47fd-bc5e-1db1ec3a7e85}"
discarded_at: NULL
priority: 3
tags: {}
meta: "{""span_context"": null}"
cancelled_at: NULL
state: retryable
sorentwo

sorentwo

Oban Core Team

And this is still happening, even after a reset? Please share the output of Oban.config(), feel free to omit queue names or cron workers.

Side note—there are many reliability fixes and diagnostic improvements in more recent versions of Oban (not to mention Pro and Web). If possible, I highly recommend upgrading.

houayang

houayang OP

iex(1)> Oban.config(Manifold.Oban)
%Oban.Config{
  dispatch_cooldown: 5,
  engine: Oban.Pro.Queue.SmartEngine,
  get_dynamic_repo: nil,
  log: false,
  name: Manifold.Oban,
  node: "M1-Max",
  notifier: Oban.Notifiers.Postgres,
  peer: Oban.Peers.Disabled,
  plugins: [
    {Oban.Web.Plugins.Stats, []},
    {Oban.Pro.Plugins.DynamicLifeline, [rescue_interval: 300000]},
    {Oban.Plugins.Gossip, [interval: 5000]},
    {Oban.Pro.Plugins.DynamicPruner,
     [
       limit: 100000,
       mode: {:max_age, {2, :days}},
       queue_overrides: [
         delete_expired: {:max_age, {1, :day}},
         upload_worker: {:max_age, {1, :day}},
         file_queue: {:max_age, {1, :day}},
         delivery_destination: {:max_age, {3, :day}},
       ]
     ]}
  ],
  prefix: "public",
  queues: [
    delivery_destination: [limit: 1],
    periodic_queue: [limit: 1],
    process: [limit: 1]
  ],
  repo: Engine.Repo,
  shutdown_grace_period: 15000,
  stage_interval: 1000,
  testing: :disabled
}
houayang

houayang OP

Also as an FYI, if I add schedule_in, to the job option, that job also is never attempted.

sorentwo

sorentwo

Oban Core Team

See if your stager is alive with Oban.Registry.whereis(Oban, Oban.Stager). If it is alive, add this telemetry to see if it reports any errors or other information.

:telemetry.attach_many(:debug, [[:oban, :plugin, :exception], [:oban, :plugin, :stop]], fn _, time, meta, _ ->
  if meta.plugin == Oban.Stager, do: IO.inspect({time, Map.delete(meta, :conf)})
end, nil)
houayang

houayang OP

iex(2)> Oban.Registry.whereis(Manifold.Oban, Oban.Stager)
#PID<0.3316.0>
iex(3)> :telemetry.attach_many(:debug, [[:oban, :plugin, :exception], [:oban, :plugin, :stop]], fn _, time, meta, _ ->
...(3)>   if meta.plugin == Oban.Stager, do: IO.inspect({time, Map.delete(meta, :conf)})
...(3)> end, nil)

20:08:17.173 level=[info] The function passed as a handler with ID :debug is a local function.
This means that it is either an anonymous function or a capture of a function without a module specified. That may cause a performance penalty when calling that handler. For more details see the note in `telemetry:attach/4` documentation.

https://hexdocs.pm/telemetry/telemetry.html#attach/4 line=112 pid=<0.3846.0> file=/Users/hyang/cars/cars_platform/deps/telemetry/src/telemetry.erl mfa=:telemetry.attach_many/4 
:ok
{%{monotonic_time: -576460289938607883, duration: 2329583},
 %{
   telemetry_span_context: #Reference<0.3291262840.1439170561.97292>,
   plugin: Oban.Stager,
   staged_count: 0,
   staged_jobs: []
 }}
{%{monotonic_time: -576460289931805058, duration: 1137269},
 %{
   telemetry_span_context: #Reference<0.3291262840.1439170561.97305>,
   plugin: Oban.Stager,
   staged_count: 0,
   staged_jobs: []
 }}
{%{monotonic_time: -576460289904803716, duration: 2169496},
 %{
   telemetry_span_context: #Reference<0.3291262840.1439170561.97316>,
   plugin: Oban.Stager,
   staged_count: 0,
   staged_jobs: []
 }}
{%{monotonic_time: -576460288931119893, duration: 1765531},
 %{
   telemetry_span_context: #Reference<0.3291262840.1439170561.97352>,
   plugin: Oban.Stager,
   staged_count: 0,
   staged_jobs: []
 }}
{%{monotonic_time: -576460288929637285, duration: 1357398},
 %{
   telemetry_span_context: #Reference<0.3291262840.1439170562.51763>,
   plugin: Oban.Stager,
   staged_count: 0,
   staged_jobs: []
 }}
{%{monotonic_time: -576460288902098684, duration: 1741655},
 %{
   telemetry_span_context: #Reference<0.3291262840.1439170562.51768>,
   plugin: Oban.Stager,
   staged_count: 0,
   staged_jobs: []
 }}
{%{monotonic_time: -576460287925900192, duration: 3010885},
 %{
   telemetry_span_context: #Reference<0.3291262840.1439170561.97430>,
   plugin: Oban.Stager,
   staged_count: 0,
   staged_jobs: []
 }}
sorentwo

sorentwo

Oban Core Team

It’s not staging anything, which means it can’t see any jobs within the time window. Will you output the console time (DateTime.utc_now()) and the scheduled_at timestamp from a scheduled job in the database?

You can also try running this:

Oban.Engine.stage_jobs(Oban.config(), Oban.Job, limit: 5000)

And see if it returns a count.

Where Next? Top

Trending in Questions Top

Blokh
Hey guys, I’ve got a huge CSV ( around 10 GB ) that needs to be processed hourly Do you guys have any suggestions what is the best prac...
New
RSP87
I’m working on a project that simulates the bumbl example in the programming phoenix book. It acts almost like an email client. We have a...
New
kszambelanczyk
Hello! Could someone please give me a help/sample code, how to delete a file from s3 using waffle/waffle_ecto from Phoenix app. I creat...
New
RemyXRenard
I’m seeing that a list inside a Kino.DataTable will be interpreted as a charlist, even if the Kino.configure() is set to charlists: :as_l...
New
velrest
So my question is quite simple and i have found no conclusive answer on forum, google or AI. Should we use :erlang.float for Integer to ...
New
samoloth
Hi, I’ve just set up an application with ash_authentication. There is only magic link strategy for now, so there is no confirmation add o...
New
FlyingNoodle
If a change or preparation module uses Ash.Changeset.get_argument/2 or Ash.Query.get_argument/2 (or any of the other get_argument functio...
New

Other Trending Topics Top

mudasobwa
I am happy to introduce the very α version of the new programming language compiled to BEAM. Welcome Cure. It has literally three kille...
New
garrison
Hobbes is a low-level distributed database for the Elixir programming language. Hobbes provides a simple, safe, and scalable storage lay...
New
marciok
Hi there! We created Gust: A task orchestrator inspired by Airflow. For those who have never heard about Aiflow, it’s a Python-based wor...
New
jimsynz
Beam Bots (or just BB for short) is a framework for building fault-tolerant robotics applications in Elixir using familiar OTP patterns. ...
New
Damirados
Hello everyone. After busy few months I am happy to announce v0.1.0 of Emerge &amp; Solve. They are GUI (Emerge) and State management (S...
New
netoum
Corex is an accessible, unstyled UI component library for Phoenix that integrates Zag.js state machines using Vanilla JavaScript and Live...
New

We're in Beta

About us Mission Statement

Options

Thread Display Mode




Thread Preview

Skip Thread Previews