aviraj

aviraj

We have elixir application deployed on two servers. The application has been configured with some quantum jobs which execute from both servers at the same time. We want to limit execution of the quantum jobs to a single server. As per documentation, we could see the following options,

  1. Add global?: true to the job configuration
  2. run_strategy: {Quantum.RunStrategy.Random, :cluster}

We tried option (1), it didnt work. The job still executed from both servers. Do both (1) and (2) need to be set at the same time in a job? or there something else missing in the configuration

config :titan, Titan.Scheduler,
  timeout: 20_000,
  global?: true,
  jobs: [
    job_schedule: [
      schedule: "20 6 * * *",
      task: {Titan.ContentSchedule, :schedule_updates, []}
    ],
]

How does this clustering work? Even without giving the list of nodes, how do the two servers communicate among themselves to co-ordinate the job execution?

Showing Posts 16 to 7

bartblast

bartblast

Creator of Hologram

There is also the Citrine package, that aims to solve the cron clustering problem.

poops

poops

Sure, there were 25 queues running. I also ran SELECT * FROM oban_jobs_id_seq to see the last ID used, and it was 2915796. So we had close to 3 million jobs, although I don’t know how many were in there at the time of the spike.

Ah ok. I assumed it was the pruner because it happened right after enabling it. It could’ve been a coincidence. These were the only three plugins running at the time:

plugins: [
  Oban.Plugins.Pruner,
  Oban.Pro.Plugins.Lifeline,
  Oban.Web.Plugins.Stats
]
sorentwo

sorentwo

Oban Core Team

Those are the primary queries used to stage scheduled jobs and to fetch them for execution. They’re fully indexed and should be very fast, < 1ms under normal load. How many queues are/were you running? I’d love to know more to help you all, or other people in a similar situation.

It would be configured in the plugins section of your config, using the Dynamic Pruning Plugin. That is all moot though, since the queries you shared aren’t from pruning.

poops

poops

Here are the 2 queries our DBA sent me:

UPDATE “public”.“oban_jobs” AS o0 SET “state” = $1 WHERE (o0.“id” IN (SELECT so0.“id” AS “id” FROM “public”.“oban_jobs” AS so0 WHERE (so0.“state” IN (?,?)) AND (so0.“queue” = $2) AND (so0.“scheduled_at” <= $3) FOR UPDATE SKIP LOCKED))

UPDATE “public”.“oban_jobs” AS o0 SET “state” = $1, “attempted_at” = $2, “attempted_by” = $3, “attempt” = o0.“attempt” + $4 WHERE (o0.“id” IN (SELECT so0.“id” AS “id” FROM “public”.“oban_jobs” AS so0 WHERE (so0.“state” = ?) AND (so0.“queue” = $5) ORDER BY so0.“priority”, so0.“scheduled_at”, so0.“id” LIMIT $6 FOR UPDATE SKIP LOCKED)) RETURNING o0.“id”, o0.“state”, o0.“queue”, o0.“worker”, o0.“args”, o0.“errors”, o0.“tags”, o0.“attempt”, o0.“attempted_by”, o0.“max_attempts”, o0.“priority”, o0.“att

I’m not sure, I didn’t originally set it up. Where would that be configured?

sorentwo

sorentwo

Oban Core Team

I haven’t had any other reports of that, I would like to have seen the queries.

That is entirely normal. Even the demo at https://getoban.pro/oban has 600,000 completed jobs sitting around.

Any idea if you were using per-worker or per-state dynamic pruning? The pruner deletes 10k records every one minute by default, which is extremely fast.

connorlay

connorlay

Fixed! Thanks for pointing that out :smile:

Exadra37

Exadra37

It’s a 404 link, maybe the repo is private?

connorlay

connorlay

Hey! I recently solved this problem of running globally unique jobs in a clustered Elixir application. The project was written using Quantum 2, which had built-in support for clustering via the global: true mode. In our experience the implementation of global jobs was unreliable and we found that many jobs would stop executing entirely.

After doing some research, we ended up moving from Quantum to periodic for the job scheduler, combined with highlander to ensure uniqueness in the cluster.

I have a proof-of-concept here that shows the basic functionality.

poops

poops

Was using oban 2.0.0, oban_pro 0.3.0, and oban_web 2.0.0. It was a recent upgrade to those versions, with the following plugins:

 Oban.Plugins.Pruner,
 Oban.Pro.Plugins.Lifeline,
 Oban.Web.Plugins.Stat

Database CPU was pinned at 100% and our DBA said it was related to some queries to the oban_jobs table.
I don’t have exact numbers as the table eventually got pruned, but I believe there were hundreds of thousands of completed job records. My guess was that the pruning plugin locked things up, but I haven’t had a chance to look into it.

poops

poops

Correct, we didn’t need the crons to be distributed. Just needed to ensure only one process was running them.

Where Next? Top

Trending in Questions Top

katta
I having some trouble figuring out if I have set myself too strict of standards for my production server. Currently I can handle 75% of r...
New
achenet
Hello, I’m trying to build a basic Phoenix web-app, and I’d like to use Tailwind. However, when I launch mix phx.server, I get an error...
New
kpanic
Hi everyone, I am toying with the idea of building a “match maker” for giving personal help to people that wants to start coding. I sta...
New
Cxx-mlr
I’m working on a small exercise involving update_in/3, and I came up with this solution: data = %{ name: "Periodic Table", category:...
New
ChrisAmelia
I’ve got trouble wrapping my head around the order in which functions are called in this snippet (from Phoenix’s authentication): toke...
New
dillonoconnor
Is there any way to avoid the Hologram compiler running when using iex? It seems like the front-end code could potentially be disregarded...
New
thiagogsr
** (ArgumentError) expected :max_attempts to be a positive integer, got: {:@, [line: 10, column: 19], [{:max_attempts, [line: 10, column:...
New

Other Trending Topics Top

GenericJam
Edit: 2026 May 15 - This post is archived. Mob is alive!! Main docs: mob v0.7.11 — Documentation A bit of explanation for the slightly c...
New
mudasobwa
I am happy to introduce the very α version of the new programming language compiled to BEAM. Welcome Cure. It has literally three kille...
New
garrison
Hobbes is a low-level distributed database for the Elixir programming language. Hobbes provides a simple, safe, and scalable storage lay...
New
budgie
A little off-topic, but I feel like people here have a good head on their shoulders. I used to be quite good at making software. Was luc...
New
KristerV
Hey. Is there anyone here who creates agents in their apps? Not talking about using agents, but creating them. I’m finding it pretty diff...
New
mcass19
ExRatatui lets you cook up rich terminal UIs in Elixir, powered by Rust’s ratatui via Rustler NIFs. Build interactive terminal applicatio...
New

We're in Beta

About us Mission Statement

Options

Thread Display Mode




Thread Preview

Skip Thread Previews