preciz
This is going on for a while now. I think it’s best described with an example:
We have an Oban queue “postgresql” with 100K jobs available.
We deploy production, the nodes start and the queues start processing the jobs.
Then the queues start to disappear (no errors in logs) and the processing is slow-ish, it is still going but on the nodes the queues appear / disappear kind of randomly and at one point usually 1/3 of them are running.
And this happens only with a few queues, usually the ones that have a larger number of jobs waiting.
Anybody ever had this problem?
Trending in Questions
Hello!
Suppose you are building workflow (order / task / payment) processing system with the following requirements:
Each workflow con...
New
I’m in search of an Elixir library that offers PDF generation capabilities similar to Ruby’s Prawn. While there have been discussions abo...
New
I’m looking to build a personal workflow to quickly deploy web applications written in elixir/phoenix, for local consumption (ie not on t...
New
Before I dive in myself, did anyone successfully sprinkle Hologram into their existing LiveView app?
Looking for hints regarding:
Addi...
New
Kia ora,
We have been using elixir-google-api to connect to Google Drive. However, with the updates to Tesla due to CVEs this is now bro...
New
Hi all, I wanted to ask how the community is dealing with post-release steps.
Today we have Ecto migrations, which make sure that the db...
New
Hello,
I have an Elixir backend that implements a custom protocol over TCP. I want to load test the backend and assess the performance o...
New
Other Trending Topics
Hey, I’m Jesse and I’m the main contributor behind Dexter, a full-featured, lightning-fast Elixir LSP optimized for large codebases. It s...
New
Beam Bots (or just BB for short) is a framework for building fault-tolerant robotics applications in Elixir using familiar OTP patterns. ...
New
ExRatatui lets you cook up rich terminal UIs in Elixir, powered by Rust’s ratatui via Rustler NIFs. Build interactive terminal applicatio...
New
Hello everyone. After busy few months I am happy to announce v0.1.0 of Emerge & Solve.
They are GUI (Emerge) and State management (S...
New
Corex is an accessible, unstyled UI component library for Phoenix that integrates Zag.js state machines using Vanilla JavaScript and Live...
New
Emily is an Elixir library that runs Nx computations on Apple’s MLX. Install it as the default Nx backend and Nx, defn, Axon, Nx.Serving,...
New
Categories:
Sub Categories:
Forums
Popular Tags
- #ecto
- #liveview
- #troubleshooting
- #learning-elixir
- #deployment
- #library
- #erlang
- #testing
- #genserver
- #mix
- #absinthe
- #remote-other
- #otp
- #plug
- #how-to-question
- #macros
- #postgres
- #channels
- #elixirconf
- #exunit
- #discussion
- #code-sync
- #javascript
- #podcasts
- #onsite
- #dialyzer
- #docker
- #authentication
- #umbrella
- #full-time-contract
- #podcasts-by-brainlid
- #ecto-query
- #elixir-ls
- #blog-post
- #phoenix_html
- #iex
- #graphql
- #ai
- #genstage
- #elixirconf-us
- #websockets
- #supervisor
- #advent-of-code
- #distillery
- #processes
- #api
- #forms
- #metaprogramming
- #security
- #hex










Showing Posts 1 to 6- Show Best Posts
- Show All Posts (oldest first)
- Show All Posts (newest first)
sorentwo
Which Oban and Pro versions? Also, which hosting provider? The only ways I’m aware of that the producers would slow down and potentially crash is from query timeouts, but that would log errors.
Are these inserted all at once, or was it accumulated over time? It shouldn’t change whether the queue stays active, but Pro v1.7 shipped with a feature called automatic spacing, which prevents a big dump of jobs from overwhelming the queue.
preciz
Latest hex version but this has been going on for a while, can’t remember when it started but this was the only issue we couldn’t figure out in private messages back then and now it got really annoying so I’m just looking for somebody who experienced the same and maybe solved it.
Our DB is a hyper optimized postgres with beefy hardware 256 GB RAM and this happens even after a full vacuum so the speed is not the problem I believe. Jobs are accumulated over time.
preciz
Looks like this, all those queues should have 52 limit (13 node * 4) and running 52 jobs but they have the limit jumping up and down (as queues disappear and restart on nodes). This slows down processing, that’s why there is a postgres2 and postgres3 queue as a mitigation for this problem (so processing goes faster).
sorentwo
Latest Oban version and latest Pro version? The latest Pro release has fine-grained telemetry around the fetch transaction (the main thing a producer does), so if you’re on v1.7+ we can get some better instrumentation in place to see what’s happening.
preciz
Okay how should I provide you the data then?
sorentwo
Attach a telemetry handler like this, which you can optionally scope to a single queue/producer, and gather metrics for a healthy instance as it starts to degrade. This shows a Logger, but you can output however you like to a CSV or something, then email it to us at support.