Fl4m3Ph03n1x
Background
I have an app that must never go down, under any circumstances. This app has a supervisor that has a ton of workers. These workers are feeble and may die … a lot. Every time a worker dies I have a metrics system that tells me so I know when something is wrong.
Problem
The problem here is that I can’t find a way to prevent my Supervisor from restarting or just outright dying. I tried the following configuration:
Supervisor.start_link([], max_restarts: :infinity, strategy: :one_for_one)
But I got the following error:
(EXIT from pid<0.104.0>) shell process exited with reason: bad supervisor configuration, invalid max_restarts (intensity): :infinity
Question
How can I tell my supervisors to never die?
Trending in Questions
I having some trouble figuring out if I have set myself too strict of standards for my production server. Currently I can handle 75% of r...
New
Hello,
I’m trying to build a basic Phoenix web-app, and I’d like to use Tailwind.
However, when I launch mix phx.server, I get an error...
New
I really like the adapter patterns that ecto, nebulex, waffle, etc. use and would love find something similar for a key management servic...
New
Hello folks!
So at work, we are seeing some situations where we have to define some “fixed” strings that are used across the codebase in...
New
I’m working on a small exercise involving update_in/3, and I came up with this solution:
data = %{
name: "Periodic Table",
category:...
New
I’ve got trouble wrapping my head around the order in which functions are called in this snippet (from Phoenix’s authentication):
toke...
New
Is there any way to avoid the Hologram compiler running when using iex? It seems like the front-end code could potentially be disregarded...
New
Other Trending Topics
Edit: 2026 May 15 - This post is archived.
Mob is alive!!
Main docs: mob v0.7.11 — Documentation
A bit of explanation for the slightly c...
New
Hobbes is a low-level distributed database for the Elixir programming language.
Hobbes provides a simple, safe, and scalable storage lay...
New
A little off-topic, but I feel like people here have a good head on their shoulders.
I used to be quite good at making software. Was luc...
New
Hey. Is there anyone here who creates agents in their apps? Not talking about using agents, but creating them. I’m finding it pretty diff...
New
I fully migrated to my own harness from Anthropic/Gemini and I think it’s time to share it. Welcome DSH, the DeepSeek Harness, fully writ...
New
ExRatatui lets you cook up rich terminal UIs in Elixir, powered by Rust’s ratatui via Rustler NIFs. Build interactive terminal applicatio...
New
Categories:
Sub Categories:
Forums
Popular Tags
- #ecto
- #liveview
- #troubleshooting
- #learning-elixir
- #library
- #deployment
- #erlang
- #testing
- #genserver
- #mix
- #absinthe
- #remote-other
- #otp
- #plug
- #how-to-question
- #macros
- #postgres
- #elixirconf
- #channels
- #exunit
- #discussion
- #code-sync
- #podcasts
- #javascript
- #onsite
- #dialyzer
- #docker
- #authentication
- #umbrella
- #full-time-contract
- #ai
- #podcasts-by-brainlid
- #ecto-query
- #blog-post
- #elixirconf-us
- #elixir-ls
- #phoenix_html
- #iex
- #graphql
- #genstage
- #websockets
- #supervisor
- #advent-of-code
- #distillery
- #processes
- #elixirconf-eu
- #api
- #forms
- #security
- #metaprogramming










Showing Posts 1 to 10- Show Best Posts
- Show All (oldest first)
- Show All (newest first)
sasajuric
The point of
max_restarts(andmax_seconds) option is to break the endless restart cycle, and move recovery to the higher level. Thus, in my opinion, an infinity restart option doesn’t make sense. You could approximate it by e.g. using some insanely large number for both options, but I’d advise against doing it.There are some alternative approaches which could be considered, but first I’d like to learn more about what do these workers do, and why can they restart so frequently?
keathley
Totally agree with @sasajuric here. The supervisor’s job is to bring up your child processes in a stable state. If the supervisor can’t achieve that steady state then you want to continue restarting at a higher level until you eventually shut down the application (because at that point something has gone terribly wrong).
I think in this case you may want to rethink your problem. We don’t really have enough information to say but if your workers really are that transient then I would probably move the workers into a dynamic supervisor and then monitor and start them from some other controlling process. But there are a lot of options here. The important thing is to think about the problem in terms of what guarantees you want your supervisor to provide.
Fl4m3Ph03n1x
These workers use
gunto open HTTP2 connections to other servers. On startup, each worker opens a connection to a given domain and then keeps that connection open. This allows us to send an insane amount of requests per second.The issue here is that when the connection times out or dies, the worker is smart enough to try and reconnect. But what if it is impossible to reconnect? Maybe there is some corrupt state in the worker or maybe there is something else. In this scenario (and other simillar scenarios) I allow the worker to die so the supervisor can create a new one with a fresh state (as advised by your book
)
But what happens when a worker suicides over and over again? The Supervisor will try to restart it over and over again until it decides this is a lost cause, and then suicides himself as well.
Workers will fail. The systems they connect to may be unavailable for hours or days at a row. If my supervisor crashes, it starts a chain reaction that causes the app to crash. Thus my approach is to make the Supervisor never die while still being able to allow workers to die so they can come back online with a clean state.
Fl4m3Ph03n1x
I don’t want this to happen. Ever.
keathley
You’re fixating on the wrong part of that statement.
^ thats the important part. You need to think about how to bring up your connection processes in a stable state and assume that they’ll eventually crash. If you aren’t bringing them up in a stable state to begin with then you’ve designed a fragile system.
If you have designed the system to handle the connections correctly and your supervisor still can’t bring up your system into a steady state then you absolutely want to surface that error. Eventually if this propagates far enough up then you’ll want to shutdown. If you’ve done your design well and you reach the point of shutting down then it means that something drastic has gone wrong outside of your control. To put it a different way, who cares if your service is up if it can’t do anything useful or always serves the wrong answers? Its functionally equivalent to being shutdown anyway.
benwilson512
Don’t design around promises you can’t keep. This is central to the erlang philosophy. Your server will go bad. The network will go bad. Design to handle these.
Fl4m3Ph03n1x
Yes, but I rather have a useless service online that keeps pinging the team with error messages rather than having it crash completely. This is a design choice.
Yes, but it doesn’t mean I have to completely shutdown the entire application. I rather have it online checking for service availability every chance it gets and posting error messages to some service constantly than having it quit and shutdown.
peerreynders
As you describe it this is expected, not exceptional behaviour. Design needs to accomodate expected behaviour.
Sounds like working against live servers and watching for dead ones to come back online are distinct responsibilities that need to be handled separately (i.e. when a worker decides that the server won’t respond it terminates normally handing matters back to the dead server watcher).
jola
It sounds like you need to manage this with some form of backoff. You can’t code to prevent anything bad from ever happening, that’s just not possible. But if you have some idea of what can go wrong (and it sounds like you do), you should handle those expected errors. This might be in the form of try/catches, exponential backoff in trying to connect with gun etc. You would in this case put that logic in the worker, not in the Supervisor, and not rely on crashes to refresh the state in cases of known errors.
If that’s for some reason not possible, set the worker strategy to transient and let a separate process handle the restarts. This is not exactly the reason for, but kind of fits into, the use case for the new
Registry.selectin 1.9, where you would be able to register your workers and then query the Registry for the status of the workers, restarting as required. You’d then be able to keep track of workers that keep crashing, backoff and notify/alert, but keep the other workers living. This can also be implemented by usingProcess.monitor. Note that this is making your life difficult for yourself, and you’re introducing a lot more risk for corrupted states.When it comes to the unknown, there’s no way for you to keep the system working in a functioning condition. For those cases crashing is correct.
kwando
You can protect your workers with circuit breakers, one straight forward way to do this is to use this tiny application/library GitHub - jlouis/fuse: A Circuit Breaker for Erlang · GitHub