smaximov
Greetings!
Tonight we had an incident in production when a somewhat important GenServer process seemingly disappeared and had not been restarted by a supervisor. This is the first incident in 2.5 years since this worker had been introduced.
The GenServer in question has basically the following structure:
defmodule MyWorker do
use GenServer, restart: :transient
def start_link(opts) do
GenServer.start_link(__MODULE__, opts, name: __MODULE__)
end
@impl GenServer
def init(opts) do
Process.send_after(self(), :do_work, 500)
{:ok, opts}
end
@impl GenServer
def handle_info(:do_work, state) do
do_some_work(state)
Process.send_after(self(), :do_work, 500)
{:noreply, state}
end
defp do_some_work(state) do
# query DB
# send HTTP request
# that's all - no OTP messaging, exits, etc.
end
end
This GenServer is started as part of a Supervisor (let’s call it MySupervisor). MySupervisor uses the :one_for_one strategy, all other init options are default.
When I connected to the remote console, I found out that the process is not alive - Process.whereis(MyWorker) returned nil, Supervisor.which_children(MySupervisor) returned the following entry for MyWorker:
{MyWorker, :undefined, :worker, [MyWorker]}
As far as I understand, this means MyWorker’s child spec was known to MySupervisor, but the worker process itself was not running. I restarted MyWorker with Supervisor.restart_child/2 and started figuring out what went wrong.
The obvious culprit is restart: :transient - as MyWorker was intended to work indefinitely and to never successfully terminate, it should have been :permanent. If we assume MyWorker terminated normally, it explains why it wasn’t restarted.
The problem is figuring out why MyWorker died, because it neither stops itself, nor it is explicitly stopped by some other app code. I have 2 hypotheses left:
- some developer connected to the remote console and stopped the worker by hand (very unlikely);
- the worker terminated with an error and was not restarted due to a bug in
Supervisor(extremely unlikely).
Am I missing anything else?
Trending in Questions
Other Trending Topics
Categories:
Sub Categories:
Forums
Popular Tags
- #ecto
- #liveview
- #troubleshooting
- #learning-elixir
- #deployment
- #library
- #erlang
- #testing
- #genserver
- #mix
- #absinthe
- #remote-other
- #otp
- #plug
- #how-to-question
- #macros
- #postgres
- #elixirconf
- #channels
- #exunit
- #discussion
- #code-sync
- #javascript
- #podcasts
- #onsite
- #dialyzer
- #docker
- #authentication
- #umbrella
- #full-time-contract
- #podcasts-by-brainlid
- #ecto-query
- #blog-post
- #elixir-ls
- #ai
- #elixirconf-us
- #phoenix_html
- #iex
- #graphql
- #genstage
- #websockets
- #supervisor
- #advent-of-code
- #distillery
- #processes
- #api
- #forms
- #hex
- #security
- #metaprogramming











Showing Posts 1 to 10- Show Best Posts
- Show All (oldest first)
- Show All (newest first)
lud
Hi
Welcome back!
Is there
exit(:normal)orexit({:shutdown, some_value})somewhere in your code that could have been called by that process?Or maybe it was linked to another process that exits with shutdown? I don’t remember how GenServer behaves in that case.
smaximov
Thanks for the reply, I think you nailed the two most possible cases that I haven’t though of!
No, there’re no explicit exits that can be called from this process (unless I missed something). I also doubt that Ecto and HTTPoison have such exits in their codebases that could leak into user’s code.
As for how GenServer behaves in this case, I think this part of
Process.link/1docs is relevant:So if my understanding is correct, if
MyWorkerwas linked to another process, and that process exited with a:shutdown(or{:shutdown, value}), then it would indeed causeMyWorkerto terminate normally as well. Unfortunately,MyWorkerisn’t linked to any other process.mudasobwa
There are two things to take into consideration as well:
① Restart strategies, specifically
:max_restarts- the maximum number of restarts allowed in a time frame. Defaults to3and:max_seconds- the time frame in which:max_restartsapplies. Defaults to5. If something went south and there were 4 attempts to restart it within 5 seconds interval (defaults,) all failing, theGenServerwould end up not restarted. This is likely your case, if e. g. the DB connection has dropped for 3 seconds. The culprit would be 500ms interval ininit/1, which makes 4 consequtive insuccessful attempts to executedo_some_work/1to fit into 5 secs and shutdown the process forever.② The mailbox overflow. This technically might be your case, but extremely unlikely, just saying for the complete picture.
smaximov
I was under the impression that if a supervised process restarts more than
:max_restartsin a:max_secondsperiod, it would lead to the restart of the whole supervisor (and not to the supervised process to end up not being restarted). Is it not the case? I didn’t find specific mentions of this behaviour in Elixir’s Supervisor docs, but Erlang docs say (emphasis mine):mudasobwa
I have no access to
iexatm, but it could be easily validated with 10 LoCs.Once you mentioned this, I recall the same experience: the sup terminates, but then it comes to
Applicationand I doubtApplicationwould terminate itself as well. This is the dark spot anyway, because too much of internal VM machinery would be involved and one God knows how3and5are calculated in the real world.I just wanted to mention it, because in my experience if you don’t really expect an infinite loop there, it’d be safer to increase
:max_restartsparameter and forget about this potential issue.LostKobrakai
The app will terminate if the root pid exits, and immediatelly on the first exit, no restarts. And if the app is
:permanentit’ll make the whole vm stop.smaximov
Good call! It seems like exceeding
:max_restartsrestarts in a:max_periodwindow does lead to the restart of the supervisor:Running this code will lead to an endless stream of errors and log messages of GenServers and Supervisor restarts:
You can see that
TestSup.Supervisorhas been restarted.If you add
max_restarts: 1tooptsinTestSup.Application, you’ll see it will eventually lead to the shutdown of the application:Yeah, it’s a solid advice which I will implement (combined with changing the
:restarttype ofMyWorkerto:permanent), I just wanted to get to the bottom of the issue, if possible.lud
Yes the supervisor will exit if it has to restart a child too many times. And it will be restarted by its own supervisor.
You can have a chain of Sup → Sup → Sup → Sup → Worker and all elements of the chain will behave like this.
A good way to think about this is that the code describing the supervisor tree (each supervisor children, and the children description in each child supervisor, and so on) describes “how the app state should be at runtime”. And the system will try its best to maintain that state, or otherwise fail and shutdown totally.
Of course that description can be altered at runtime by calling start_child, delete_child to add/remove parts of the tree, and of course having
:transient,:temporaryor:significantchildren.But there is no way it will skip some parts of the tree unless you tell it to. Max-restarts and al’ will only have a temporary impact on the state of the system, but the app will always get back to the desired state.
If the child spec for MyWorker has
restart: :transientthen yes! (:shutdownis not “normally” but the supervisor will not restart it indeed).Any progress on finding the cause?
:transientis replaced with:temporary; but I guess this is unlikely.Supervisor.terminate_childorGenServer.stopin your codebase?:significantwith:any_significantor:all_significant?anuaralfetahe
Are there any logs available that could help identify why the process shut down?
I might be wrong, but I believe that if the process terminated multiple times in quick succession, the supervisor should have eventually shut down as well, which would have brought the entire VM down. However, that doesn’t seem to be the case here — the supervisor somehow stayed up but stopped restarting the process.
It’s also possible that someone accidentally connected to the VM and manually stopped the process.
smaximov
Unfortunately, no progress so far. Also, “no” to all of your points here.
For now we decided to change the restart type to
:permanentand monitor this issue closely. I will follow up with more information if it arrives.We didn’t find any suspicious entries in the logs (we know the approximate time the issue occurred from the application metrics). There were also no relevant errors reported in our Sentry instance in that time frame.
It’s possible, but very unlikely. There are maybe 2 people in our company other than me with the necessary knowledge required to gracefully terminate a supervised process, and the issue happened in the middle of the night. We also asked our Kubernetes admins to check if anyone accessed the pods in that time frame, just to be sure.