manukall
Ecto/Postgres/Poolboy time outs after a few queries
Since a few days ago I keep getting the following error on my dev machine when making requests to a phoenix app:
exited in: :gen_server.call(#PID<0.656.0>, {:checkout, #Reference<0.1406881335.1305214977.237415>, true, 15000}, 5000) ** (EXIT) time out
It seems like the app works for some time and then the error happens for every request. Sometimes restarting Postgres or the app fixes the problem again for some time, but sometimes not. I suspect it maybe just works again after something times out.
I’m not sure what changed since I’ve last worked on this project some weeks ago. I’ve tried downgrading Postgresql and Elixir and also completely removed Postgresql, including the data directory.
The Phoenix logs for the affected requests show several successful queries before the error happens. It feels like connections are not checked in to the pool again, but I’ve no idea if that’s really true or where to start debugging. The same app works on a different laptop without those issues, though.
I’m on Arch Linux with Postgres 10.4, Ecto 2.2.10 and postgres 0.13.5.
Most Liked
amagdas
This looks to be fixed by Erlang: Timeouts on Poolboy.checkout · Issue #127 · elixir-ecto/db_connection · GitHub
manukall
Update: Seems to only happen with Erlang 20.3. After downgrading to 20.2.2, everything works again. My other project on the same machine was using 20.2.2. When I upgrade Erlang there to 20.3 it’s broken, too.
I’d like to raise this as a bug somewhere, but I’ve no idea where. I don’t know if it’s an issue with Erlang, Ecto, Postgrex, Poolboy, …
hubertlepicki
Oh that’s interesting. I am not using 20.3 yet but it seems like a good opportunity to update one of the projects and see if I can replicate your problem. I have got one that’s matching Ecto and postgrex versions to your set up, and is reasonably big so it will be a good test.
It might be a bug in poolboy or database_connection on this specific version of Erlang.
Last Post!
axelson
@amagdas Thanks for updating this thread!
This was all fascinating reading:
- Elixir bug report with initial investigation: Timeouts on Poolboy.checkout · Issue #127 · elixir-ecto/db_connection · GitHub
- Erlang bug report with c/assembly-level investigation: Redirecting…
- The patch that fixes the issue: Fix a race condition when generating async operation ids · jhogberg/otp@e27b98b · GitHub
A high-level summary is that erlang ports updated a shared global counter async_ref in a way that was not atomic. This race-condition has existed for quite a while but on pre GCC 8.1 this variable would only get read once in each thread. But GCC 8.1 (especially with some optimizations enabled) exposed the race-condition and caused prim_inet:recv0/3 to hang because it was waiting for the wrong reference (since it was waiting for a reference from a different thread). The bug was fixed by changing the async_ref to be unique per-port (since it doesn’t actually need to be shared between ports). This fix is included in Erlang/OTP 21.0.2.
If I got any of that wrong please feel free to correct it ![]()
Popular in Questions
Other popular topics
Categories:
Sub Categories:
Forums
Popular Tags
- #ecto
- #liveview
- #troubleshooting
- #learning-elixir
- #deployment
- #library
- #erlang
- #testing
- #genserver
- #mix
- #absinthe
- #remote-other
- #otp
- #plug
- #how-to-question
- #macros
- #postgres
- #channels
- #elixirconf
- #exunit
- #discussion
- #code-sync
- #javascript
- #podcasts
- #onsite
- #dialyzer
- #docker
- #authentication
- #umbrella
- #full-time-contract
- #podcasts-by-brainlid
- #ecto-query
- #elixir-ls
- #phoenix_html
- #iex
- #blog-post
- #graphql
- #genstage
- #ai
- #websockets
- #supervisor
- #elixirconf-us
- #advent-of-code
- #distillery
- #processes
- #forms
- #api
- #metaprogramming
- #hex
- #security









