quatermain
Hi,
We have a regular Phoenix project using Ecto and Postgresql. But recently we are seeing a lot of connection timeouts or idle/closed connections. All looks like problem that DB pool is exhausted. Traffic has been pretty low, queries look ok, no big queries running all the time, no big Oban jobs. But I’m pretty sure we’re missing something.
So is it possible to debug which processes and code places are using the DB pool? Thanks to telemetry, we’re able to run code when these errors occur, but the question is whether it’s possible to get details about pool usage.
Telemetry handler:
def handle_event(
[:enaia, :repo, :query],
measurements,
%{result: {:error, %DBConnection.ConnectionError{} = error}} = metadata,
_config
) do
# check and log details about Pool usage
end
I can get repo’s child, but I’m not sure how to proceed.
repo_pid = Process.whereis(MyApp.Repo)
children = Supervisor.which_children(repo_pid)
Trending in Questions
I having some trouble figuring out if I have set myself too strict of standards for my production server. Currently I can handle 75% of r...
New
Documentation
While reading the Scoped Routes section, I noticed that the documentation currently refers to a problem without explainin...
New
Hello,
I’m trying to build a basic Phoenix web-app, and I’d like to use Tailwind.
However, when I launch mix phx.server, I get an error...
New
Hi everyone,
I am toying with the idea of building a “match maker” for giving personal help to people that wants to start coding.
I sta...
New
I recently noticed that Elixir’s Logger defaults its primary log level to :debug when no :logger, :level application configuration is pre...
New
I’m working on a small exercise involving update_in/3, and I came up with this solution:
data = %{
name: "Periodic Table",
category:...
New
I’ve got trouble wrapping my head around the order in which functions are called in this snippet (from Phoenix’s authentication):
toke...
New
Other Trending Topics
Edit: 2026 May 15 - This post is archived.
Mob is alive!!
Main docs: mob v0.7.11 — Documentation
A bit of explanation for the slightly c...
New
Hey, I’m Jesse and I’m the main contributor behind Dexter, a full-featured, lightning-fast Elixir LSP optimized for large codebases. It s...
New
I am happy to introduce the very α version of the new programming language compiled to BEAM.
Welcome Cure.
It has literally three kille...
New
Hobbes is a low-level distributed database for the Elixir programming language.
Hobbes provides a simple, safe, and scalable storage lay...
New
Hi everyone!
The first release candidate for the Expert language server project is now available!
We’ve published a press release detai...
New
A little off-topic, but I feel like people here have a good head on their shoulders.
I used to be quite good at making software. Was luc...
New
Categories:
Sub Categories:
Forums
Popular Tags
- #ecto
- #liveview
- #troubleshooting
- #learning-elixir
- #library
- #deployment
- #erlang
- #testing
- #genserver
- #mix
- #absinthe
- #remote-other
- #otp
- #plug
- #how-to-question
- #macros
- #postgres
- #elixirconf
- #channels
- #exunit
- #discussion
- #code-sync
- #podcasts
- #javascript
- #onsite
- #dialyzer
- #docker
- #authentication
- #umbrella
- #full-time-contract
- #podcasts-by-brainlid
- #ai
- #ecto-query
- #elixirconf-us
- #blog-post
- #elixir-ls
- #phoenix_html
- #iex
- #graphql
- #genstage
- #websockets
- #supervisor
- #advent-of-code
- #distillery
- #processes
- #elixirconf-eu
- #api
- #forms
- #metaprogramming
- #hex











Showing Posts 1 to 10- Show Best Posts
- Show All (oldest first)
- Show All (newest first)
josevalim
First step is to see your data for queue times, query times, and idle times, as those will show how busy the pool is. Also, what do you mean by connection timeouts?
The idle/closed connections may happen if the database or a proxy are closing them and are not necessarily an indicator of a problem (although we do ping the connection every second).
quatermain
We get these errors:
how you can see, waiting time in queue is wild. Sometime it’s in hundreds and sometimes thousands of ms.
Our configuration:
Other than that is as default. Bigger timeout is because sometimes per day we run some long running queries. But these were not running at that time.
Database is running on Fly in cluster of 3 machines with size
performance-8x-cpu@16384MB. App itself it doesn’t do big stuff. At that time when this happen last time a few hours ago we had max 20 users online and we were doing some background jobs. See stats from AppSignal.Issue happen at 7:48am, you can see that ~10 minutes before we have “bigger” traffic than in time crash happen.
see peak before
what do you suggest to track when this error happen to help us to remove these errors?
Really thank you for you time.
josevalim
In this case there is something overloaded in the system. I would follow the suggestions in the message, in particular tracking queue_time and query_time, both in telemetry events, finding slow queries in the database, or checking if the machine is overloaded, by measuring the run queues or the machine cpu/memory stats.
quatermain
From fly dashboard only this shows some kind of peak, other metrics are consistents:
So basically machine is not overloaded when this event happen, both app and database to be correct. When this event happen we have some struggles with DB immediately after “pool is overloaded”. It means that pool is fully used, but it feels more like some pools are “abandoned”. I mean no big queries are running, most of queries went to ConnectionError and database start closing pools. Sometimes this even mean that master machine goes to zombie mode and new master has to be elected. Election happen but zombie has to be cleaned manually.
So question is, when I will have query_time and queue_time. I’m pretty sure that queue_ time will be long, but query_time first regular, later bigger. At least that execution time in AppSignal looks like, graph went up when this even happen.
My goal was to see what is using all pool connections when error happen, when query is timeout because it’s waiting too long in pool. But I think it’s not so simple how I expected.
Also when I add custom metric to every ecto query which we track in AppSignal, it will not help us, because it will be hard to identify as problematic query. We see execution time, so adding query_time will tell us that it was not affected by full pool, but execution time can be affected by event or primary cause.
I will try to select all queries around the event and see execution time and see slow queries. But we have not seen anything big there.
josevalim
Other things that could be helpful:
Investigate if this relates to deployments somehow
Chart p90% and p99% not averages, especially for queue and query times, but you should be able to track queries that take too long. Even if by using custom telemetry handlers. If anything takes more than a second, make sure to log that
quatermain
We will soon release where we will track these numbers on own dashboard and per every sample.
Code from custom telemetry handler for better reference:
Hope it will clarify behaviour of our app or database.
We had two small incidents meanwhile. No big queries were executed based on telemetry.
But we had some long running Oban jobs but based on telemetry they were long because 3rd party api response was super slow. But that should not affect DB pool, because only output is stored in DB and no db transaction is called around that api call. But maybe it could affect network, but…
pdgonzalez872
I was going to ask if you were on Fly and read the post above. The follow up question is: are you using fly postgres? If so, something similar happened to me a few weeks ago. In my case it was that one of the machines (I have 3 for the db like they suggest) for no reason crashed and was stuck in a boot loop, never becoming available completely. I was looking at connections like you were, I’d see the same errors. After contacting support, they mentioned the bad machine and, they were correct! You can see it in the dashboard, 2/3 healthy machines. I’m not sure this is your problem, but it was for me. In case it is, I have some notes on how to fix it and can share, hope this helps.
quatermain
We already have had this issue with crashed primary machine multiple times this week, we changed machine with more resources maybe a week ago. so it’s new machine. In case this is the issue, something has to broke that machine even with new ones
quatermain
please share it, maybe it will help us to see our problem. Thanks
pdgonzalez872
Hmm, interesting.
Absolutely, here are the steps (if it doesn’t help you, maybe it will help someone else):
One of the things support said was to make sure the machine you cloned was a good one. Could it be you cloned a bad one?
I feel like I’m pulling you into a rabbit hole, sorry about that, but I’ve had to do this twice now (that’s why I had notes
)
Have you reached out to Fly’s support?