cjk
Hi there,
I have a problem with ecto timeouts (so it seems). In a Elixir application I wrote a module for importing data. It gets a CSV file, goes through that file line by line (with a Stream.map()), looks if this dataset already exists in the database, updates it or inserts it as a new dataset. Pretty basic.
To make that operation restartable I wrap this whole process in a database transaction. The amount of datasets is pretty large and it can take up to an hour.
I know of Ecto timeouts, and thus I start that transaction with timeout: :infinity to avoid timeout problems. It works in development, but in prod I get strange errors:
16:02:34.435 [error] Postgrex.Protocol (#PID<0.423.0>) disconnected: ** (DBConnection.ConnectionError) ssl send: closed
or
** (exit) an exception was raised:
** (DBConnection.ConnectionError) ssl send: closed
(ecto_sql) lib/ecto/adapters/sql.ex:624: Ecto.Adapters.SQL.raise_sql_call_error/1
(ecto_sql) lib/ecto/adapters/sql.ex:557: Ecto.Adapters.SQL.execute/5
(ecto) lib/ecto/repo/queryable.ex:147: Ecto.Repo.Queryable.execute/4
(ecto) lib/ecto/repo/queryable.ex:18: Ecto.Repo.Queryable.all/3
(ecto) lib/ecto/repo/queryable.ex:66: Ecto.Repo.Queryable.one/3
(termitool) lib/termitool/meta/meta.ex:332: Termitool.Meta.get_user_by/1
(termitool) lib/termitool/meta/meta.ex:439: Termitool.Meta.get_by_username/1
(termitool) lib/termitool/meta/meta.ex:465: Termitool.Meta.username_password_auth/2
Basically connection errors at random places in the application. I can avoid this problem by setting timeout: :infinity in the repo configuration, but this seems wrong and dangerous to me.
The code is basically:
Repo.transaction(fn ->
stream
|> Stream.map(fn {row, idx} -> update_or_create_row(row) end)
end,
timeout: :infinity)
I am using a stream because that’s what I get from the CSV parsing library.
What am I doing wrong?
Best regards,
CK
Trending in Questions
Other Trending Topics
Categories:
Sub Categories:
Forums
Popular Tags
- #ecto
- #liveview
- #troubleshooting
- #learning-elixir
- #library
- #deployment
- #erlang
- #testing
- #genserver
- #mix
- #absinthe
- #remote-other
- #otp
- #plug
- #how-to-question
- #macros
- #postgres
- #elixirconf
- #channels
- #exunit
- #discussion
- #code-sync
- #podcasts
- #javascript
- #onsite
- #dialyzer
- #docker
- #authentication
- #umbrella
- #full-time-contract
- #podcasts-by-brainlid
- #ecto-query
- #elixirconf-us
- #ai
- #blog-post
- #elixir-ls
- #phoenix_html
- #iex
- #graphql
- #genstage
- #websockets
- #supervisor
- #advent-of-code
- #distillery
- #processes
- #api
- #forms
- #metaprogramming
- #hex
- #security











Showing Posts 18 to 9- Show Best Posts
- Show All (oldest first)
- Show All (newest first)
cjk
OK, after downgrading everything just works. Thanks for the hint, @outlog and @engineeringdept!
cjk
Huh…
I actually get this error nightly since a few days. I had not yet the time to look into this. So it seems very likely that you are right.
Phew, you just solved me some headaches…
outlog
think you are hitting this 21.3 OTP bug that is just emerging:
and mentioned here Workaround for Docker with Erlang/OTP 21.3 and
Redirecting… (identical but with patch)
can you downgrade to OTP 21.2.7 ? or wait for patch release (presumably later today)..
cjk
I’m not using distillery (had not yet the time to wrap my head around that), I just checked out the source and compiled it on the server. Using
mix phx.servervia systemd to start the service.cjk
Nope. Bare metal with Ubuntu 18.04 LTS
outlog
fishing expedition here..
but check OTP and elixir versions and make sure that you are on latest patch release of what you are using.. (recently there was a transient bug with ConnectionError fixed in later OTP..)
also can you check/post mix hex.outdated for the relevant dependencies..
also is this using distillery and what version?
engineeringdept
Are you running in a Docker container? I discovered a bug in the the official Elixir 1.8.1 container, which is backed by Erlang 21.3, which caused SSL problems like these in production. I had to switch to a custom image running Erlang 21.2.7.
cjk
Other applications using ecto on the same server work like a charm… none of my queries seem to take more than 20ms, most of them below 10ms. The
:pool_sizeis 50. I’m seriously confused.cjk
Dammit.
The errors appeared again, random connection drops by Postgrex:
And this time there was not even an import job running, I just removed the
timeout: :infinityfrom the repo configuration.What’s going on?!
cjk
Hm. Race conditions on the data are a non-issue in this case. But I will overhaul my import pipeline, thanks for your input…
That said, I guess I found the underlying cause for this problem. Your hint with concurrently importing the data gave me the idea: the CSV library uses workers to parallelize the CSV reading and parsing. That means that my
Stream.map()executions are parallelized as well.Parallelized execution means: different ecto processes are used. And since the server has traffic and more cores than my workstation this means: more Stream workers, more user connections and thus the pool (I was using the default size of 15) could be exhausted pretty fast.
And indeed, if I reduce the pool size on my workstation to 2 and the timeout to 1 second, I get the same errors all over the place.
To avoid that I now use the
:calleroption on all repo calls, and now it works like a charm on my dev machine. Have still to test it in production, thoughThis also means: my import did not run in the transaction anyways…