warmwaffles
Author of the exqlite library here. I need some help from experienced NIF devs in hunting down the culprit of this SIGSEGV error. This bug has been keeping me awake at night for a few weeks now.
Ya ya ya, it’s C, and I should be using Rust, but right now, I don’t want to make that change just yet for this library. Let’s ignore that for now.
What I am running into is that when I have an sqlite database, and a pool of DBConnections to a single sqlite database resource on disk. SQLite has a write ahead log functionality that works pretty well in most of my use cases that I’ve been hammering it with.
Except for an issue where database connections timeout. In order to simulate the timeouts in a reproduce-able manner, we shortened the timeout window to 1 second, write 1,000 new rows to a single table and queue up 5,000 async tasks to execute the database query.
Failing test implemented here:
Can be ran by checking that branch out and running the following
mix test test/exqlite/timeout_segfault_test.exs
Trending in Questions
Other Trending Topics
Categories:
Sub Categories:
Forums
Popular Tags
- #ecto
- #liveview
- #troubleshooting
- #learning-elixir
- #library
- #deployment
- #erlang
- #testing
- #genserver
- #mix
- #absinthe
- #remote-other
- #otp
- #plug
- #how-to-question
- #macros
- #postgres
- #elixirconf
- #channels
- #exunit
- #discussion
- #code-sync
- #podcasts
- #javascript
- #onsite
- #dialyzer
- #docker
- #authentication
- #umbrella
- #full-time-contract
- #podcasts-by-brainlid
- #ecto-query
- #elixirconf-us
- #ai
- #blog-post
- #elixir-ls
- #phoenix_html
- #iex
- #graphql
- #genstage
- #websockets
- #supervisor
- #advent-of-code
- #distillery
- #processes
- #api
- #forms
- #hex
- #security
- #metaprogramming










Showing Posts 1 to 10- Show Best Posts
- Show All (oldest first)
- Show All (newest first)
warmwaffles
Is there a way I can have valgrind with an
asdfinstallation or am I going to need to use some other erlang / elixir version manager to do this?jstimps
I’ve debugged my share of C NIFs in the past, so I hope I can provide some help. I pulled your branch and reproduced the segfault. Locally I added some fprintf-based logging to
exqlite_closeand noticed something that made my raise an eyebrow, so I thought I’d share. If I’m off base here, please let me know.An fprintf statement in
exqlite_closeprevents the segfault on my system, so there is definitely a timing element here (race condition). I printed the memory addresses forconnandconn->dband saw many duplicate addresses, which made me wonder ifexqlite_closeis being called multiple times on the same resource, resulting in a double-free-like segfault when running with high concurrency.BTW, thanks for this library. I am a fan.
warmwaffles
This was my observations as well. Which makes me believe this has something to do with the garbage collector.
jstimps
I’m not familiar with
DBConnection, but I’ve written similar libraries in the past. When you doDBConnection.start_link, is it calling yourconnectone time and then sharing the result among the pool of 50? If that’s the case, I wonder if creating a single enif resource and allowing it to be “disconnected” from multiple concurrent processes in the pool is the source of the race condition. TheDBConnectionREADME says it calls disconnect automatically:(It looks like DBConnection was written with network sockets in mind. Due to the nature of SQLite, obviously there is no network socket, so the disconnect behavior may not be optimal for your use case.)
In my experience, it’s safest to ensure that an enif resource gets a dedicated BEAM process so that you can carefully control the access to it. At the very least, I assume you only want a maximum of 1
exqlite_closeon a given resource; You can build that guarantee by serializing calls toConnection.closeon a dedicated BEAM process. Otherwise, you can’t guarantee there are not 2 scheduler threads freeing the memory simultaneously.We get spoiled by the BEAM scheduler and garbage collection, but when it comes to NIFs sadly we have to worry about all these issues again. I don’t believe the BEAM garbage collector is causing any problems in this case.
Edit: in my haste I probably jumped to some wrong conclusions about DBConnection. I’m going to spend some more time and try to get more concrete results.
warmwaffles
The way I’ve written it is that every time a new connection to the pool is created, a new sqlite instance is opened and no sharing of sqlite databases between connections.
dimitarvp
Is
db_connectioneven necessary in SQLite’s case? When I was trying to get my library off the ground a few years ago I planned on using:poolboyor, as I learned lately, the:jobsErlang library.Plus if you open only one handle and set it up as fully multithreaded you can have only one stateful OTP process per SQLite connection string and return it each time a connection is checked out – and only close it if the processes is terminated – which you can do manually with an API e.g.
Exqlite.close(handle)or simply when the app itself shuts down.jstimps
Yes, I missed the mark on that detail. I see now that the pool is starting a process per enif resource.
However, the DBConnection docs sound as if they’re exposing the “socket” (in your case the enif resource handle) to client processes as an optimization. The lib is pretty sophisticated, so it’s challenging for me to grok quickly, but the docs suggest this is the case.
If I have this right, then you may find that multiple processes (the client procs) are calling
exqlite_closeconcurrently on the same enif resource.warmwaffles
Not necessarily, but I tied this pretty closely to being used by
ecto_sqlite3.I could use poolboy or nimble pool for pooled resources, but when it comes to Ecto, you have to implement the
DBConnectioninterface in order to have it play nice withecto_sql.warmwaffles
This certainly could be the case. I just don’t have a good way to prove that a double
closeis being called.dimitarvp
I get it and I’m not judging, simply saying that if I encounter something like this and it takes me more than a few hours to troubleshoot I’d likely just eliminate the problem by drastically changing approach.
I even have most of the Rust code that maintains a per-connection-string SQLite handles cache (a parallel hash map); IIRC I only needed to add tests there but it’s when I stopped everything due to personal and work life turmoil.
I will be starting a new job soon and after I settle a bit I’ll finally return to making my library workable. Maybe you can take inspiration from my Rust code; I’ll be happy to help when I have some more free time as well.
Feel free to ping me anytime!
Oh. I haven’t looked into that. I definitely will now, thanks.