vegabook

vegabook

adbc seems to use a huge amount more memory than “native” drivers.

I’m dumping large amounts of data each minute, about 600k data points across 10k different identifiers. I need column store for querying so I have gone to duckdb via adbc. But it’s using tons of memory even though I am specifiying an on disk file. I have written a small test to show what I mean in this repo: GitHub - vegabook/adbc_leak_demo · GitHub . You’ll see that adbc basically accumulates memory, especially for duckdb, but also for sqlite, at a much higher pace than native drivers.

I’d like to stick to “offical” adbc if possible as it will also allow me more easily to swap out backends. So is this a bug or some implementation issue? What can I do to avoid this behaviour? As it stands it basically makes adbc difficult to use right now for append workflows as memory just keeps on growing forever. Even using sqlite backend the memory growth is kind of a concern. duckdbex is kind of acceptable as memory growth seems slow enough that a once-every-few-hours restart or something should keep thins manageable. But really only exqlite currently is fit for a long-running ingest, but it has the major downside of not doing column store so data science queries will be ultra slow.

Any advice on how to solve the “lots of fast ingest, must be column store, mustn’t leak memory” problem?

The table and chart show process memory growth as rows get ingested. Each iter is 100 rows. All are on-disk workflows so apart from a bit of buffer delta early on, you’d expect stability.

Where Next? Top

Trending in Questions Top

Blokh
Hey guys, I’ve got a huge CSV ( around 10 GB ) that needs to be processed hourly Do you guys have any suggestions what is the best prac...
New
kszambelanczyk
Hello! Could someone please give me a help/sample code, how to delete a file from s3 using waffle/waffle_ecto from Phoenix app. I creat...
New
Onor.io
I have what I’ve heard referred to as a “lookup table” in my database. This is a way of assigning codes to common values. One common lo...
New
Trolleger
What approach to take when sending live updates to “random” users Hi! I have a question, I have a little chat app, and when I create a DM...
New
RemyXRenard
I’m seeing that a list inside a Kino.DataTable will be interpreted as a charlist, even if the Kino.configure() is set to charlists: :as_l...
New
matt-savvy
Anyone here using Honeybadger? My Honeybadger account is being overwhelmed with noise from some bots. Seeing a lot of Bandit.HTTPError...
New
samoloth
Hi, I’ve just set up an application with ash_authentication. There is only magic link strategy for now, so there is no confirmation add o...
New

Other Trending Topics Top

mudasobwa
I am happy to introduce the very α version of the new programming language compiled to BEAM. Welcome Cure. It has literally three kille...
New
garrison
Hobbes is a low-level distributed database for the Elixir programming language. Hobbes provides a simple, safe, and scalable storage lay...
New
mcass19
ExRatatui lets you cook up rich terminal UIs in Elixir, powered by Rust’s ratatui via Rustler NIFs. Build interactive terminal applicatio...
New
Damirados
Hello everyone. After busy few months I am happy to announce v0.1.0 of Emerge & Solve. They are GUI (Emerge) and State management (S...
New
netoum
Corex is an accessible, unstyled UI component library for Phoenix that integrates Zag.js state machines using Vanilla JavaScript and Live...
New
wintermeyer
There are three potential reasons for members of this forum to have a look at https://vutuv.de You are tired or annoyed of LinkedIn. Yo...
New

We're in Beta

About us Mission Statement

Options

Thread Display Mode




Thread Preview

Skip Thread Previews