Blokh
Hey guys,
I’ve got a huge CSV ( around 10 GB ) that needs to be processed hourly
Do you guys have any suggestions what is the best practice in that scenario?
Any useful mixes?
Thanks in advance.
Daniel
Trending in Questions
I having some trouble figuring out if I have set myself too strict of standards for my production server. Currently I can handle 75% of r...
New
Documentation
While reading the Scoped Routes section, I noticed that the documentation currently refers to a problem without explainin...
New
Hello,
I’m trying to build a basic Phoenix web-app, and I’d like to use Tailwind.
However, when I launch mix phx.server, I get an error...
New
Hi everyone,
I am toying with the idea of building a “match maker” for giving personal help to people that wants to start coding.
I sta...
New
I recently noticed that Elixir’s Logger defaults its primary log level to :debug when no :logger, :level application configuration is pre...
New
I’m working on a small exercise involving update_in/3, and I came up with this solution:
data = %{
name: "Periodic Table",
category:...
New
I’ve got trouble wrapping my head around the order in which functions are called in this snippet (from Phoenix’s authentication):
toke...
New
Other Trending Topics
Edit: 2026 May 15 - This post is archived.
Mob is alive!!
Main docs: mob v0.7.11 — Documentation
A bit of explanation for the slightly c...
New
Hey, I’m Jesse and I’m the main contributor behind Dexter, a full-featured, lightning-fast Elixir LSP optimized for large codebases. It s...
New
I am happy to introduce the very α version of the new programming language compiled to BEAM.
Welcome Cure.
It has literally three kille...
New
Hobbes is a low-level distributed database for the Elixir programming language.
Hobbes provides a simple, safe, and scalable storage lay...
New
Hi everyone!
The first release candidate for the Expert language server project is now available!
We’ve published a press release detai...
New
A little off-topic, but I feel like people here have a good head on their shoulders.
I used to be quite good at making software. Was luc...
New
Categories:
Sub Categories:
Forums
Popular Tags
- #ecto
- #liveview
- #troubleshooting
- #learning-elixir
- #library
- #deployment
- #erlang
- #testing
- #genserver
- #mix
- #absinthe
- #remote-other
- #otp
- #plug
- #how-to-question
- #macros
- #postgres
- #elixirconf
- #channels
- #exunit
- #discussion
- #code-sync
- #podcasts
- #javascript
- #onsite
- #dialyzer
- #docker
- #authentication
- #umbrella
- #full-time-contract
- #podcasts-by-brainlid
- #ai
- #ecto-query
- #elixirconf-us
- #blog-post
- #elixir-ls
- #phoenix_html
- #iex
- #graphql
- #genstage
- #websockets
- #supervisor
- #advent-of-code
- #distillery
- #processes
- #elixirconf-eu
- #api
- #forms
- #metaprogramming
- #hex











Showing Posts 1 to 5- Show Best Posts
- Show All (oldest first)
- Show All (newest first)
akoutmos
This conference talk by José Valim sounds like it may be of interest to you:
It goes from processing data in an eager way (reading everything up front), to a lazy way (leveraging Streams), to concurrent processing. Depending on how complex your data+processing is it sounds like Streams would be more than sufficient.
If you provide more context my opinion could change though
brettbeatty
You may be aware of it already, but Plataformatec created a library called NimbleCSV that I would look it. I’ve never used it on anything close to a 10 GB file, but it was pretty easy to use, and it supports streaming pretty well. I would think that as long as you handled it in appropriately-sized chunks you should be fine.
As far as scheduling the task goes, the only app-level crons I’ve done in Elixir used Quantum. I don’t remember it well enough to endorse, but I also don’t remember hating it.
tme_317
I process large CSVs also… In addition to
nimble_csvwhen I need to do the reading/writing in Elixir I also shell out from my Elixir app toxsvextensively for pre-processing… it’s written in Rust and super fast. Not sure what processing you need to do with the hourly files but it could speed you along.https://github.com/BurntSushi/xsv
Blokh
Thank you very much @brettbeatty @tme_317 and @BurntSushi.
You helped me a lot.
jeffyzooom
I know this is an old thread, but it still appears in searches for processing large CSV files. I maintain RustyCSV, a NimbleCSV-compatible parser and encoder using precompiled Rust NIFs. It supports bounded-memory streaming for files like this.
More details here: