leifg
Hey there,
a while ago I implemented an Excel Parser Library. So far it’s working but it has the problem that it doesn’t scale with file size.
Background: The (2000) Excel format is just a bunch of compressed (zip) XML files. In order to get access to the underlying XML, I extract the whole archive in memory:
:zip.extract spreadsheet_filename, [:memory]
This of course is very bad in terms of memory usage. In other languages I have found examples where a compressed archive is directly read as a stream, filtered by filename and then processed.
Is there a way to do this in the Erlang/Elixir world?
Trending in Questions
I’m working on a project that simulates the bumbl example in the programming phoenix book. It acts almost like an email client. We have a...
New
Hello,
I know there is an approach for handling lists that allows for optimized traversal, but I can’t recall the specific method (somet...
New
Documentation
While reading the Scoped Routes section, I noticed that the documentation currently refers to a problem without explainin...
New
Hi everyone,
I am toying with the idea of building a “match maker” for giving personal help to people that wants to start coding.
I sta...
New
So my question is quite simple and i have found no conclusive answer on forum, google or AI.
Should we use :erlang.float for Integer to ...
New
I recently noticed that Elixir’s Logger defaults its primary log level to :debug when no :logger, :level application configuration is pre...
New
I’m new to elixir and just tried to install the elixirLS extension for VScode(ium) and it is throwing some errors that I would like help ...
New
Other Trending Topics
Edit: 2026 May 15 - This post is archived.
Mob is alive!!
Main docs: mob v0.7.11 — Documentation
A bit of explanation for the slightly c...
New
Hey, I’m Jesse and I’m the main contributor behind Dexter, a full-featured, lightning-fast Elixir LSP optimized for large codebases. It s...
New
I am happy to introduce the very α version of the new programming language compiled to BEAM.
Welcome Cure.
It has literally three kille...
New
Hobbes is a low-level distributed database for the Elixir programming language.
Hobbes provides a simple, safe, and scalable storage lay...
New
Hi there! We created Gust: A task orchestrator inspired by Airflow.
For those who have never heard about Aiflow, it’s a Python-based wor...
New
Hi everyone!
The first release candidate for the Expert language server project is now available!
We’ve published a press release detai...
New
Categories:
Sub Categories:
Forums
Popular Tags
- #ecto
- #liveview
- #troubleshooting
- #learning-elixir
- #library
- #deployment
- #erlang
- #testing
- #genserver
- #mix
- #absinthe
- #remote-other
- #otp
- #plug
- #how-to-question
- #macros
- #postgres
- #elixirconf
- #channels
- #exunit
- #discussion
- #code-sync
- #podcasts
- #javascript
- #onsite
- #dialyzer
- #docker
- #authentication
- #umbrella
- #full-time-contract
- #podcasts-by-brainlid
- #ecto-query
- #ai
- #elixirconf-us
- #blog-post
- #elixir-ls
- #phoenix_html
- #iex
- #graphql
- #genstage
- #websockets
- #supervisor
- #advent-of-code
- #distillery
- #processes
- #api
- #forms
- #elixirconf-eu
- #metaprogramming
- #hex











Showing Posts 1 to 6- Show Best Posts
- Show All (oldest first)
- Show All (newest first)
bbense
Looking at the documentation for the :zip library , I think you can get what you want using the :zip.open interface. Open the handle with the :memory option, list all the files and then use :zip.get
to get each individual file.
You could wrap this with Stream.resource if you want to get all elixiry on it.
leifg
Thanks for the info. Played around with it but I’m not sure this is the way to go.
the
:zip.zip_getfunction always blocks and returns the whole content of the referenced file. What I actually need is a function that returns chunks of the content so that I can process it right away.Side Note: I peeked at the zip source code in OTP. It seems as if Erlang is reading the contents in chunks. So from research it seems to me that Erlang does not expose this API.
effinbanjos
Did you ever find a solution for this issue? I have a use-case where I’d need to read from a very large compressed file (large for me, that is…22G compressed / 65G+ uncompressed).
leifg
For this particular problem I had, I did not find a solution.
However I recently started another project and found this library: GitHub - ne-sachirou/stream_gzip: Gzip or gunzip an Elixir stream · GitHub
I’m not sure if it works with
.zipfiles.If it doesn’t and you have control over the compression used, you could stream process with that:
effinbanjos
Excellent, thanks
effinbanjos
A little follow-up many months later:
I finally had a little time to play with this and found that StreamGzip would die after a bit of processing. After casting about for other solutions, I found that the core File.stream! supports compressed files - you just need to feed it the right mode:
File.stream! also defaults to a line-by-line output mode instead of to a byte chunk output mode. I ended-up having to used chunk_while to re-order bytes chunks into line chunks with the StreamGzip library since I wanted to emit individual records into RabbitMQ.
HTH