csokun
What is the fastest way to split a large string into multiple chunks by size e.g breaking 10MB long string into multiple chunks of 5KB each?
Trending in Questions
I’m working on a project that simulates the bumbl example in the programming phoenix book. It acts almost like an email client. We have a...
New
Hello,
I know there is an approach for handling lists that allows for optimized traversal, but I can’t recall the specific method (somet...
New
Documentation
While reading the Scoped Routes section, I noticed that the documentation currently refers to a problem without explainin...
New
Hi everyone,
I am toying with the idea of building a “match maker” for giving personal help to people that wants to start coding.
I sta...
New
So my question is quite simple and i have found no conclusive answer on forum, google or AI.
Should we use :erlang.float for Integer to ...
New
I recently noticed that Elixir’s Logger defaults its primary log level to :debug when no :logger, :level application configuration is pre...
New
I’m new to elixir and just tried to install the elixirLS extension for VScode(ium) and it is throwing some errors that I would like help ...
New
Other Trending Topics
Edit: 2026 May 15 - This post is archived.
Mob is alive!!
Main docs: mob v0.7.11 — Documentation
A bit of explanation for the slightly c...
New
Hey, I’m Jesse and I’m the main contributor behind Dexter, a full-featured, lightning-fast Elixir LSP optimized for large codebases. It s...
New
I am happy to introduce the very α version of the new programming language compiled to BEAM.
Welcome Cure.
It has literally three kille...
New
Hobbes is a low-level distributed database for the Elixir programming language.
Hobbes provides a simple, safe, and scalable storage lay...
New
Hi there! We created Gust: A task orchestrator inspired by Airflow.
For those who have never heard about Aiflow, it’s a Python-based wor...
New
Hi everyone!
The first release candidate for the Expert language server project is now available!
We’ve published a press release detai...
New
Categories:
Sub Categories:
Forums
Popular Tags
- #ecto
- #liveview
- #troubleshooting
- #learning-elixir
- #library
- #deployment
- #erlang
- #testing
- #genserver
- #mix
- #absinthe
- #remote-other
- #otp
- #plug
- #how-to-question
- #macros
- #postgres
- #elixirconf
- #channels
- #exunit
- #discussion
- #code-sync
- #podcasts
- #javascript
- #onsite
- #dialyzer
- #docker
- #authentication
- #umbrella
- #full-time-contract
- #podcasts-by-brainlid
- #ecto-query
- #ai
- #elixirconf-us
- #blog-post
- #elixir-ls
- #phoenix_html
- #iex
- #graphql
- #genstage
- #websockets
- #supervisor
- #advent-of-code
- #distillery
- #processes
- #api
- #forms
- #elixirconf-eu
- #metaprogramming
- #hex










Showing Posts 1 to 10- Show Best Posts
- Show All (oldest first)
- Show All (newest first)
hauleth
Use binary pattern matching together with list comprehension:
dimitarvp
This could produce invalid Unicode strings though?
csokun
@hauleth I did a quick test and it seems like this method will throughway a leftover chunk.
csokun
I found an answer from StackOverflow which goes like this:
However, I not sure is there any performance implication here that I should take into consideration.
kelvinst
well, you could use
Stream.chunk_every, so it would lazily chunk your string instead of eagerly. For long strings like you said you are working with, it should improve performance a lot already.One question though, where are you getting this string from?
cmkarlsson
The stackoverflow is likely to be quite a bit slower.
handling utf8 is slow, but if not needed:
Here is a version based on the list comprehension but which takes the leftover into account/
And here is the benchee run:
Here is the result:
csokun
Hi @kelvinst I tried to port a method createFileFromText from
azure-storagenpm to my ex_azure_storage lib. So to be honest I don’t know where the string comes fromevadne
Random thought here but if how UTF-8 uses a variable number of bytes bothers you or causes issues, you could use UTF-32 which is fixed-length!
Alternatively you might have to implement something which is smart enough to know how many bytes each grapheme requires. You might want to check out elixir/lib/elixir/unicode/unicode.ex at v1.11.4 · elixir-lang/elixir · GitHub
Basically read a chunk of bytes and consume graphemes repeatedly off it (and emit chunks whenever you have one that is large enough), store the rest in the buffer and continue reading, repeat until done
csokun
@cmkarlsson impressive I’ve learnt something today thanks to you
kelvinst
I see, yeah, I was just curious cause normally big amounts of data like that come from files or uploads, that could be streamed themselves, avoiding to load the whole thing to memory. But in your case, as it’s a ported lib function that takes a loaded string, you don’t much control over that, so yeah, binary pattern matching seems like the way to go then.