csokun
How to split string into multiple chunks by size
What is the fastest way to split a large string into multiple chunks by size e.g breaking 10MB long string into multiple chunks of 5KB each?
Most Liked
hauleth
Use binary pattern matching together with list comprehension:
for <<chunk::size(chunk_size)-binary <- input>>, do: chunk
bdarla
A relevant library that came to my attention recently is text_chunker_ex by Revelry (thanks to @hugobarauna for sharing this in Elixir Radar 420).
The aim of the library is to split text to be fed to AI models. As I understood, it mimics some functionality of the LangChain. I have not used it (yet), but seems interesting!
cmkarlsson
The stackoverflow is likely to be quite a bit slower.
handling utf8 is slow, but if not needed:
Here is a version based on the list comprehension but which takes the leftover into account/
defmodule Chunker do
alias Chunker
def chunk(string, size \\ 5), do: chunk(string, size, [])
defp chunk(<<>>, size, acc), do: Enum.reverse(acc)
defp chunk(string, size, acc) when byte_size(string) > size do
<<c::size(size)-binary, rest::binary>> = string
chunk(rest, size, [c | acc])
end
defp chunk(leftover, size, acc) do
chunk(<<>>, size, [leftover | acc])
end
def stackoverflow(string, size \\ 5) do
string
|> String.codepoints
|> Enum.chunk_every(size)
|> Enum.map(&Enum.join/1)
end
def withstream(string, size \\ 5) do
string
|> String.codepoints
|> Stream.chunk_every(size)
|> Enum.map(&Enum.join/1)
end
end
And here is the benchee run:
{:ok, data} = File.read("./data/alice29.txt")
Benchee.run(
%{
"chunk" => fn -> Chunker.chunk(data, 5000) end,
"overflow" => fn -> Chunker.stackoverflow(data, 5000) end
})
Here is the result:
Operating System: Linux
CPU Information: Intel(R) Xeon(R) CPU E3-1245 v3 @ 3.40GHz
Number of Available Cores: 8
Available memory: 15.55 GB
Elixir 1.10.4
Erlang 22.3
Benchmark suite executing with the following configuration:
warmup: 2 s
time: 5 s
memory time: 0 ns
parallel: 1
inputs: none specified
Estimated total run time: 21 s
Benchmarking chunk...
Benchmarking overflow...
Benchmarking withstream...
Name ips average deviation median 99th %
chunk 586726.39 0.00170 ms ±1225.64% 0.00145 ms 0.00350 ms
withstream 17.50 57.15 ms ±14.33% 54.27 ms 81.87 ms
overflow 15.71 63.64 ms ±16.49% 61.00 ms 94.93 ms
Comparison:
chunk 586726.39
withstream 17.50 - 33528.78x slower +57.14 ms
overflow 15.71 - 37339.36x slower +63.64 ms
Last Post!
jpc-ae
This is a bit of an old one, but here’s my version that isn’t as memory-friendly as streaming, but it is unicode-safe and better than splitting+joining graphemes. Overall its quite performant:
Regex.split(~r/(.{5000}[\p{Mn}\p{Me}]*|.+)/su, string, include_captures: true, trim: true)
It’ll return a list of strings chunked into, in this case, 5000 characters (plus any combining marks to ensure they don’t get split up from what they’re attached to), with the remaining bit at the end. This won’t be exactly 5kb depending on how many multi-byte characters you have, but should do fairly well for most purposes. You can always adjust the match string to better reflect the input data structure.
Popular in Questions
Other popular topics
Categories:
Sub Categories:
Forums
Popular Tags
- #ecto
- #liveview
- #troubleshooting
- #learning-elixir
- #deployment
- #library
- #erlang
- #testing
- #genserver
- #mix
- #absinthe
- #remote-other
- #otp
- #plug
- #how-to-question
- #macros
- #postgres
- #channels
- #elixirconf
- #exunit
- #discussion
- #code-sync
- #javascript
- #podcasts
- #onsite
- #dialyzer
- #docker
- #authentication
- #umbrella
- #full-time-contract
- #podcasts-by-brainlid
- #ecto-query
- #elixir-ls
- #phoenix_html
- #iex
- #blog-post
- #graphql
- #genstage
- #ai
- #websockets
- #supervisor
- #elixirconf-us
- #advent-of-code
- #distillery
- #processes
- #api
- #forms
- #metaprogramming
- #security
- #hex










