csokun

csokun

What is the fastest way to split a large string into multiple chunks by size e.g breaking 10MB long string into multiple chunks of 5KB each?

Showing Posts 1 to 10

hauleth

hauleth

Use binary pattern matching together with list comprehension:

for <<chunk::size(chunk_size)-binary <- input>>, do: chunk
dimitarvp

dimitarvp

This could produce invalid Unicode strings though?

csokun

csokun OP

@hauleth I did a quick test and it seems like this method will throughway a leftover chunk.

for <<chunk::size(5)-binary ← “hello world”>>, do: chunk
[“hello”, " worl"]

csokun

csokun OP

I found an answer from StackOverflow which goes like this:

 "hello world" 
|> String.codepoints
|> Enum.chunk_every(5)
|> Enum.map(&Enum.join/1)
# ["hello", " worl", "d"]

However, I not sure is there any performance implication here that I should take into consideration.

kelvinst

kelvinst

well, you could use Stream.chunk_every, so it would lazily chunk your string instead of eagerly. For long strings like you said you are working with, it should improve performance a lot already.

One question though, where are you getting this string from?

cmkarlsson

cmkarlsson

The stackoverflow is likely to be quite a bit slower.

handling utf8 is slow, but if not needed:

Here is a version based on the list comprehension but which takes the leftover into account/

defmodule Chunker do

  alias Chunker

  def chunk(string, size \\ 5), do: chunk(string, size, [])

  defp chunk(<<>>, size, acc), do: Enum.reverse(acc)
  defp chunk(string, size, acc) when byte_size(string) > size do
    <<c::size(size)-binary, rest::binary>> = string
    chunk(rest, size, [c | acc])
  end
  defp chunk(leftover, size, acc) do
    chunk(<<>>, size, [leftover | acc])
  end


  def stackoverflow(string, size \\ 5) do
    string
    |> String.codepoints
    |> Enum.chunk_every(size)
    |> Enum.map(&Enum.join/1)
  end

  def withstream(string, size \\ 5) do
    string
    |> String.codepoints
    |> Stream.chunk_every(size)
    |> Enum.map(&Enum.join/1)
  end

end

And here is the benchee run:

{:ok, data} = File.read("./data/alice29.txt")
Benchee.run(
  %{
    "chunk" => fn -> Chunker.chunk(data, 5000) end,
    "overflow" => fn -> Chunker.stackoverflow(data, 5000) end
  })

Here is the result:



Operating System: Linux
CPU Information: Intel(R) Xeon(R) CPU E3-1245 v3 @ 3.40GHz
Number of Available Cores: 8
Available memory: 15.55 GB
Elixir 1.10.4
Erlang 22.3

Benchmark suite executing with the following configuration:
warmup: 2 s
time: 5 s
memory time: 0 ns
parallel: 1
inputs: none specified
Estimated total run time: 21 s

Benchmarking chunk...
Benchmarking overflow...
Benchmarking withstream...

Name                 ips        average  deviation         median         99th %
chunk          586726.39     0.00170 ms  ±1225.64%     0.00145 ms     0.00350 ms
withstream         17.50       57.15 ms    ±14.33%       54.27 ms       81.87 ms
overflow           15.71       63.64 ms    ±16.49%       61.00 ms       94.93 ms

Comparison:
chunk          586726.39
withstream         17.50 - 33528.78x slower +57.14 ms
overflow           15.71 - 37339.36x slower +63.64 ms
csokun

csokun OP

Hi @kelvinst I tried to port a method createFileFromText from azure-storage npm to my ex_azure_storage lib. So to be honest I don’t know where the string comes from :slight_smile:

evadne

evadne

Random thought here but if how UTF-8 uses a variable number of bytes bothers you or causes issues, you could use UTF-32 which is fixed-length!

Alternatively you might have to implement something which is smart enough to know how many bytes each grapheme requires. You might want to check out elixir/lib/elixir/unicode/unicode.ex at v1.11.4 · elixir-lang/elixir · GitHub

Basically read a chunk of bytes and consume graphemes repeatedly off it (and emit chunks whenever you have one that is large enough), store the rest in the buffer and continue reading, repeat until done

csokun

csokun OP

@cmkarlsson impressive I’ve learnt something today thanks to you :smiley:

kelvinst

kelvinst

I see, yeah, I was just curious cause normally big amounts of data like that come from files or uploads, that could be streamed themselves, avoiding to load the whole thing to memory. But in your case, as it’s a ported lib function that takes a loaded string, you don’t much control over that, so yeah, binary pattern matching seems like the way to go then.

Where Next? Top

Trending in Questions Top

stjefim
Hello! Suppose you are building workflow (order / task / payment) processing system with the following requirements: Each workflow con...
New
Blokh
Hey guys, I’ve got a huge CSV ( around 10 GB ) that needs to be processed hourly Do you guys have any suggestions what is the best prac...
New
roeland
Kia ora, We have been using elixir-google-api to connect to Google Drive. However, with the updates to Tesla due to CVEs this is now bro...
New
kszambelanczyk
Hello! Could someone please give me a help/sample code, how to delete a file from s3 using waffle/waffle_ecto from Phoenix app. I creat...
New
Onor.io
I have what I’ve heard referred to as a “lookup table” in my database. This is a way of assigning codes to common values. One common lo...
New
jaybe78
Hello, I’m developing a online persistent chat system (what’s app) like using elixir/dynamodb/aws for a mobile app(flutter). The diffic...
New
Trolleger
What approach to take when sending live updates to “random” users Hi! I have a question, I have a little chat app, and when I create a DM...
New

Other Trending Topics Top

garrison
Hobbes is a low-level distributed database for the Elixir programming language. Hobbes provides a simple, safe, and scalable storage lay...
New
jimsynz
Beam Bots (or just BB for short) is a framework for building fault-tolerant robotics applications in Elixir using familiar OTP patterns. ...
New
mcass19
ExRatatui lets you cook up rich terminal UIs in Elixir, powered by Rust’s ratatui via Rustler NIFs. Build interactive terminal applicatio...
New
Damirados
Hello everyone. After busy few months I am happy to announce v0.1.0 of Emerge &amp; Solve. They are GUI (Emerge) and State management (S...
New
netoum
Corex is an accessible, unstyled UI component library for Phoenix that integrates Zag.js state machines using Vanilla JavaScript and Live...
New
wintermeyer
There are three potential reasons for members of this forum to have a look at https://vutuv.de You are tired or annoyed of LinkedIn. Yo...
New

We're in Beta

About us Mission Statement

Options

Thread Display Mode




Thread Preview

Skip Thread Previews