stefanluptak
Hi all!
I would like to ask for an ideas how to solve the different way JavaScript and Elixir are approaching Strings and indexes of characters.
I want to store the same (user entered) text on the Elixir and also the JavaScript side. I also want to exchange messages describing operations like {:delete, from, to}, {:insert, position, text}, etc. to ensure the same text is on both sides.
The thing is, that this becomes intersting as soon as the text contains some emojis, characters with puncutation and so on.
Working with strings (binaries) in Elixir is pretty straightforward. Unfortunately (but not surprisingly
), the JavaScript behavior is (in my opinion) a bit unintuitive.
// JavaScript
> String.fromCharCode(97, 769).slice(0, 1)
'a'
> String.fromCharCode(225).slice(0, 1)
'á'
# Elixir
iex> [97, 769] |> to_string() |> String.slice(0, 1)
"á"
iex> [225] |> to_string() |> String.slice(0, 1)
"á"
What strategy should I use? Should I work with the text as a charlist on the Elixir side or should I normalize all strings everywhere? Or is there a better strategy I should take a look at?
Thank you all for you advices.
Trending in Questions
Other Trending Topics
Categories:
Sub Categories:
Forums
Popular Tags
- #ecto
- #liveview
- #troubleshooting
- #learning-elixir
- #library
- #deployment
- #erlang
- #testing
- #genserver
- #mix
- #absinthe
- #remote-other
- #otp
- #plug
- #how-to-question
- #macros
- #postgres
- #elixirconf
- #channels
- #exunit
- #discussion
- #code-sync
- #podcasts
- #javascript
- #onsite
- #dialyzer
- #docker
- #authentication
- #umbrella
- #full-time-contract
- #podcasts-by-brainlid
- #ecto-query
- #elixirconf-us
- #ai
- #blog-post
- #elixir-ls
- #phoenix_html
- #iex
- #graphql
- #genstage
- #websockets
- #supervisor
- #advent-of-code
- #distillery
- #processes
- #api
- #forms
- #hex
- #security
- #metaprogramming










Showing Posts 14 to 5- Show Best Posts
- Show All (oldest first)
- Show All (newest first)
stefanluptak
Ah OK. Thanks a lot. So if I understand this correctly, charlist is a standard Elixir
Listcontainer implemented as linked list in this case containing integers that represent code points andStringis piece of memory and to “iterate” over it’s content we just need to increment the memory address. Am I right?axelson
Charlists are lists which are implemented as linked lists. And to iterate through the linked list you need to follow a bunch of pointers which is O(n).
stefanluptak
That make sense. I thought that charlist is in fact the same, so I don’t understand the speed difference.

I just started watching Johanna’s talk about string processing performance. I hope I will learn something new and useful.
hauleth
Binaries are in reality just byte arrays and
binary_partis justmake_binary(old + start, length), so I highly doubt that you can get any faster than that.stefanluptak
Yes, I meant that one. Sorry for not being clear enough. OK, good to know. Thanks.
benwilson512
Do what, binary_part? binary_part is about as fast as it gets on the BEAM for that specific operation I think.
stefanluptak
I tried to do a little benchmark here and I am quite surprised, that doing
binary_part(binary, from, length)is 46x faster thanEnum.slice(charlist, from, to)Of course doing
String.slice(string, from, to)is extremely slow. That’s not surprising.Do you have some tips to do that even faster?
stefanluptak
Thanks a lot. Now I understand.
hauleth
Operate on bytes it is the safest way, so you do not use
Stringmodule in short. This will provide you independence form encoding of your data. So it would look like thisAnd then in Elixir:
Alternatively use LSEQ mentioned earlier which is representation independent (as it generates it’s own indices instead of using string positions).
NobbZ
Do you want to be
aand^be considered as separate or asâ?If the latter, work on graphemes.
Read docs of string functions carefully to know if they work on graphemes or codepoints.
The same is true for the JS functions and methods you use.
Though I have to disappoint you. It will be complicated, no matter what. Most developers of string handling libraries either do not care or even understand the differences.
Anyway, try to avoid random access of strings by codepoint or grapheme, it’s O(n) operation!