stefanluptak

stefanluptak

Hi all!

I would like to ask for an ideas how to solve the different way JavaScript and Elixir are approaching Strings and indexes of characters.

I want to store the same (user entered) text on the Elixir and also the JavaScript side. I also want to exchange messages describing operations like {:delete, from, to}, {:insert, position, text}, etc. to ensure the same text is on both sides.

The thing is, that this becomes intersting as soon as the text contains some emojis, characters with puncutation and so on.

Working with strings (binaries) in Elixir is pretty straightforward. Unfortunately (but not surprisingly :slightly_smiling_face:), the JavaScript behavior is (in my opinion) a bit unintuitive.

// JavaScript
> String.fromCharCode(97, 769).slice(0, 1)
'a'
> String.fromCharCode(225).slice(0, 1)
'á'
# Elixir
iex> [97, 769] |> to_string() |> String.slice(0, 1)
"á"
iex> [225] |> to_string() |> String.slice(0, 1)    
"á"

What strategy should I use? Should I work with the text as a charlist on the Elixir side or should I normalize all strings everywhere? Or is there a better strategy I should take a look at?

Thank you all for you advices.

Showing Posts 1 to 10

hauleth

hauleth

Unicode is hard.
Unicode is very hard.
Unicode is enormously hard.


If you want to have consistent behaviour then operate on bytes. So you will use Blob or Uint8Array on JS side and binaries on Elixir side. I think it would be the simplest way. I hope that what you are describing (as I assume you want to created distributed concurrent text editor) is CRDT like LSEQ.

stefanluptak

stefanluptak OP

Thank you for your reply @hauleth. I am not sure I understand your solution. I will try to write pseudo-code example:

// JavaScript message to Elixir, somebody wrote "áb"
{:insert, 0, String.fromCharCode(97, 769, 98)} 
// JavaScript message to Elixir, somebody deleted "á" (but not b)
{:delete, 0, 2}

This will delete the “b” too, because Elixir is considering “á” to be 1 character and doesn’t care how many codepoints it has:

text_from_javascript = to_string([97, 769, 98])
deleted_text = String.slice(text, 0, 2)
# deletes "áb"

From my point of view, it might be safe (on Elixir side) to convert all the strings to charlists and then use the List / Enum operations. Am I wrong?

P.S.: Yes, the concurrent/distributed operations are handled with CRDT. It’s just this unicode stuff I am trying to solve.

NobbZ

NobbZ

The big question is, do you want to work on “graphemes” or on “codepoints”? If the latter, do you do normalize first?

The Javascript seems to operate on “codepoints”, though the current normalisation is unknown, if it normalises at all, instead of just taking what it gets from the operating system.

stefanluptak

stefanluptak OP

:slight_smile: I want to work on whatever is simpler and more consistent or safe respectively. I don’t do any normalization (yet?). But maybe it will be necessary. I am not sure. That’s why I am asking. I don’t want to overcomplicate it.

NobbZ

NobbZ

Do you want to be a and ^ be considered as separate or as â?

If the latter, work on graphemes.

Read docs of string functions carefully to know if they work on graphemes or codepoints.

The same is true for the JS functions and methods you use.

Though I have to disappoint you. It will be complicated, no matter what. Most developers of string handling libraries either do not care or even understand the differences.

Anyway, try to avoid random access of strings by codepoint or grapheme, it’s O(n) operation!

hauleth

hauleth

Operate on bytes it is the safest way, so you do not use String module in short. This will provide you independence form encoding of your data. So it would look like this

// JavaScript message to Elixir, somebody wrote "áb"
["insert", 0, Uint8Array.of(97, 204, 129, 98)]
// JavaScript message to Elixir, somebody deleted "á" (but not b)
["delete", 0, 3]

And then in Elixir:

text_from_javascript = <<97, 204, 129, 98>>
deleted_text = binary_part(text_from_javascript, 0, 3)

Alternatively use LSEQ mentioned earlier which is representation independent (as it generates it’s own indices instead of using string positions).

stefanluptak

stefanluptak OP

Thanks a lot. Now I understand. :+1:

stefanluptak

stefanluptak OP

I tried to do a little benchmark here and I am quite surprised, that doing binary_part(binary, from, length) is 46x faster than Enum.slice(charlist, from, to)
Of course doing String.slice(string, from, to) is extremely slow. That’s not surprising.
Do you have some tips to do that even faster?

benwilson512

benwilson512

Author of Craft GraphQL APIs in Elixir with Absinthe

Do what, binary_part? binary_part is about as fast as it gets on the BEAM for that specific operation I think.

stefanluptak

stefanluptak OP

Yes, I meant that one. Sorry for not being clear enough. OK, good to know. Thanks. :ok_hand:

Where Next? Top

Trending in Questions Top

katta
I having some trouble figuring out if I have set myself too strict of standards for my production server. Currently I can handle 75% of r...
New
brecabral
Documentation While reading the Scoped Routes section, I noticed that the documentation currently refers to a problem without explainin...
New
achenet
Hello, I’m trying to build a basic Phoenix web-app, and I’d like to use Tailwind. However, when I launch mix phx.server, I get an error...
New
kpanic
Hi everyone, I am toying with the idea of building a “match maker” for giving personal help to people that wants to start coding. I sta...
New
asweet-confluent
I recently noticed that Elixir’s Logger defaults its primary log level to :debug when no :logger, :level application configuration is pre...
New
Cxx-mlr
I’m working on a small exercise involving update_in/3, and I came up with this solution: data = %{ name: "Periodic Table", category:...
New
ChrisAmelia
I’ve got trouble wrapping my head around the order in which functions are called in this snippet (from Phoenix’s authentication): toke...
New

Other Trending Topics Top

GenericJam
Edit: 2026 May 15 - This post is archived. Mob is alive!! Main docs: mob v0.7.11 — Documentation A bit of explanation for the slightly c...
New
JesseHerrick
Hey, I’m Jesse and I’m the main contributor behind Dexter, a full-featured, lightning-fast Elixir LSP optimized for large codebases. It s...
New
mudasobwa
I am happy to introduce the very α version of the new programming language compiled to BEAM. Welcome Cure. It has literally three kille...
New
garrison
Hobbes is a low-level distributed database for the Elixir programming language. Hobbes provides a simple, safe, and scalable storage lay...
New
mhanberg
Hi everyone! The first release candidate for the Expert language server project is now available! We’ve published a press release detai...
New
budgie
A little off-topic, but I feel like people here have a good head on their shoulders. I used to be quite good at making software. Was luc...
New

We're in Beta

About us Mission Statement

Options

Thread Display Mode




Thread Preview

Skip Thread Previews