BradS2S
Looking over how String.valid? implements :fast_ascii option, what’s the significance is pattern matching in groups of 7 bytes?
when Bitwise.band(0x80808080808080, a) == 0 (source)
I think I understand using Bitwise.band/2 to check for the value of the leading bit, but since all ascii characters are 1 byte, what does checking in groups of 7 bytes do?
Does it have to do with CPU design?
Trending in Questions
I’m working on a project that simulates the bumbl example in the programming phoenix book. It acts almost like an email client. We have a...
New
Hello,
I know there is an approach for handling lists that allows for optimized traversal, but I can’t recall the specific method (somet...
New
I’m seeing that a list inside a Kino.DataTable will be interpreted as a charlist, even if the Kino.configure() is set to charlists: :as_l...
New
So my question is quite simple and i have found no conclusive answer on forum, google or AI.
Should we use :erlang.float for Integer to ...
New
Hi, I’ve just set up an application with ash_authentication. There is only magic link strategy for now, so there is no confirmation add o...
New
If a change or preparation module uses Ash.Changeset.get_argument/2 or Ash.Query.get_argument/2 (or any of the other get_argument functio...
New
Documentation
While reading the Scoped Routes section, I noticed that the documentation currently refers to a problem without explainin...
New
Other Trending Topics
I am happy to introduce the very α version of the new programming language compiled to BEAM.
Welcome Cure.
It has literally three kille...
New
Hi there! We created Gust: A task orchestrator inspired by Airflow.
For those who have never heard about Aiflow, it’s a Python-based wor...
New
Beam Bots (or just BB for short) is a framework for building fault-tolerant robotics applications in Elixir using familiar OTP patterns. ...
New
Xamal is a deployment tool for Elixir apps that deploys native releases to bare metal servers over SSH. It’s a port of GitHub - basecamp/...
New
Corex is an accessible, unstyled UI component library for Phoenix that integrates Zag.js state machines using Vanilla JavaScript and Live...
New
Aludel - LLM Evaluation Workbench
Aludel is an embeddable Phoenix LiveView dashboard for evaluating and comparing LLM prompts across mult...
New
Categories:
Sub Categories:
Forums
Popular Tags
- #ecto
- #liveview
- #troubleshooting
- #learning-elixir
- #library
- #deployment
- #erlang
- #testing
- #genserver
- #mix
- #absinthe
- #remote-other
- #otp
- #plug
- #how-to-question
- #macros
- #postgres
- #elixirconf
- #channels
- #exunit
- #discussion
- #code-sync
- #podcasts
- #javascript
- #onsite
- #dialyzer
- #docker
- #authentication
- #umbrella
- #full-time-contract
- #podcasts-by-brainlid
- #ecto-query
- #ai
- #elixirconf-us
- #blog-post
- #elixir-ls
- #phoenix_html
- #iex
- #graphql
- #genstage
- #websockets
- #supervisor
- #advent-of-code
- #distillery
- #processes
- #api
- #forms
- #hex
- #security
- #metaprogramming










Showing Posts 1 to 9- Show Best Posts
- Show All (oldest first)
- Show All (newest first)
adw632
On 64bit systems a word size on the BEAM is 8 bytes. Now because the BEAM uses some bits to indicate the type, and binaries are boxed types, the actual available payloa in the first word is 7 bytes, but beyond that it ideally should be 8 bytes at a time but the constant used for the AND limits each batch to 7 bytes.
Doing computation a word at a time is efficient vs byte at a time. Going beyond this using an 8 byte constant in a BIF and even using SSE instructions would be even faster so it appears there is opportunity for further optimisation.
adw632
Further to this there is an article that covers the fast decoding in terms of cycles per known. Validating UTF-8 bytes using only 0.45 cycles per byte (AVX edition) – Daniel Lemire's blog
There is also the fastest JSON parser here that can validate UTF-8 and ascii at 13GB/s. The JSON processing is 25x faster than the fastest known C++ implementation.
I would say base on those numbers there would be room to improve Elixir/Erlang performance.
BradS2S
Is there an agreed upon best practice for doing the first iteration differently?
For example my understanding is that it’s better to use atoms besides true and false.
BradS2S
also would
:erlang.system_info(:wordsize)work to check if BEAM was running on 32 vs. 64 ?Thanks again for the great answer.
al2o3cr
The issue is the “extract the first X bits into
a” part not the representation of the binary, so subsequent iterations should be the same size. Additional discussion:https://github.com/elixir-lang/elixir/pull/12354#issuecomment-1398821451
BradS2S
read through that thread. This gives me all the context I was looking for plus some. Thanks!
mtrudel
Author of this implementation here.
We can easily test for the ‘common’ case of single-byte UTF-8 code points (ie: ASCII) by looking at the high bit of each byte to see if it’s clear by ANDing it with 0x80 (ie: b10000000). For continuous runs of such bytes, we can also check if ALL of the bytes’ high bits are clear by ANDing them with 0x80808080… in a similar manner.
On a 64 bit system, you would expect the optimal stride value to be 8 bytes. However, experimentation doesn’t seem to bear that out, and 7 byte strides seem to be optimal on ARM and x64 platforms (my hunch is that the VM does tagging with some of the higher bits in a word). Details of all this is in the PR thread that got linked above.
hauleth
It does. 4 bits of every integer are reserved for tag.
BradS2S
That was a ton of work and it made for a very interesting read. Can’t beat empirical data.