preciz
I had a lot of back and forth with Codex CLI & GPT 5.6 Sol to use a SWAR optimization for URI.encode_www_form/1.
I now have a version that on my machine has no performance regression and has a 1.1X-7x speedup (depending on the input).
This topic is not really close to me and I don’t want to waste the Elixir maintainers time because I might make some trivial mistakes (like I did recently) so I will not submit it as a Pull Request, instead I think it would be great if a member of the community, who understands it much better than me would take this further. Attribution is not important for me, I just want the code to run faster.
Trending in Thoughts On...
I created an app that uses Phoenix LiveView and is deployed on Gigalixir. The app manages the state of multiple players (for example, a ...
New
Other Trending Topics
Edit: 2026 May 15 - This post is archived.
Mob is alive!!
Main docs: mob v0.7.11 — Documentation
A bit of explanation for the slightly c...
New
Hey, I’m Jesse and I’m the main contributor behind Dexter, a full-featured, lightning-fast Elixir LSP optimized for large codebases. It s...
New
I am happy to introduce the very α version of the new programming language compiled to BEAM.
Welcome Cure.
It has literally three kille...
New
Hobbes is a low-level distributed database for the Elixir programming language.
Hobbes provides a simple, safe, and scalable storage lay...
New
Hi everyone!
The first release candidate for the Expert language server project is now available!
We’ve published a press release detai...
New
A little off-topic, but I feel like people here have a good head on their shoulders.
I used to be quite good at making software. Was luc...
New
Categories:
Sub Categories:
Forums
Popular Tags
- #ecto
- #liveview
- #troubleshooting
- #learning-elixir
- #library
- #deployment
- #erlang
- #testing
- #genserver
- #mix
- #absinthe
- #remote-other
- #otp
- #plug
- #how-to-question
- #macros
- #postgres
- #elixirconf
- #channels
- #exunit
- #discussion
- #code-sync
- #podcasts
- #javascript
- #onsite
- #dialyzer
- #docker
- #authentication
- #umbrella
- #full-time-contract
- #podcasts-by-brainlid
- #ai
- #ecto-query
- #elixirconf-us
- #blog-post
- #elixir-ls
- #phoenix_html
- #iex
- #graphql
- #genstage
- #websockets
- #supervisor
- #advent-of-code
- #distillery
- #processes
- #elixirconf-eu
- #api
- #forms
- #metaprogramming
- #hex










Showing Posts 1 to 2- Show Best Posts
- Show All (oldest first)
- Show All (newest first)
videsnelson
Hello there!
Oh I believe it was me who some time ago took the first SWAR optimisation into OTP
:jsonmodule. So disclaimer, I think it was me who opened this pandora boxThis is hardcoded at compile time, so compiling on a 64bit and then running the BEAM files on a 32bit machine, or the other way around, would get the whole optimisation wrong. The weirdest edge case, but still worth keeping in mind. I believe I just hardcoded 7b ints in both OTP and Elixir, so actually maybe I even made 32bit machines a bit slower
I’d make that
8also an attribute with a name, something like@64bitor whatever, just to avoid magic numbers.Now something cooler, the original code. I checked the generated assembly, and actually, the biggest difference is in the loop through the binary. When the code is looping byte by byte, it is doing that snippet above, which compiles to an Erlang’s bitstring comprehension, with a function call, a new allocated binary (the output of
percent, and acase-docomparison, that the new code is not doing, and that is actually the largest performance saving and something you could write on the@swar_wordsize != 8 or byte_size(string) < @swar_thresholdclause as well and it would also improve massively. Because you didn’t try to do a for-loop with bigger binary slots but instead you pattern-matched those bigger binary slots into clauses, herethe function becomes first of all tail-recursive without intermediate function calls nor intermediate garbage.
When it comes to the SWAR code itself, well, all those indexes are the most confusing thing on earth so can’t say I’ve carefully checked right now, but might do if you put this into a PR!
encode_unreservedgets ugly with so many quite repetitive clauses, but I guess that’s the trade-off of an obsessive optimisation. Maybe the clauses browsing the bad byte can be autogenerated? Claude suggested me this:preciz
Thank you for your feedback, I see if I can make this better based on this.