mudasobwa
Creator of Cure
AFAIU, Elixir delegates regular expression evaluation to Erlang.
Erlang claims to support Unicode general category matchers.
“ is declared in Unicode spec as LEFT DOUBLE QUOTATION MARK under General punctuation. Both Ruby and Perl do recognize this symbol as opening punctuation:
[0x201C].pack('U*').match /\p{Pi}/
#⇒ #<MatchData "“">
Both Elixir and Erlang, unfortunately, do not:
Regex.scan(~r/[\p{Pi}\p{Pf}\p{Ps}\p{Pe}]/, "'\"“”‘’«»")
#⇒ [[<<171>>], [<<187>>]] # these are « and »
What am I missing and/or what should I tune to make the regex engine to work properly?
Trending in Questions
I having some trouble figuring out if I have set myself too strict of standards for my production server. Currently I can handle 75% of r...
New
Documentation
While reading the Scoped Routes section, I noticed that the documentation currently refers to a problem without explainin...
New
Hello,
I’m trying to build a basic Phoenix web-app, and I’d like to use Tailwind.
However, when I launch mix phx.server, I get an error...
New
Hi everyone,
I am toying with the idea of building a “match maker” for giving personal help to people that wants to start coding.
I sta...
New
So my question is quite simple and i have found no conclusive answer on forum, google or AI.
Should we use :erlang.float for Integer to ...
New
I recently noticed that Elixir’s Logger defaults its primary log level to :debug when no :logger, :level application configuration is pre...
New
So i have been using ash framework for a while and i love it. However currently the issue im having with ash framework is the error handl...
New
Other Trending Topics
Edit: 2026 May 15 - This post is archived.
Mob is alive!!
Main docs: mob v0.7.11 — Documentation
A bit of explanation for the slightly c...
New
Hey, I’m Jesse and I’m the main contributor behind Dexter, a full-featured, lightning-fast Elixir LSP optimized for large codebases. It s...
New
I am happy to introduce the very α version of the new programming language compiled to BEAM.
Welcome Cure.
It has literally three kille...
New
Hobbes is a low-level distributed database for the Elixir programming language.
Hobbes provides a simple, safe, and scalable storage lay...
New
Hi everyone!
The first release candidate for the Expert language server project is now available!
We’ve published a press release detai...
New
A little off-topic, but I feel like people here have a good head on their shoulders.
I used to be quite good at making software. Was luc...
New
Categories:
Sub Categories:
Forums
Popular Tags
- #ecto
- #liveview
- #troubleshooting
- #learning-elixir
- #library
- #deployment
- #erlang
- #testing
- #genserver
- #mix
- #absinthe
- #remote-other
- #otp
- #plug
- #how-to-question
- #macros
- #postgres
- #elixirconf
- #channels
- #exunit
- #discussion
- #code-sync
- #podcasts
- #javascript
- #onsite
- #dialyzer
- #docker
- #authentication
- #umbrella
- #full-time-contract
- #podcasts-by-brainlid
- #ai
- #ecto-query
- #elixirconf-us
- #blog-post
- #elixir-ls
- #phoenix_html
- #iex
- #graphql
- #genstage
- #websockets
- #supervisor
- #advent-of-code
- #distillery
- #processes
- #elixirconf-eu
- #api
- #forms
- #metaprogramming
- #hex










Showing Posts 1 to 10- Show Best Posts
- Show All (oldest first)
- Show All (newest first)
NobbZ
You forgot to enable unicode mode:
mudasobwa
Thanks! I wonder how come it’s disabled by default.
NobbZ
Yeah, thats a thing I do not understand as well. Elixir has very good unicode support, but having to remember to explicitely enable it for a regex everytime is a bit…
OvermindDL1
Because parsing unicode is muuuuuuch slower than parsing ascii in a few common cases, and if you know you don’t need direct unicode matching then there is no need.
mudasobwa
Nah. People who do parse looooooong texts with regular expressions or require a μs speed-up on parsing the natural string consisting of a dozen of characters, are indeed aware of all the flags.
The gain for, say, parsing [pun intended] emails, is negligible though. For the sake of saving keystrokes (I could buy new RAM stick for every 100K “u” typed) the switch should be turned on by default IMHO.
OvermindDL1
Except you can’t always know if they are parsing a hundred thousand emails for example. But still, if you think it should be on by default, make up a proposal to the elixir core mailing list after a forum thread dedicated to it first.
mudasobwa
Fair enough.
This is a matter of personal preference and I cordially avoid wasting core team time with such requests
OvermindDL1
You can always make your own delegating
sigil_rmacro that auto-adds the unicode tag too and import it where you need.NobbZ
If one does this, I’d be happy if he did not use
sigil_r, it might be confusing when copy pasting into a module that does not have that import… Perhapssigil_ufor unicode?mudasobwa
C’mon
sigil_ฤ.