scouten
Unicode-savvy `letter_or_digit?` function?
I’m parsing a file format by recursing on a charlist. One of the challenges I have in parsing this format is determining whether the current Unicode codepoint is a letter or digit. So far I’ve written this:
# HELP: This is not Unicode-savvy. Is there such a thing?
defp letter_or_digit?(c) when c >= ?0 and c <= ?9, do: true
defp letter_or_digit?(c) when c >= ?A and c <= ?Z, do: true
defp letter_or_digit?(c) when c >= ?a and c <= ?z, do: true
defp letter_or_digit?(_), do: false
This is called via Enum.split_while(&letter_or_digit?/1) and similar constructions.
What I’ve written is obviously not savvy about non-ASCII letters or digits. Is there such a thing handy? My Google-fu has failed me so far.
The obvious answer would be to parse via Regex match. Unfortunately, some other parts of the file format are far better consumed via charlist recursion and I don’t want to pay the cost of bouncing back and forth between strings and charlists.
First Post!
hauleth
- You absolutely can run regexes on charlists as in Erlang charlists are the strings.
- If you want to check if code point is digit, then here you have full list, you need to check all these groups. For rest of the character groups check here.
BTW you do not need to change binary (as assume that you get binary from the file) to iterate over it, you can pattern match on binary fragments as well by <<byte, rest :: binary>> = "abba", byte == ?a and rest == "bba".
Most Liked
Nicd
Try Erlang regular expressions for charlists instead, I think this is what hauleth meant: re — OTP 29.0.2 (stdlib 8.0.1)
You probably want to use regex to match for character property Number.
kip
I’ve just published the first version of ex_cldr_unicode that might be helpful to you. It builds functions at compile time using data from the Unicode database.
Theres a bunch of fun stuff you can do, but to your use case it includes some guards that you may find useful. For example:
defmodule MyModule do
require Cldr.Unicode.Guards
alias Cldr.Unicode.Guards
def my_function(codepoint) when Guards.is_upper(codepoint) do
IO.puts "Its Uppercase!"
end
See hex docs for further info.
I’ve defined only the following guards so far but it’s trivial to add more so let me know if its useful:
is_upperis_loweris_digitis_currency_symbol
Oh, and it’s more than twice as fast as using a regex for this kind of matching.
There are a bunch of classifier functions as well. From my understanding of your use case the following may also apply:
iex> Cldr.Unicode.Property.alphanumeric? "1234"
true
iex> Cldr.Unicode.Property.alphanumeric? "KeyserSöze1995"
true
iex> Cldr.Unicode.Property.alphanumeric? "3段"
true
iex> Cldr.Unicode.Property.alphanumeric? "dragon@example.com"
false
kip
Updated to version 0.2.0. Main changes are:
-
Moves the public API to the
Cldr.Unicodemodule. -
Updates and adds documentation to all public functions.
-
Removes the text annotations from the compiled functions which materially reduces the size of the beam files.
Feedback welcome as are feature requests and PRs.
Popular in Questions
Other popular topics
Categories:
Sub Categories:
Forums
Popular Tags
- #ecto
- #liveview
- #troubleshooting
- #learning-elixir
- #deployment
- #library
- #erlang
- #testing
- #genserver
- #mix
- #absinthe
- #remote-other
- #otp
- #plug
- #how-to-question
- #macros
- #postgres
- #channels
- #elixirconf
- #exunit
- #discussion
- #code-sync
- #javascript
- #podcasts
- #onsite
- #dialyzer
- #docker
- #authentication
- #umbrella
- #full-time-contract
- #podcasts-by-brainlid
- #ecto-query
- #elixir-ls
- #phoenix_html
- #iex
- #blog-post
- #graphql
- #genstage
- #ai
- #websockets
- #supervisor
- #elixirconf-us
- #advent-of-code
- #distillery
- #processes
- #api
- #forms
- #metaprogramming
- #security
- #hex









