kip
ex_cldr Core Team
I’ll shortly be launching Text, a nascent text analysis library.
Current functionality
In this early version (not ready for prime time) it includes:
- Word counting
- N-gram generation
- Language detection (of about 250 languages with pluggable vocabularies and pluggable correlation models)
- An English inflector (singular to plural) using a non-regex algorithmic approach
Future functionality
-
A language stemmer - as soon as I finishing writing the snowball compiler
-
Parts of speech tagger
Collaboration encouraged
-
Contributions in all areas are most welcome
-
Non-english speakers who would like to contribute to non-english inflectors are particularly welcome
Next steps
After some polishing this weekend I will publish a version to hex.
Trending in Announcing
WebAuthnLiveComponent WebAuthnComponents
See this post about renaming the package.
Passwordless authentication for Phoenix LiveView app...
New
Edit: 2026 May 15 - This post is archived.
Mob is alive!!
Main docs: mob v0.7.11 — Documentation
A bit of explanation for the slightly c...
New
I released Doggo, a collection of unstyled Phoenix components.
https://github.com/woylie/doggo
Features
Unstyled Phoenix components....
New
Hi everyone,
I’ve been working on this protobuf library for 3 years. We use it in the company I work for, EasyMile, to communicate with ...
New
Hobbes is a low-level distributed database for the Elixir programming language.
Hobbes provides a simple, safe, and scalable storage lay...
New
I’ll shortly be launching Text, a nascent text analysis library.
Current functionality
In this early version (not ready for prime time) ...
New
Following on from my CLDR lbraries I started work on Unicode transforms. But like everything related to CLDR there is a lot of yak-shavin...
New
Other Trending Topics
I am happy to introduce the very α version of the new programming language compiled to BEAM.
Welcome Cure.
It has literally three kille...
New
A little off-topic, but I feel like people here have a good head on their shoulders.
I used to be quite good at making software. Was luc...
New
Hey. Is there anyone here who creates agents in their apps? Not talking about using agents, but creating them. I’m finding it pretty diff...
New
With AI doing more of the implementation work, I’ve been wondering how much coding I should deliberately keep doing myself.
My main conc...
New
I love Elixir. It’s one of 2 programming languages I’ve ever fallen in love with.
But I don’t use it anymore.
Serverless was the promis...
New
I just stumbled on a newly redesigned elixir-lang.org. :tada: It looks like @Software_Mansion did the work, and I think it is generally a...
New
Categories:
Sub Categories:
Forums
Popular Tags
- #ecto
- #liveview
- #troubleshooting
- #learning-elixir
- #library
- #deployment
- #erlang
- #testing
- #genserver
- #mix
- #absinthe
- #remote-other
- #otp
- #plug
- #how-to-question
- #macros
- #postgres
- #elixirconf
- #channels
- #exunit
- #discussion
- #code-sync
- #podcasts
- #javascript
- #onsite
- #dialyzer
- #docker
- #authentication
- #umbrella
- #full-time-contract
- #podcasts-by-brainlid
- #ai
- #ecto-query
- #elixirconf-us
- #blog-post
- #elixir-ls
- #phoenix_html
- #iex
- #graphql
- #genstage
- #websockets
- #supervisor
- #advent-of-code
- #distillery
- #processes
- #elixirconf-eu
- #api
- #forms
- #metaprogramming
- #hex










First Post!- Show Best Posts
- Show All (oldest first)
- Show All (newest first)
kip
Based on the positive feedback it seems this project has merit. I didn’t get finished all the work I planned for the weekend but I’m clearing a backlog so I can give this project some greater attention. The inflector is finished (nouns, pronouns, verbs). I’ve started work on adding the Metaphone 2 algorithm but its a slow slog because there is no description of the algorithm I can find - just imperative code. And then some more testing and verification on language detection.
All in all, likely a one week delay.
Most Liked
kip
Thought I’d share the near term roadmap in a little more detail. Feedback is most definitely welcome on the capabilities you would find most useful. Or any areas you’d like to contribute to.
Step 1: Language recognition
Most natural language processing is language dependent. So identifying the source language is important. The primary way of identifying languages is to split the text into n-grams and then perform various statistical analysis of the source text versus the same analysis of a standard corpora in multiple different languages. The Universal Declaration of Human Rights is a standard text published in a lot of languages so this is the corpora I’m using. There are different ways to correlate source text versus a corpora. I am primarily using the algorithms in Language Identification from Text Using N-gram Based Cumulative Frequency Addition.
This is the due now for delivery on 28th June.
Step 2: Text segmentation
No matter what analysis is required, segmenting the text into grapheme clusters, words and sentences is required. This is very language dependent. Elixir’s
String.graphemes/1implements the Unicode segmentation algorithm for grapheme clusters so thats taken care of. Elixir’sString.split/1implements the Unicode segmentation algorithm for words.String.split/1is great for a default case but its not sufficient for language-specific segmentation. And we still need sentence segmentation too. Therefore I am implementing the CLDR Segmentation rules which provide language-specific customisation for text segmentation. This is another rules parser (I think so far I have implemented 8 different rules parsers and “compilers” in various parts of the ex_cldr project).The text segmentation algorithms will be implemented as part of the unicode_string library.
Step 3: Parts of Speech Tagging
Now we have segments of text we can proceed to understanding what is being expressed. The starting point for this is called “parts of speech tagging”. Because I want a good native Elixir implementation that supports a wide variety of languages with good (but not necessarily the absolutely best) tagging I’m using A Rule-based Part-of-Speech and Morphological Tagging Toolkit which provides a fully trained corpora for ~90 languages using the open source data maintained in the Universal Dependencies treebanks. The trained models are maintained in the RDRPOSTagger project which also defines a rules engine that I will implement in Elixir. Another parser/compiler
Step 4: Sentiment Analysis
Now that we have a grammatical breakdown of the target source we can start to identify meaning. Wikipedia says:
The implementation approach is not yet defined and feedback and suggestions are warmly welcomed.
Step 5
To be determined. It will take a 2-4 months to get through the first 4 steps so thats plenty of time for feedback and collaboration
kip
I’ve published Text 0.4.0 today with seven new NLP modules (all native Elixir, no NIF or ML).
They cover the kinds of preprocessing you might reach for once your sentiment / classification / search pipeline outgrows
String.split/1.This release represents a largely feature complete text library from my perspective. Happy to take feature suggestions though.
A few of the more immediately useful additions in this release:
Text.Clean— pipeline-style normalisationWhitespace, control characters, smart quotes, mojibake, NFC/NFKC. Composable; defaults are sensible.
Text.Truecase— restore casing for ALL-CAPS or lowercased textPOS-aware heuristics for proper nouns, acronyms, and sentence starts. Useful when an upstream system has destroyed the casing (chat logs, OCR, screaming customer feedback).
Text.Emoji— detection, stripping, counting, conversionBacked by the
:unicodepackage’s emoji property tables, so it recognises every codepoint flagged emoji in the current Unicode release — no shipped JSON.Text.Hyphenation— Knuth–Liang TeX-pattern hyphenationShips en-US patterns baked in (~5 000). Other languages load from any standard
hyph-*.texfile.Text.PII— detect & redact common identifiersPhone, email, credit-card-shaped digits, IBANs, IPv4/IPv6, US SSN. Pattern-based — fast and deterministic. The right tool for “please don’t paste this into the LLM” preflight; pair with a stricter checker if you need legal-grade accuracy.
Text.Spell— Norvig-style spelling suggestionsEdit-distance candidates ranked by frequency in
Text.WordFreq(the 30,000-word English frequency table that also ships in 0.4.0).Text.Summarize— extractive summarisation via TextRankSentence-graph TextRank with configurable similarity (
:cosineor:jaccard) and target length.kip
text version 2.0 has been published today, just on schedule. In addition the library text_corpus_udhr is also published today - it provides a corpus to support natural language detection.
Language Detection
textcontains 3 language classifiers to aid in natural language detection. However it does not include any corpora; these are contained in separate libraries. The available classifiers are:Text.Language.Classifier.CommulativeFrequencyText.Language.Classifier.NaiveBayesianText.Language.Classifier.RankOrderAdditional classifiers can be added by defining a module that implements the
Text.Language.Classifierbehaviour.The library text_corpus_udhr implements the
Text.Corpusbehaviour for the United National Declaration of Human Rights which is available for download in 423 languages from Unicode.Examples:
Word Counting
textcontains an implementation of word counting that is oriented towards large streams of words rather than discrete strings. Input toText.Word.word_count/2can be aString.t,File.Stream.torFlow.tallowing flexible streaming of text.English Pluralization
textincludes an inflector for the English language that takes an approach based upon An Algorithmic Approach to English Pluralization. See the moduleText.Inflect.Enand the functions:Text.Inflect.En.pluralize/2Text.Inflect.En.pluralize_noun/2Text.Inflect.En.pluralize_verb/1Text.Inflect.En.pluralize_adjective/1N-Gram generation
The
Text.Ngrammodule supports efficient generation of n-grams of length2to7. SeeText.Ngram.ngram/2.Language detection accuracy
Detection accuracy is reliable at text lengths of 150 characters or more, reasonable at 100 characters and may not be considered acceptable at shorter lengths.
The results are consistent for the range of tested languages with German being a clear exception where the results are unacceptable for now.
Further details are contained in the github repo in the
analysisdirectory.English language with Naive Bayesian classifier
Text.Language.detect/2withclassifier: Text.Classifier.NaiveBayesianand three different vocabularies.Accuracy for German language detection
German is an exception to the consistent accuracy of most languages and the results are poor. Further analysis is required to understand the underlying cause.
Last Post!
kip
Today I’ve published Text version 1.0. There are no API changes from the currently published version but it does add:
Text.Extract.Link, implementing UTS #58 §3.5.1 termination against the full Unicode repertoire - 65 bracket pairs and 129 ranges of soft terminators, against the 4 pairs and 7 characters handled before.Text.Extract.Escape, implementing UTS #58 §4.1 minimal escaping, which rewrites a URL into its most readable form without changing where link detection ends it.minimal/1accepts either a serialised URL or a keyword list of already-parsed parts, the latter for when a syntax character is data rather than structure.URL detection now accepts hosts and paths that are wholly non-ASCII, all four UTS #46 label separators (so
普遍适用测试。我爱你is a two-label host), hosts carrying an explicit root label (foo.example.com./path), andmailto:addresses. Script-internal punctuation such as the Tibetan tsheg is accepted in host labels, since UTS #46 processing validates the host afterwards.