yoavgeva
Hey everyone,
I’ve been working on something I want to share and get feedback on.
The itch
Every web app I build ends up with the same stack: a database, a cache layer, and the app itself. Three things to deploy, monitor, pay for, and hope they stay connected to each other.
For caching, the options are:
Redis — great protocol, great ecosystem, but in production you need a managed service (ElastiCache, Upstash, etc.) because self-hosted Redis persistence is… optimistic. AOF rewrite can lose your last few seconds. RDB snapshots lose everything since the last dump. And now you’re paying for a whole separate instance just to hold keys.
Java solutions* (Hazelcast, Infinispan) — powerful but painful to operate. Config files longer than your app code. JVM tuning. Cluster formation issues. Not exactly “easy to maintain.”
ETS/Mnesia — built into the BEAM, fast, but volatile. Node goes down, data’s gone. Mnesia replication exists but comes with its own set of surprises.
What I actually wanted: use the disk I’m already paying for on my instance, speak a protocol I already know (Redis), get real durability without a separate service, and keep my stack simple.
What I built with AI
[FerricStore]( GitHub - yoavgeva/ferricstore: Distributed crash-safe key-value store with Redis wire protocol. Durable by default — every write is Raft-committed and fsync'd. Embeddable Elixir library or standalone server. Built with Elixir + Rust NIFs. · GitHub ) is a persistent key-value store written in Elixir + Rust that:
Speaks RESP3 — connect with redis-cli, or any Redis client library that support RESP3. Your existing code mostly just works.
Every write is durable by default — Raft consensus + Bitcask append-only log + fsync. When you get OK back, your data is on disk. Not “eventually.” Not “if the AOF rewrite finishes.” On disk, right now.
Runs embedded in your app — add it as a dependency, FerricStore.set("key", "value"). No separate process, no network hop, no connection pool. Your cache lives in the same BEAM node as your app.
Or runs standalone — `docker run -p 6379:6379 yoavgeva/ferricstore` and connect with redis-cli. Drop-in for apps that already speak Redis.
50+ Redis commands — strings, hashes, lists, sets, sorted sets, TTL, MULTI/EXEC, pub/sub, and more.
Native Elixir commands — CAS (compare-and-swap), distributed locks, rate limiting, FETCH_OR_COMPUTE — things you’d build on top of Redis but get out of the box here.
Probabilistic data structures — Bloom filters, Cuckoo filters, Count-Min Sketch, TopK, HyperLogLog, T-Digest. All built-in.
Vector search — HNSW index for similarity search, built into the storage engine.
The architecture in 30 seconds
Client (redis-cli / Redix / your app)
|
v
RESP3 Parser (pure Elixir, zero-copy)
|
v
Command Dispatcher
|
v
Raft Consensus (via ra library — same one RabbitMQ uses)
|
v
Bitcask Storage Engine (Rust NIF — append-only log, CRC-checked)
|
v
ETS Hot Cache (recent values served from memory, cold values read from disk)
Writes go through Raft for consistency, then to Bitcask for persistence. Reads hit ETS first (microseconds), fall back to disk for cold data. The Rust NIF handles the low-level I/O — pure functions, no dirty schedulers, proper consume_timeslice yielding so the BEAM scheduler stays happy.
Embedded mode — the thing I’m most excited about
# mix.exs
{:ferricstore, "\~> 0.1"}
# your app
FerricStore.set("user:123:session", session_data)
{:ok, data} = FerricStore.get("user:123:session")
# with TTL
FerricStore.set("rate:api:123", "1", ttl: :timer.seconds(60))
# atomic operations
{:ok, new_count} = FerricStore.incr("page_views")
# compare-and-swap
:ok = FerricStore.cas("inventory:sku42", "10", "9")
No Redis connection. No connection pool config. No “what happens when Redis is down.” It’s just a function call that persists to disk. Your Phoenix app, your cache, one deployment, one thing to monitor.
Standalone mode — drop-in Redis replacement
If you have apps in other languages, or you just want a Redis-compatible server with real durability:
bash
# Docker
docker run -p 6379:6379 -v ferricstore_data:/data yoavgeva/ferricstore
# Then use any Redis client
redis-cli SET mykey "hello"
redis-cli GET mykey
Or from your Elixir app via Redix:
elixir
{:ok, conn} = Redix.start_link("redis://localhost:6379")
Redix.command!(conn, \["SET", "user:42", "alice"\])
Redix.command!(conn, \["GET", "user:42"\])
# => "Alice"
Everything you’d expect from Redis works — MULTI/EXEC transactions, pub/sub, pipelining, HELLO 3 (RESP3). Plus you get a built-in health endpoint (GET /health on port 6380), Prometheus metrics, ACL authentication, and TLS support.
The dashboard gives you live visibility into shards, key counts, memory pressure, hit rates, and slow queries — no Grafana setup needed.
What it’s NOT
Not a database replacement — it’s a cache/store. Great for sessions, rate limits, feature flags, job queues, counters, leaderboards. Not for your users table.
Not a sharded cluster (yet) — scales to 3-5 nodes via Raft replication for high availability and read scaling (every node serves reads from local ETS). Storage capacity scales with disk, not RAM — a 500GB NVMe gives you 500GB of cache. But there’s no hash-slot sharding across nodes yet, so every node holds all the data.
Not battle-tested in production — this is v0.1. It passes 8000+ tests including shard-kill recovery and multi-node cluster tests, but it hasn’t seen real production traffic yet. That’s where you come in.
Not going to beat Redis on raw throughput — Redis keeps everything in RAM and doesn’t fsync by default. But it’s persistent store where every write hits disk,
Why Elixir + Rust?
Elixir gives us the BEAM’s supervision trees, distribution primitives, and ETS for the hot cache. Rust gives us a memory-safe, zero-copy storage engine that doesn’t GC-pause or mess with the BEAM scheduler.
The NIF boundary is clean — Rust functions are pure and stateless. No Mutex, no shared state, no dirty schedulers. Just `v2_append_batch(path, entries)` and `v2_pread_at(path, offset)`.
Looking for
Feedback on the API, the architecture, the docs — what’s confusing, what’s missing, what would make you try it?
Early adopters willing to try it in a side project or staging environment and report what breaks.
Use case ideas — what would you use a durable, embedded Redis-compatible cache for in your Elixir apps?
Links:
- Hex: ferricstore | Hex
- Docs: ferricstore v0.5.7 — Documentation
- Docker: yoavgeva/ferricstore - Docker Image
- GitHub:
https://github.com/yoavgeva/ferricstore
Happy to answer any questions. And yes, the name is a pun — Ferric (iron/Rust) + Store.
Trending in Announcing
Other Trending Topics
Categories:
Sub Categories:
Forums
Popular Tags
- #ecto
- #liveview
- #troubleshooting
- #learning-elixir
- #library
- #deployment
- #erlang
- #testing
- #genserver
- #mix
- #absinthe
- #remote-other
- #otp
- #plug
- #how-to-question
- #macros
- #postgres
- #elixirconf
- #channels
- #exunit
- #discussion
- #code-sync
- #podcasts
- #javascript
- #onsite
- #dialyzer
- #docker
- #authentication
- #umbrella
- #full-time-contract
- #podcasts-by-brainlid
- #ai
- #ecto-query
- #elixirconf-us
- #blog-post
- #elixir-ls
- #phoenix_html
- #iex
- #graphql
- #genstage
- #websockets
- #supervisor
- #advent-of-code
- #distillery
- #processes
- #elixirconf-eu
- #api
- #forms
- #metaprogramming
- #hex











Showing Posts 29 to 20- Show Best Posts
- Show All (oldest first)
- Show All (newest first)
yoavgeva
Yes, this is what I mean by durable here.
If any workflow command return success, the flow state was already written through FerricStore durable path before the ack.
So if the app / worker restart, the workflow is still there.
If FerricStore server restart, it should also recover the current flow state from the data dir / volume
and rebuild the hot indexes / projections from the durable state.
Important details:
The model is durable workflow state (imagine persistence state machine), not durable code replay.
One caveat: you need to run FerricStore with persistent data dir / docker volume / disk. If you run on
tmpfs or delete the volume, of course there is nothing to recover from.
dimitarvp
Can you clarify whether durable work flows in this case is truly durable between application / database restarts?
yoavgeva
Hey everyone,
Small update from my side.
Current version is `0.7.1` on Hex:
The repo also moved to the org repo:
The direction also changed a lot from the first post.
The biggest change is that I moved from the Redis/RESP wire protocol in the standalone server to the Ferric native protocol.
That was intentional, because I want to support both things better: durable KV / Redis-style data structures, and FerricFlow workflows.
FerricStore still has the KV/data-structure side. The goal is still to support the useful Redis-style primitives: strings, hashes, lists, sets, sorted sets, TTLs, counters, locks, rate limits and so on.
But the main promise is not “any Redis client can connect and everything is a Redis drop-in” anymore. The current promise is more like: FerricStore is a durable native store with Redis-style data structures, and FerricFlow adds durable workflows and queues on top.
The thing I am more excited about now is FerricFlow: durable workflows on top of the same store.
The reason I started to move there is that I wanted a high throughput workflow engine, and I didnt find the thing I wanted.
In many projects I always ended up using queues for everything. Queue for this, retry queue for that, delayed queue, dead-letter queue, some DB table for status, some script to repair stuck jobs. It works, but after a while the real system is not the queue anymore, it is all the workflow state you build around it.
What I wanted was something with the speed and simple worker model of queues, but with the workflow state built in from the start.
The things FerricFlow is trying to solve are the things I always see people build around queues:
- durable workflow state
- leases and fencing
- retries and retry exhaustion
- history and audit
- signals
- value refs for big payloads
- fanout / child workflows
- indexed workflow metadata
- retention and repair
- governance controls
The native protocol gives me more control on the things both sides need: request ids, multiplexing, routing, backpressure, typed SDKs, leader hints, ACL behavior, KV commands, and workflow specific commands.
So KV is still supported and important. The direction is FerricStore native protocol with Redis-style data structures + FerricFlow workflows.
I also started using the HA cluster in my own projects already. In my company we also started using it for small stuff, and the feedback until now is very positive. It is not me saying “this is finished 1.0 production database”, but it is also not only a toy repo anymore. I am dogfooding it in real places, and that is helping me see what need to be better.
The numbers that are published now are for FerricFlow workflow/queue workloads. On the Azure 16-vCPU runs, the best balanced 32-shard runs show roughly:
- `54,060` workflows/s end-to-end (3 states), in the workflow-worker benchmark shape
These are not universal numbers, only the result for this workload and hardware.
I still need to rerun KV SET/GET with the current native SDK/protocol before I publish latest KV numbers. I also want better Elixir-facing comparisons with what people actually use here: Cachex, Nebulex, ETS based caches, and DB backed solutions.
What changed in the current release line:
- KV/data structures are still supported through the Ferric native protocol and embedded API.
- The direction is Redis-style primitives, not Redis wire-protocol compatibility as the main promise.
- FerricFlow is now the main focus: durable workflows built into FerricStore.
- Flow state has leases, fencing, retries, history, signals, value refs, fanout, retention, and indexed metadata.
- There is now a governance layer for workflows:
- budgets
- human approvals
- external effect tracking
- distributed limits
- circuit breakers
- governance ledger/debugging
The governance layer is the part I think is most intresting now.
The problem is not only “how do I run a job later?”.
It is:
- should this work retry?
- how many times?
- how much budget can it spend?
- does a human need to approve the next step?
- did an external side effect already happen?
- should a circuit breaker stop this provider call?
- is one tenant using too much capacity?
- why did this workflow pause, fail, or get denied?
If you also have this pain in your projects, give it a try you might like it
yoavgeva
I agree with what you said overall, I never said that I won’t benchmark, I will compare with all relevant products as you see fit, with your own tests if you want.
I am saying I don’t expect that the results will be amazing benchmark wise for example the result be 10-20% worse than the solution for one node, people usually choose products because of this, I am trying to give something product wise which is SLA of 99.99% with simplicity of setup and ops, this features come at a cost of performance.
I agree that I need to polish the readme, guides and code quality, it’s only the start, but it nice to hear from you that this kind of solution have some merit in today world, that’s what I was searching for.
ziinc
Benchmarks do not just “help with marketing”. You are asking busy developers to trust your solution on production projects, but won’t bother to do up even a simple Benchee benchmark to compare lookup speeds against existing libraries?
FWIW your library does target a very relevant use case for my day job, but I would not touch this with a 10 feet pole unless there is some effort made on code quality and performance assurance.
Some pointers:
If you want your caching solution to succeed in the Elixir space and community, it also has to address the concerns raised in such threads critically, as some have probably hand rolled similar solutions to what you are building in some shape or form.
yoavgeva
I didn’t run benchmark yet, because it’s not fully ready, so I can’t say numbers, I will tell you the arch comparison between both:
Both use Redis spec, which mean for read concurrency they are not optimized, because the spec is tuned for Redis which is one thread, which mean each command arrived must return in the same order per connection.
Ferricstore has embedded mode for elixir/erlang which also mean it is not limited by rspec, so I expect even higher concurrency there.
Need to remember that the numbers of the benchmark also depend on machines, runtimes, etc.., and I don’t understand how they help, because in usage you don’t only do set/get you do a lot more, you need more features that allow more different stuff to do, so for example maybe Redis will be faster for your kind of work, but other kind of work Ferricstore or kvrocks will be better and faster, well I do understand that benchmark help in marketing
syepes
It would be interesting to know how does it compare with KvRocks in terms of write & read performance. For the moment I have not found another K/V db that supports disk persistent that is faster.
https://github.com/apache/kvrocks
AstonJ
A few days ago we updated the flag system after having changed the post composition tip to the following a few weeks ago:
The flag system:
We’ve been monitoring AI generated posts in a log since Feb of last year and determined this is the best way forward for now.
If people want to post AI generated text/conversations, then the new section we created last year may be a good fit as it would ensure the flow of the main thread remains in line with expectations/the norm: https://forum.elixirforum.com/t/new-ai-conversations-section/72455 (we could probably update those guidelines now).
To the OP, you should be able to edit the post and submit a new version (just please make sure it’s in your own words - translated to English with the help of an AI is fine if need be).
Asd
Yeah, I agree that your post was on-topic. It was automatically hidden when it received a number of flags from regular users. In any way, this forum doesn’t have a rule which says “You can’t write or edit posts with help of AI”, and I hope that moderators will make it visible once they have time to review it.
yoavgeva
Well hard not to take hard, when someone flag my post, if it was read it’s actually what I was doing with the lib, and I did review to post before sending, as for sending fixes fast, AI love Elixir, and this arch already in my mind, so I know what to feed it and how to work with, FYI I work on 4 projects in the same time, you can call it AI slop, I call it to maximize potential, I agree they are some who use it lazily more than other, but you can say the other also that people become lazy to read and understand what the user try to bring, English is not my first language I am trying to publish a Library which is a little bigger than other seen here, so it hard to explain it with clarity, I need AI for better clarity, because the way you guys responded made it me feel that I was not clear.