alexcastano
Using libraries like singleton you can create applications which are the only one in the whole cluster. This application may be a GenServer which initializes an ETS table and delegates the requests to this ETS. It is like a thin wrapper to convert a local ETS into a global one.
I have this doubt because I read the following in the hammer documentation:
There may come a time when ETS just doesn’t cut it, for example if we end up load-balancing across many nodes and want to keep our rate-limiter state in one central store. Redis is ideal for this use-case, and fortunately Hammer supports a Redis backend.
If you need a shared cache for an elixir cluster, what do you prefer to use a Redis server or a singleton ETS or even a singleton cachex?
I can think advantages and disadvantages of both methods.
- With singleton ETS or Cachex you can save Elixir structs or any erlang term, in redis you have to
castdata from database to memory, unless you use very basic data structures like strings or numbers. - Singleton ETS won’t be as performant as a local one, but I think in most of the case could be enough. Maybe is it faster than Redis anyway?
- With Redis you have to create a bigger wrapper to use it like cache.
- You lose your cache if the node dies. This can be acceptable in some scenarios.
- Less architectural dependencies.
- Redis has built-in some interesting features like TTL.
I didn’t test it, because of that my question is: Is a singleton ETS a good idea or it has his own drawbacks? Or would you prefer Redis?
Trending in Questions
Other Trending Topics
Categories:
Sub Categories:
Forums
Popular Tags
- #ecto
- #liveview
- #troubleshooting
- #learning-elixir
- #deployment
- #library
- #erlang
- #testing
- #genserver
- #mix
- #absinthe
- #remote-other
- #otp
- #plug
- #how-to-question
- #macros
- #postgres
- #elixirconf
- #channels
- #exunit
- #discussion
- #code-sync
- #podcasts
- #javascript
- #onsite
- #dialyzer
- #docker
- #authentication
- #umbrella
- #full-time-contract
- #podcasts-by-brainlid
- #ecto-query
- #blog-post
- #elixirconf-us
- #elixir-ls
- #ai
- #phoenix_html
- #iex
- #graphql
- #genstage
- #websockets
- #supervisor
- #advent-of-code
- #distillery
- #processes
- #api
- #forms
- #hex
- #security
- #metaprogramming











Showing Posts 1 to 10- Show Best Posts
- Show All (oldest first)
- Show All (newest first)
mbuhot
Sounds like it would create a single process bottleneck for reads.
Maybe mnesia would be a good fit for distributed cache state?
alexcastano
Yes, you are right, but If I’m not wrong, Redis is mono thread. I don’t know the performance of this bottleneck vs Redis, maybe it is still better, maybe not. Other question could be how to avoid this bottleneck and still making it global.
It seems like a more complex solution: Caching: ETS, Mnesia, Redis - #7 by cmkarlsson
cmkarlsson
The
singletonlinked is just a wrapper around theglobalmodule. It has its own set of problems, mostly that it can’t handle split brain scenarios and it is comparatively slow.I don’t think you should write off
mnesiabased on that, it is not written to discourage but to show some things you may need to handle depending on use case. It is a great solution for distributed caching and the reasons behind these points is that there are well known “quirks” you may want to be aware of. Better the devil you know as they say. Most other distributed setups will have some set of problems but you may not be aware of them. If you go down theglobalover ets you may end up re-inventingmnesia.This externalizes the problem. Instead of dealing with net-splits, high availability and performance in your BEAM cluster you needs an OPs team handling in externally. You may need to setup multiple redis nodes and either have a client being able to swap over to another node in case the master node goes down or setup a virtual IP to handle it which requires more infrastructure.
The question is why you need the distributed cache. Normally it is either because of performance or high availability. If it is for performance then anything cached on the local node will beat redis. If it is for high availability then it is either internal or external complexity.
If you don’t require the HA setup for redis you can just run a cache on a single node (with mnesia for example) and connect directly to the node-name. Then you have a similar setup as a single instance of redis.
alexcastano
Agree.
This has other problems. It is known that
cacheis a complex matter, if each node has its own cache could be even more complex:Hammer.Plug, you need the data to be centralized.HA is not my case.
That’s actually a very good idea. I’ll take a look. Anyway, mnesia seems more complex than ETS for a key value store.
My point is distributed cache is a complicated matter and there are many situations where a single cache for the whole cluster it is enough and much simpler.
alexcastano
I was thinking about this, so I created a very simple benchmark where the
GenServerreply to call asynchronously:However, the results are almost identical, so I suppose the bottleneck still exists, or the benchmark is not similar to a real situation:
dom
This is very similar to the built-in :rpc mechanism, which also uses a single proxy process.
If you need to scale beyond that you can spawn processes directly on the remote node via :erlang.spawn/4.
mbuhot
The bottleneck isn’t in performing the ETS lookup, it’s in queueing all the requests in the mailbox of the GenServer process.
Compare with doing an ETS lookup directly in the client code and you should see a difference.
jordiee
I have personally had some success with distributing ets actions across nodes using rpc.multicall and then all reads just happen on the local node handling the call. This obviously has many areas where it is not great like handling netsplit or servers going down and then coming back(missing writes/deletes) but depending on what you are using the ETS store for could be acceptable(a cache that falls back to hitting db if key not in cache).
This way you do not have a read bottleneck but your writes are slightly slower. I have also recently just not “cared” about the replication and ran it in a Task.start. If it works out great. If not it will be built up when that cache is hit.
For what its worth I originally used mnesia for distribution…but honestly its an absolute pain for dynamic node membership and elegantly handling netsplit so I looked for something simpler in my use case.
I actually added this pattern in yesterday to handle distribution in my auth solution.
https://github.com/jpiepkow/accesspass
Relevant distribution file:
https://github.com/jpiepkow/accesspass/blob/master/lib/access_pass/helpers/ets_distributed.ex
alexcastano
You are right, just for the record I paste my results:
That’s what I thought at least for my case.
Very nice example, it seems much simpler that I thought at first.However, If I’m not wrong, your code provokes a loop infinite loop of calls:
Node A inserts and then it replicates to Node B. Node B inserts and calls again to Node A. If you have more than two nodes the thing is going to be worse, you multiply the calls every time
It can be solved with a no replicate version of each function and being called only replicating (or something similar):
jordiee
It looks like it at first if you quickly glance but you can see that the rpc calls forward to AccessPass.Ets. This file does not include the replication on calls.
https://github.com/jpiepkow/accesspass/blob/master/lib/access_pass/helpers/ets.ex
^ is the file containing the functions RPC calls.