vertti
Channels - Where to store state?
I’m building a sort of highly customized chat application with Phoenix, with some functionality similar with XMPP MUC.
Now, I have both persistent as well as transient state for both, users and channels. While the persistent data is comfortably saved in Ecto/pg, my problem is the transient state. Particularly for the channels, as user-specific transient state can be stored in socket.assigns.
Where to store channel-specific transient state in Phoenix?
For example, each channel has a quiet mode. With quiet mode on, users below certain rank cannot send messages to the channel (=topic in Phoenix slang).
Currently, my approach is to have a table in mnesia which holds all transient data for each channel. However, I don’t feel it’s quite the right approach to query mnesia every single time any user sends a message to check if quiet mode is on. I understand mnesia queries are cheap, but I was thinking if there’s a more sort of Phoenix way to handle this. I guess the most ideal way would be to carry around a channel-specific attribute, much like socket.assigns is.
Moreover, some parts of the channel-specific transient state have an expiry time, meaning that after a period of time they need to be removed from the state. What’s the best approach to handle this - should I spin up another process just to monitor them and when the expiry time hits, remove them from the state in mnesia?
Trending in Questions
Other Trending Topics
Categories:
Sub Categories:
Forums
Popular Tags
- #ecto
- #liveview
- #troubleshooting
- #learning-elixir
- #deployment
- #library
- #erlang
- #testing
- #genserver
- #mix
- #absinthe
- #remote-other
- #otp
- #plug
- #how-to-question
- #macros
- #postgres
- #channels
- #elixirconf
- #exunit
- #discussion
- #code-sync
- #javascript
- #podcasts
- #onsite
- #dialyzer
- #docker
- #authentication
- #umbrella
- #full-time-contract
- #podcasts-by-brainlid
- #ecto-query
- #elixir-ls
- #phoenix_html
- #iex
- #blog-post
- #graphql
- #genstage
- #ai
- #elixirconf-us
- #websockets
- #supervisor
- #advent-of-code
- #distillery
- #processes
- #api
- #forms
- #metaprogramming
- #hex
- #performance










First 10 of 18 Posts
MrTortoise
Have you tried stockings?
More serious note. Persistent storage is a trade off between speed and scale. As soon as you make that trade off you end up moving data from in memory state of day a gen server into a persistence layer.
You then also end up in your specific domains problems.
If you do not want to hit your persistence layer on every request you want caching or to not make the trade off in the first place. Bear in mind cache invalidation is hard and involves some very tricky trade off decisions that will not lead you to a single best answer.
vertti
Sorry I don’t entirely understand your point here. I have persistent data, such as user accounts and some channel specific data in Postgres. However, this data is queried quite rarely, inserted/updated even more rarely.
What I’m having trouble with is deciding the architectural pattern for maintaining transient channel specific state. For example, that quiet mode state currently resides in mnesia and is purely transient.
dom
Honestly the way you’re doing it sounds totally reasonable. Queries to ram or disc copy mnesia tables go through ETS which is fast. In fact I think pg2 phoenix pubsub hits ETS a couple times on every broadcast anyway.
I have a process like you suggest to expire keys, although it’s for postgresql rather than mnesia, and it works great. It doesn’t get much load anyway so I never bothered optimizing or sharding its work.
vertti
Thanks for giving me some peer validation for my approach! I’m using mnesia purely in RAM.
How do you actually work out the whole expiration → delete in your application? You set up a process to track and do it, or is there perhaps some better way?
dom
In my case the process is a gen_server that wakes up periodically (using send_after), deletes old stuff from PG using a query like expiry_timestamp <= now, then broadcasts to whatever topics were affected. The process is locally registered, i.e. it’s unique within a node but not within the cluster, since it doesn’t really harm if two nodes try to expire old stuff at the same time. But it wouldn’t be hard to make it a singleton (GitHub - arjan/singleton: Global, supervised singleton processes for Elixir · GitHub).
If you need more accuracy, you could have the process use a send_after or timer per item rather than expiring stuff in bulk. Timers are less efficient than send_after but can be canceled.
Another approach is to check the expiry timestamp when you read the table (i.e. when sending a message). Then you don’t need a dedicated process.
Cachex has an example of the send_after approach: cachex/lib/cachex/services/janitor.ex at main · whitfin/cachex · GitHub
vertti
Cool, I’m intending to only build for a single node for now so I have some less worries.
Funnily enough, I was also thinking first about having some cleanup function invoke whenever somebody would send a message and the expiry timestamp was past the current time.
Since I’m new to Elixir and Phoenix, I didn’t even know about send_after but it looks like a really good fit for me. I’ll probably try to imitate your approach with the gen_server and see if I can build it up successfully.
sasajuric
It’s a bit hard to recommend the concrete approach, because the description is quite vague, but I’ll give it a try
To avoid ambiguity with the term channel, I’ll use the term topic instead.
My first take on this would be to use a separate
GenServerfor each topic. You could hold all the transient data of the topic in that server. The server would also need to keep track of whether it’s in the quite mode. So now, whenever someone wants to send a message to the topic, they need to do it through the corresponding server. Since the server knows about its mode, it can easily decide whether to send the message or not.Adding expiry logic can now be fairly simple with the help of :timer.send_interval or Process.send_after which would be invoked in the topic server. Either approach will result in some message sent to the server, which you’d need to handle in the corresponding
handle_infoand remove expired items.This approach is very consistent, and when done properly, you’ll have no race conditions or strange behaviour. However, the topic server might turn out to be a bottleneck if you have frequent activity on a topic (frequent messages, mode changes, or other state changes). In that case, you could consider using an ETS table. ETS tables can boost your performance significantly, but they are appropriate only for some situations, so you need to think it through. In some cases a hybrid approach is needed, where all the writes are serialized through the server process, while reads are performed directly from the table.
If you go for ETS tables, and want to expire items, you could look into some of the caching libraries, such as CachEx, or my own con_cache.
Cachexhas more features and higher activity, so as the author ofcon_cache, I’d myself recommend looking intoCachexfirst.If the transient state is really simple, then I’d consider ETS from the start. For example, if we only need to deal with the quiet mode, then I’d just have one ETS table where I’d store that info for each topic.
In more complex cases, I’d start with
GenServerand do some testing to see if it can handle the desired load.Again, the problem is vaguely defined, so it’s hard to precisely say which of the two would be a more suitable choice in this particular case.
vertti
Yes I understand I didn’t specify the domain very accurately here, thanks for giving your input nevertheless! It’s actually very interesting building a sort of modern administrated Multi-User Chat on Elixir/Phoenix, without the headaches of XMPP.
The nature of my application is:
The transient state each topic has:
So my trouble has been trying to decide whether to make one sort of global state tree (like
:mnesia) or whether to prefer more local approaches for each topic.Your
GenServerapproach was also my first thought, but I read somewhere that in big topics (lots of subscribers) the GenServer could become the bottleneck, as it is just one process. Now I’m thinking to use aGenServerfor the keyword list and maybe for the message history, andETS/CachExor yourcon_cachefor the rest.By the way, is there any reason not to prefer
:mnesiaoverETS? From what I understand, if I use the transactions:mnesiacould be a more safe choice, although it does carry some performance penalty. However, that way, I wouldn’t need to serialize the writes through a channel specificGenServer, but I could still perform fastdirty_reads(if that makes sense).vertti
Just to add a bit; one problem I could seriously use advice with is spam prevention.
Is there anyone who has built their own spam prevention solution to Phoenix’s channels? My approach now is to just have something in socket.assigns to track the messages sent, then use that to determine whether the user is allowed to send a new message yet or not.
Somehow my approach just feels too clunky. I hope there is a better way in handling this.
sasajuric
Notice that I suggested
GenServerper each topic, so it’s not just one process. Although admittedly, it is one per topic, so it can still be a bottleneck for highly active topics with many users. But before doing some complex optimizations, this needs to be proved by measuringGiven how your state looks, I’d advise starting with the
GenServerapproach. It looks like it might not scale past some point, but it will give you a consistent starting point.Then, I’d implement a load tester, and analyze the behaviour of the single topic server. This should help you discover the capacity of the topic server. Perhaps it’s going to be good enough for the projected load.
If not, then you need to find the bottlenecks, and tackle them. You can use a variety of approaches here:
Mnesia requires a bit more ceremony, and you need to be careful with transactions, because on conflict, they will be restarted. If you want a fast in-memory key-value, I think ETS is your simplest option, and that’s where I’d start first.
You could do the same with GenServer that serializes writes, and an ETS table which acts as the snapshot of last writes. The writer server keeps the last state in its own state. When the write request arrives, it updates the state, and then stores it to the ETS table. All other clients read from that table, so they don’t have to wait for the writer to finish.
The benefit here is that you don’t need to use locks and transactions, so it’s easier to reason about the behaviour, especially in a highly concurrent system. This approach also paves way for actively managing the load, and applying backpressure, for example using GenStage.