Harrisonl
We have an ECS cluster with 4 services, where each task joins a single cluster, via discovery ECS discovery service.
Currently when I deploy, new tasks will attempt to join the old cluster as they a being brought up (potential red flag). I think (but not sure) that this is causing the above message to happen a lot on deploy, as once new tasks reach a healthy state, old tasks drop off causing incomplete partitions as some tasks may be notified first.
This is causing many issues, especially with Horde as it tries to connect to the old nodes (something between Libcluster and Horde not communicating correctly) resulting in every deploy having our Horde registry be reset and orphaned processes.
I’m not sure if this is actually the issue though or not or if there is something else I’m missing. I’m using LibCluster DNSPoller (1 second interval) for node membership and OTP 26.
My questions are:
- Is this expected for this setup (dynamic cluster with nodes joining / leaving frequently) or am I missing something?
- I couldn’t find a concrete answer on this, but is it better to have the new nodes not join the old cluster when they are brought up (e.g. ensure they all have the same version etc.)?
Trending in Questions
Other Trending Topics
Categories:
Sub Categories:
Forums
Popular Tags
- #ecto
- #liveview
- #troubleshooting
- #learning-elixir
- #deployment
- #library
- #erlang
- #testing
- #genserver
- #mix
- #absinthe
- #remote-other
- #otp
- #plug
- #how-to-question
- #macros
- #postgres
- #channels
- #elixirconf
- #exunit
- #discussion
- #code-sync
- #javascript
- #podcasts
- #onsite
- #dialyzer
- #docker
- #authentication
- #umbrella
- #full-time-contract
- #podcasts-by-brainlid
- #ecto-query
- #elixir-ls
- #blog-post
- #phoenix_html
- #iex
- #graphql
- #ai
- #genstage
- #elixirconf-us
- #websockets
- #supervisor
- #advent-of-code
- #distillery
- #processes
- #api
- #forms
- #metaprogramming
- #security
- #hex











Showing Posts 1 to 3- Show Best Posts
- Show All Posts (oldest first)
- Show All Posts (newest first)
Harrisonl
It also might be related to this - which occurs before those messages start ~1 min before (however not always present before the disconnect messages)
Which is weird because it will then connect straight after that. To remove security group issues, I’ve also allowed the tasks to communicate with each other on every port as well.
schneebyte
If you don’t need
:globalthen i would recommend to disableprevent_overlapping_partitionsgfviegas
Had the exactly same issue deploying in a Fargate ECS cluster with 2 tasks (each with one container with a phoenix app using libcluster dns poller).
Besides desabling the
prevent_overlapping_partitionsas recommended by @schneebyte I’ve also decreased theMin running tasks %in my ECS Service to 50% to ensure a task will still be up to converge a replicated cache in the newly created nodes.It fixed for my use case.
To give a little bit of ‘how to’ for newbies (like me) in clustering in elixir, if you are using releases and wish to disable the prevent overlapping partitions option, you should add a new line in
rel/remote.vm.args.eex(or similar) with the followng content:If you are not using releases and need a quick fix, you can also set an environment variable
ELIXIR_ERL_OPTIONSwith this setting, such asELIXIR_ERL_OPTIONS='-kernel prevent_overlapping_partitions false'.