Harrisonl

Harrisonl

We have an ECS cluster with 4 services, where each task joins a single cluster, via discovery ECS discovery service.

Currently when I deploy, new tasks will attempt to join the old cluster as they a being brought up (potential red flag). I think (but not sure) that this is causing the above message to happen a lot on deploy, as once new tasks reach a healthy state, old tasks drop off causing incomplete partitions as some tasks may be notified first.

This is causing many issues, especially with Horde as it tries to connect to the old nodes (something between Libcluster and Horde not communicating correctly) resulting in every deploy having our Horde registry be reset and orphaned processes.

I’m not sure if this is actually the issue though or not or if there is something else I’m missing. I’m using LibCluster DNSPoller (1 second interval) for node membership and OTP 26.

My questions are:

  1. Is this expected for this setup (dynamic cluster with nodes joining / leaving frequently) or am I missing something?
  2. I couldn’t find a concrete answer on this, but is it better to have the new nodes not join the old cluster when they are brought up (e.g. ensure they all have the same version etc.)?

Showing Posts 1 to 3

Harrisonl

Harrisonl OP

It also might be related to this - which occurs before those messages start ~1 min before (however not always present before the disconnect messages)

[warning] [libcluster:xxxx] unable to connect to :“xxxxx@10.0.X.X”

Which is weird because it will then connect straight after that. To remove security group issues, I’ve also allowed the tasks to communicate with each other on every port as well.

schneebyte

schneebyte

If you don’t need :global then i would recommend to disable prevent_overlapping_partitions

gfviegas

gfviegas

Had the exactly same issue deploying in a Fargate ECS cluster with 2 tasks (each with one container with a phoenix app using libcluster dns poller).

Besides desabling the prevent_overlapping_partitions as recommended by @schneebyte I’ve also decreased the Min running tasks % in my ECS Service to 50% to ensure a task will still be up to converge a replicated cache in the newly created nodes.

It fixed for my use case.

To give a little bit of ‘how to’ for newbies (like me) in clustering in elixir, if you are using releases and wish to disable the prevent overlapping partitions option, you should add a new line in rel/remote.vm.args.eex (or similar) with the followng content:

-kernel prevent_overlapping_partitions false

If you are not using releases and need a quick fix, you can also set an environment variable ELIXIR_ERL_OPTIONS with this setting, such as ELIXIR_ERL_OPTIONS='-kernel prevent_overlapping_partitions false'.

— All posts loaded —

Where Next? Top

Trending in Questions Top

RSP87
I’m working on a project that simulates the bumbl example in the programming phoenix book. It acts almost like an email client. We have a...
New
nseaSeb
Hello, I know there is an approach for handling lists that allows for optimized traversal, but I can’t recall the specific method (somet...
New
brecabral
Documentation While reading the Scoped Routes section, I noticed that the documentation currently refers to a problem without explainin...
New
velrest
So my question is quite simple and i have found no conclusive answer on forum, google or AI. Should we use :erlang.float for Integer to ...
New
asweet-confluent
I recently noticed that Elixir’s Logger defaults its primary log level to :debug when no :logger, :level application configuration is pre...
New
samoloth
Hi, I’ve just set up an application with ash_authentication. There is only magic link strategy for now, so there is no confirmation add o...
New
FlyingNoodle
If a change or preparation module uses Ash.Changeset.get_argument/2 or Ash.Query.get_argument/2 (or any of the other get_argument functio...
New

Other Trending Topics Top

JesseHerrick
Hey, I’m Jesse and I’m the main contributor behind Dexter, a full-featured, lightning-fast Elixir LSP optimized for large codebases. It s...
New
mudasobwa
I am happy to introduce the very α version of the new programming language compiled to BEAM. Welcome Cure. It has literally three kille...
New
marciok
Hi there! We created Gust: A task orchestrator inspired by Airflow. For those who have never heard about Aiflow, it’s a Python-based wor...
New
mhanberg
Hi everyone! The first release candidate for the Expert language server project is now available! We’ve published a press release detai...
New
jimsynz
Beam Bots (or just BB for short) is a framework for building fault-tolerant robotics applications in Elixir using familiar OTP patterns. ...
New
Dmk
Xamal is a deployment tool for Elixir apps that deploys native releases to bare metal servers over SSH. It’s a port of GitHub - basecamp/...
New

Latest on Elixir Forum

We're in Beta

About us Mission Statement

Options

Thread Display Mode




Thread Preview

Skip Thread Previews