tfwright

tfwright

I have a mysterious problem running my app on multiple nodes I hope someone can advise on.

A while back I set up my deployment script to build 2 separate releases of my app two serve behind nginx for load balancing/zero down time deployment purposes (with RELEASE_NAME set to something like “my_app” and “my_app2”). I am managing each app process as a systemd service. This all works great–after a deployment I can tail the logs of each server and observe them both serving requests, and use systemd to stop 1 app instance and observe that nginx is still responsive.

The problem is that when I go to connect to my_app using the remote command I get the error Could not contact remote node app@127.0.0.1, reason: :nodedown. Aborting... Connecting to my_app2 works fine.

I can fix this issue by running service my_app restart. After that, I can connect to both instances.

However, if I run service my_app2 restart the app node again appears to be down.

I am using libcluster to manage the nodes using the following config:

    [
      my_app: [
        strategy: Cluster.Strategy.LocalEpmd
      ]
    ]

my_app env.sh:

export RELEASE_DISTRIBUTION=name
export RELEASE_NODE=my_app@127.0.0.

my_app2 env.sh:

export RELEASE_DISTRIBUTION=name
export RELEASE_NODE=my_app2@127.0.0.

Thanks in advance for any clues or suggestions of things to try!

Showing Posts 1 to 3

kokolegorille

kokolegorille

No RELEASE_COOKIE?

You should have one shared by all the nodes…

And provide it as well when using remote.

tfwright

tfwright OP

I set the RELEASE_COOKIE sys var when I built the releases and it appears to be set properly:

iex(my_app@127.0.0.1)1> (System.get_env()
...(my_app@127.0.0.1)1> |> Enum.map(fn {k, v} -> "#{k}=#{v}" end)
...(my_app@127.0.0.1)1> |> Enum.filter(&String.starts_with?(&1, "RELEASE_")))
["RELEASE_BOOT_SCRIPT_CLEAN=start_clean",
 "RELEASE_ROOT=/home/my_app/apps/my_app/releases/20220812151723/api_v2/_build/prod/rel/my_app",
 "RELEASE_SYS_CONFIG=/home/deployer/apps/my_app/releases/20220812151723/api_v2/_build/prod/rel/my_app/releases/0.1.0/sys",
 "RELEASE_VSN=0.1.0", "RELEASE_DISTRIBUTION=name",
 "RELEASE_COOKIE=[redacted]",
 "RELEASE_VM_ARGS=/home/my_app/apps/my_app/releases/20220812151723/api_v2/_build/prod/rel/my_app/releases/0.1.0/vm.args",
 "RELEASE_BOOT_SCRIPT=start",
 "RELEASE_TMP=/home/my_app/apps/my_app/releases/20220812151723/api_v2/_build/prod/rel/my_app/tmp",
 "RELEASE_COMMAND=start", "RELEASE_MODE=embedded", "RELEASE_NAME=my_app",
 "RELEASE_NODE=my_app@127.0.0.1"]

Note: the output is exactly the same in both nodes, which I can connect to after the restart I mentioned above, aside from the node name/paths.

After restarting the first instance, I can see both nodes:

sudo service  my_app restart
erts-11.1.8/bin/epmd -names
epmd: up and running on port X with data:
name my_app at port Y
name my_app2 at port Z

But after restarting the second instance, the first is gone?

sudo service my_app2 restart
erts-11.1.8/bin/epmd -names
epmd: up and running on port X with data:
name my_app2 at port Z
tfwright

tfwright OP

Forgot to update this, but I eventually discovered this was due to Quantum trying to run jobs on a connected node that itself is not running Quantum (see How to setup quantum against libcluster · Issue #485 · quantum-elixir/quantum-core · GitHub). Switching to the Local run strategy resolved the issue.

— All posts loaded —

Where Next? Top

Trending in Questions Top

katta
I having some trouble figuring out if I have set myself too strict of standards for my production server. Currently I can handle 75% of r...
New
brecabral
Documentation While reading the Scoped Routes section, I noticed that the documentation currently refers to a problem without explainin...
New
achenet
Hello, I’m trying to build a basic Phoenix web-app, and I’d like to use Tailwind. However, when I launch mix phx.server, I get an error...
New
kpanic
Hi everyone, I am toying with the idea of building a “match maker” for giving personal help to people that wants to start coding. I sta...
New
velrest
So my question is quite simple and i have found no conclusive answer on forum, google or AI. Should we use :erlang.float for Integer to ...
New
asweet-confluent
I recently noticed that Elixir’s Logger defaults its primary log level to :debug when no :logger, :level application configuration is pre...
New
mnkhod
So i have been using ash framework for a while and i love it. However currently the issue im having with ash framework is the error handl...
New

Other Trending Topics Top

GenericJam
Edit: 2026 May 15 - This post is archived. Mob is alive!! Main docs: mob v0.7.11 — Documentation A bit of explanation for the slightly c...
New
JesseHerrick
Hey, I’m Jesse and I’m the main contributor behind Dexter, a full-featured, lightning-fast Elixir LSP optimized for large codebases. It s...
New
mudasobwa
I am happy to introduce the very α version of the new programming language compiled to BEAM. Welcome Cure. It has literally three kille...
New
garrison
Hobbes is a low-level distributed database for the Elixir programming language. Hobbes provides a simple, safe, and scalable storage lay...
New
mhanberg
Hi everyone! The first release candidate for the Expert language server project is now available! We’ve published a press release detai...
New
budgie
A little off-topic, but I feel like people here have a good head on their shoulders. I used to be quite good at making software. Was luc...
New

We're in Beta

About us Mission Statement

Options

Thread Display Mode




Thread Preview

Skip Thread Previews