janp

janp

I am wondering whether it is possible to use distributed Erlang plus libcluster in kubernetes? :thinking:

I tried it with the following setup:

StatefulSet (with 3 replicas)

  • example@example-0.example-headless.default.svc.cluster.local
  • example@example-1.example-headless.default.svc.cluster.local
  • example@example-2.example-headless.default.svc.cluster.local

Cluster.Strategy.Kubernetes

  • kubernetes_ip_lookup_mode: :pods
  • mode: :hostname
  • kubernetes_namespace: default
  • kubernetes_selector: app=app
  • kubernetes_service_name: example-headless

config/runtime.exs

config :kernel,
  sync_nodes_optional: [
    :"example@example@example-0.app-headless.default.svc.cluster.local",
    :"example@example@example-1.app-headless.default.svc.cluster.local",
    :"example@example@example-2.app-headless.default.svc.cluster.local"
  ],
  sync_nodes_timeout: 5000

mix.exs

def releases do
  [
    example: [
     reboot_system_after_config: true,
     # ...
    ]
  ]
end

My problem with this setup is, that the nodes cannot connect to each other before starting the application, because their DNS records will only be available when the pods are Ready, which requires the readinessProbe (GET /health/readyz) to be successful, but the readyinessProbe already requires the application to be started :thinking:

So a chicken and egg problem, but maybe I am just doing something wrong :eyes:

Currently, when all pods are started after exceeding the sync_nodes_timeout, the application is running in a cluster, but I would like to have the nodes being connected with each other before the application gets started.

Showing Posts 1 to 2

mudasobwa

mudasobwa

Creator of Cure

I did not check it in Kubernetes, but my cloister does essentially this in AWS ECS.

The idea behind is simple: the library declares the application that connects nodes in its start_phase/3 callback.

You might do the same in your app’s start_phase/3 with libcluster, which would postpone application’s start/2 to return until the nodes are connected.

janp

janp OP

@mudasobwa Thanks for your answer, but I think I found the issue with my setup :slight_smile:


TLDR;

  • changed from Cluster.Strategy.Kubernetes to Cluster.Strategy.Kubernetes.DNSSRV
  • set publishNotReadyAddresses: true for the headless service
  • set podManagementPolicy: Parallel for the StatefulSet

$ kubectl explain service.spec.publishNotReadyAddresses
KIND:     Service
VERSION:  v1

FIELD:    publishNotReadyAddresses <boolean>

DESCRIPTION:
     publishNotReadyAddresses indicates that any agent which deals with
     endpoints for this Service should disregard any indications of
     ready/not-ready. The primary use case for setting this field is for a
     StatefulSet's Headless Service to propagate SRV DNS records for its Pods
     for the purpose of peer discovery. The Kubernetes controllers that generate
     Endpoints and EndpointSlice resources for Services interpret this to mean
     that all endpoints are considered "ready" even if the Pods themselves are
     not. Agents which consume only Kubernetes generated endpoints through the
     Endpoints or EndpointSlice resources can safely assume this behavior.

Setting publishNotReadyAddresses: true allows the pods to find each other via the headless-service before they are ready.

As mentioned in the explaination for the publishNotReadyAddresses option, I also switched from Cluster.Strategy.Kubernetes to Cluster.Strategy.Kubernetes.DNSSRV.

Additionally I set podManagementPolicy to Parallel which allows the StatefulSet to schedule all pods in parallel. The default is OrderedReady which requires one pod to be ready before Kubernetes schedules the next pod, so they are only started one after the other.


Now my application is clustered and started within 15 seconds :partying_face:

— All posts loaded —

Where Next? Top

Trending in Questions Top

RSP87
I’m working on a project that simulates the bumbl example in the programming phoenix book. It acts almost like an email client. We have a...
New
nseaSeb
Hello, I know there is an approach for handling lists that allows for optimized traversal, but I can’t recall the specific method (somet...
New
RemyXRenard
I’m seeing that a list inside a Kino.DataTable will be interpreted as a charlist, even if the Kino.configure() is set to charlists: :as_l...
New
velrest
So my question is quite simple and i have found no conclusive answer on forum, google or AI. Should we use :erlang.float for Integer to ...
New
brecabral
Documentation While reading the Scoped Routes section, I noticed that the documentation currently refers to a problem without explainin...
New
samoloth
Hi, I’ve just set up an application with ash_authentication. There is only magic link strategy for now, so there is no confirmation add o...
New
FlyingNoodle
If a change or preparation module uses Ash.Changeset.get_argument/2 or Ash.Query.get_argument/2 (or any of the other get_argument functio...
New

Other Trending Topics Top

mudasobwa
I am happy to introduce the very α version of the new programming language compiled to BEAM. Welcome Cure. It has literally three kille...
New
marciok
Hi there! We created Gust: A task orchestrator inspired by Airflow. For those who have never heard about Aiflow, it’s a Python-based wor...
New
jimsynz
Beam Bots (or just BB for short) is a framework for building fault-tolerant robotics applications in Elixir using familiar OTP patterns. ...
New
Dmk
Xamal is a deployment tool for Elixir apps that deploys native releases to bare metal servers over SSH. It’s a port of GitHub - basecamp/...
New
netoum
Corex is an accessible, unstyled UI component library for Phoenix that integrates Zag.js state machines using Vanilla JavaScript and Live...
New
webofbits
With AI doing more of the implementation work, I’ve been wondering how much coding I should deliberately keep doing myself. My main conc...
#ai
New

We're in Beta

About us Mission Statement

Options

Thread Display Mode




Thread Preview

Skip Thread Previews