janp
I am wondering whether it is possible to use distributed Erlang plus libcluster in kubernetes? ![]()
I tried it with the following setup:
StatefulSet (with 3 replicas)
example@example-0.example-headless.default.svc.cluster.localexample@example-1.example-headless.default.svc.cluster.localexample@example-2.example-headless.default.svc.cluster.local
Cluster.Strategy.Kubernetes
kubernetes_ip_lookup_mode: :podsmode: :hostnamekubernetes_namespace: defaultkubernetes_selector: app=appkubernetes_service_name: example-headless
config/runtime.exs
config :kernel,
sync_nodes_optional: [
:"example@example@example-0.app-headless.default.svc.cluster.local",
:"example@example@example-1.app-headless.default.svc.cluster.local",
:"example@example@example-2.app-headless.default.svc.cluster.local"
],
sync_nodes_timeout: 5000
mix.exs
def releases do
[
example: [
reboot_system_after_config: true,
# ...
]
]
end
My problem with this setup is, that the nodes cannot connect to each other before starting the application, because their DNS records will only be available when the pods are Ready, which requires the readinessProbe (GET /health/readyz) to be successful, but the readyinessProbe already requires the application to be started ![]()
So a chicken and egg problem, but maybe I am just doing something wrong ![]()
Currently, when all pods are started after exceeding the sync_nodes_timeout, the application is running in a cluster, but I would like to have the nodes being connected with each other before the application gets started.
Trending in Questions
Other Trending Topics
Categories:
Sub Categories:
Forums
Popular Tags
- #ecto
- #liveview
- #troubleshooting
- #learning-elixir
- #library
- #deployment
- #erlang
- #testing
- #genserver
- #mix
- #absinthe
- #remote-other
- #otp
- #plug
- #how-to-question
- #macros
- #postgres
- #elixirconf
- #channels
- #exunit
- #discussion
- #code-sync
- #podcasts
- #javascript
- #onsite
- #dialyzer
- #docker
- #authentication
- #umbrella
- #full-time-contract
- #podcasts-by-brainlid
- #ecto-query
- #elixirconf-us
- #ai
- #blog-post
- #elixir-ls
- #phoenix_html
- #iex
- #graphql
- #genstage
- #websockets
- #supervisor
- #advent-of-code
- #distillery
- #processes
- #api
- #forms
- #hex
- #security
- #metaprogramming










Showing Posts 1 to 2- Show Best Posts
- Show All (oldest first)
- Show All (newest first)
mudasobwa
I did not check it in Kubernetes, but my
cloisterdoes essentially this in AWS ECS.The idea behind is simple: the library declares the application that connects nodes in its
start_phase/3callback.You might do the same in your app’s
start_phase/3withlibcluster, which would postpone application’sstart/2to return until the nodes are connected.janp
@mudasobwa Thanks for your answer, but I think I found the issue with my setup
TLDR;
Cluster.Strategy.KubernetestoCluster.Strategy.Kubernetes.DNSSRVpublishNotReadyAddresses: truefor the headless servicepodManagementPolicy: Parallelfor theStatefulSetSetting
publishNotReadyAddresses: trueallows the pods to find each other via the headless-service before they are ready.As mentioned in the explaination for the
publishNotReadyAddressesoption, I also switched fromCluster.Strategy.KubernetestoCluster.Strategy.Kubernetes.DNSSRV.Additionally I set
podManagementPolicytoParallelwhich allows theStatefulSetto schedule all pods in parallel. The default isOrderedReadywhich requires one pod to be ready before Kubernetes schedules the next pod, so they are only started one after the other.Now my application is clustered and started within 15 seconds