<turbo-stream action="append" target="posts_list"><template>    <div class="postbit" id="339438" data-post-id="339438">
  <section>
    <div class="post-wrap">


					<div class="post-header">
		        <div class="user-avatar">
		          <img alt="MarthinL" src="/assets/icons/user-9f439610.png" width="120" height="120" />
		        </div>
					
						<div class="user-details">
		          <div class="user-name">
		            <h3>
                  MarthinL
                    <span class="op-star" title="Thread Starter">
                      <img alt="OP" class="op-star-icon" src="/assets/thread-icons/thread-icon-thread-starter-df91e872.png" />
                    </span>
                  </h3>
		          </div>
						
						</div>
					
					</div>

	        <div class="thread-main">
	            <div class="post-body" data-turbo="false">
								<p>Appologies, I found this reply apparently un-sent when I visited forum again. Don’t know if it was sent and cancelled again because it came from me or I never pushed the button.</p>
<p>Hey <a class="mention" href="/u/julismz" rel="nofollow">@julismz</a>,</p>
<p>It certainly is a journey, not one event or decision. I’m sory to hear of your performance hurdles and poor results, but I’d need to get quite a bit deeper into what you’ve really been experimenting with.</p>
<p>There never was and still isn’t any nginx in my solution. How/why did nginx enter the mix for you?</p>
<p>As for the customers per node question you’re right on the money identifying that as the key question. It’s a lot like engineering a race-car’s drag coefficient. You can have an idea of the Cd for simple shapes but when many shapes overlap and interact it becoes a rabbit hole which if you do make it out of will leave you with unreliable predictions anyway.</p>
<p>The only application that performs without bugs is the null application. From there, everything your app has to do in order to add value will lower it’s performance. It’s up to you as designer to try avoid giving up large chunks of performance for no or questionionable value to the end-user. You cannot increase performance, only decrease it less by making smarter choices which avoids unnecessary work. In most cases, unnecessary work equates duplicated work so that’s usually a good place to start.</p>
<p>Perhaps your application is somple enough for you to know in advance what every user will be doing which means for you the only parameter becoes how many users are active. If that’s the case we can discuss strategies for that but there are probably many oters that can help you more than I can.</p>
<p>Each of my users use my application in their own unique way which I can neither predict nor directly and it varies day by day. That variance impacts my “optimal” ration between nodes and users far more than my code does so as long as I don’t waste computing cycles I’m doing as well as I can.</p>
<p>That then engages the third element we can call pre-scaling if you want. As you might have seen above in this thread I’ve been criticised rather heavily for it as premature optimisation. The essense of it is to design your application assuming it will need to scale bigger than you could have imagined and then implement it at the smallest possible scale you can pull off. The rationale is simple. Even if you cannot predict how many live customers a single node will service you can measure how it’s doing and if it starts to struggle you can add additional nodes without delay. You can only do that if even for your initial version you’ve made provision for the application to run on multiple nodes and in multiple regions. If you don’t do that from the start you will get trapped. Your users will demand better performance but to give them that you need to redesign your app or find a way to give your server more resources (vertical scaling) which will in the end put you exactly where the cloud providers want you - dependent on them for scaling.</p>
<p>I’ve seen too many designers and managerial types with unrealistic expectations of the agile and “first make it work, then make it fast” to not say anything about it when asked. I’ve even worked with a company selling an “Agile product” that declared war against Architects as the enemy of Agile. But as an Architect using Agile I did things with their product they couldn’t believe, so I developed this analogy: When tasked with bulding a dam it’s easy to find a hole in the ground, fill it with water and say that wasn’t hard at all, now we just need to evolve it into a dam at the desired scale. If you cannot see the issues with evolving a filled dam or the simple matter of where the water will come from if your initial hole in the ground is not in an opportune geographical location, you deserve the disappointment that will follow. If you need a dam, your starting point is to find the right location, then to design whole dam, then the way to divert the water while you build the dam and only then to start construction of the dam even if you build it in phases on top of the full foundation started in the right spot.</p>
<p>Design your application’s value to your customers and let that guide you as to how to deliver that value at a lower cost than what your customers pay you for it.</p>
<p>A lot of Kubernetes’ facilities are there to accomodate legacy applications such as the monolythic web servers of yesteryear and “micro service” stuff written in C++, Java and go. Don’t get distracted by all that noise. Erlang, Elixir and Phoenix already provides a tonne of facilities which those applications never had and had to either implement at application level or get from a container environment when they needed to head for the clouds. As a result, a lot of the concepts like nginx (as Ingress controller) and NodePorts   have wide-ranging capabilities in order to adapt to the enormous variance with which application designers had gone about they business before they involved Kubernetes. Your Phoenix app needs surprisingly little mangling to do well in Kubernetes.</p>
<p>Away from the public cloud providers on your own clusters MetalLB is your friend and you have a few good options as to how to create your kubernetes cluster. Once you are using a public cloud, use their kubernetes stacks and load balancers directly. It’s silly to implement your own stack on cloud-based servers.</p>
<p>You also need to choose a database strategy and a storage strategy to suit your cloud-agnostic ideals. It largely depends on your applications’s actual needs and the capabilities of your own hardware. For my purposes I chose to a “standalone” PostgreSQL deployment in Kubernetes using Percona’s Operator for PostgreSQL local to each regional kubernetes cluster and share data betweeen them at the application layer only. As for storage I’ve gone and stuck with Ceph because that’s what best suits my hardware but because Kubernetes abstracts that so well through storage classes and providers I’d use whatever the cloud provider offers where I need to use a public cloud’s kubernetes stack.</p>
<p>I don’t know if I’m really answering your questions and I don’t even know if all your questions have the kind of answers you’d like them to have. Both cloud-native and distributed computing (very different things, but related anyway) can overwhelm you exactly like the monsters they’re reported to be. They certainly are powerful but they’re mostly misunderstood which makes them scary. Once you know enough about them they’re incrediby good friends to have on your side. It just takes a little work and empathy to get to know and understand them.</p>
<p>Good luck with your journey and dpon’t hesitate to reach out (again) if you feel stuck or uncertain. Just bear in mind that the bulk of the information you will encounter on the interwebs have been written for purposes other than your own. This too. My purpose is the hope spark a kinship among fellow travellers on the good ship Phoenix sailing towards large scale independent federated systems such as my own. At some point we’re going to need others who look at the problems we’ll face through the same colour lenses.</p> 
	            </div>

	            <div class="base-line">
	                <div class="thread-counters">
	                    <span class="thread-count count-likes js-likers-trigger" title="Likes" data-post-id="339438" data-batch-url="/posts/batch_likers">
                        0
                      </span>
                      <!-- <span class="thread-count js-solved-indicator" title="Marked as solution"></span> -->
	                </div>
	                <div class="go-to-post">
	                  <a title="Go to post" alt="Go to post" href="https://forum.elixirforum.com/t/production-phoenix-liveview-on-cloud-agnostic-kubernetes/58988/12">Post #11</a>
	                </div>
	            </div>
              <div id="likers-container-339438" 
                   class="likers-container"
                   data-first-post="false"
                   data-batch-url="/posts/batch_likers">
                   <div class="likers-placeholder" 
                     data-likers-post-id="339438"
                     data-batch-url="/posts/batch_likers">
                  <div class="post-likers"></div>
                </div>
              </div>
	        </div>
			

    </div>

    <div class="triangle-top-right type-standard-post cat-standard-post" title="Post #11"></div>
  </section>
</div>
    <div class="postbit" id="339442" data-post-id="339442">
  <section>
    <div class="post-wrap">


					<div class="post-header">
		        <div class="user-avatar">
		          <img alt="MarthinL" src="/assets/icons/user-9f439610.png" width="120" height="120" />
		        </div>
					
						<div class="user-details">
		          <div class="user-name">
		            <h3>
                  MarthinL
                    <span class="op-star" title="Thread Starter">
                      <img alt="OP" class="op-star-icon" src="/assets/thread-icons/thread-icon-thread-starter-df91e872.png" />
                    </span>
                  </h3>
		          </div>
						
						</div>
					
					</div>

	        <div class="thread-main">
	            <div class="post-body" data-turbo="false">
								<aside class="quote no-group" data-username="DaAnalyst" data-post="11" data-topic="58988">
<div class="title">
<div class="quote-controls"></div>
<img alt="" width="24" height="24" src="https://forum.elixirforum.com/letter_avatar_proxy/v4/letter/d/278dde/48.png" class="avatar"> DaAnalyst:</div>
<blockquote>
<p>Providing support for an arbitrary identifier (a <code>token_id</code> or a <code>user_id</code> in our case) for stickiness would solve this issue, but K8s doesn’t seem to care much about stateful apps.</p>
</blockquote>
</aside>
<p>Keeping state in a web app is a taboo from the monolithic web server era where it was deemed simply too resource intensive for anything beyond perhaps a corporate intranet application. It’s no surprise that K8s’ default model is aligned with stateless web services.</p>
<p>By virtue of the BEAM’s extremely efficient lightweight processes with states as small as you like Phoenix in Elixir has made a mockery of that old taboo. LiveView exploited it even further to the point where it took things even corporate intranet applications with limited users wouldn’t dream of and made it easy and effective for hundreds of thousands of users per server node.</p>
<p>Not everyone in the industry has recognised the paradigm shift and/or adapted to its implications. Kubernetes wasn’t conceived with Erlang/OTP/Cowbow/Phoenix/LiveView applications and backends in mind but rather for a world where stateless applications are the norm.</p>
<p>Luckily the story does not end there. Kubernetes also provides for “Stateful Sets” which as the name suggests are specifically for applications that maintains state. In traditional web-service terms it was only the database that was “qualified” to carrry state so in many minds and articles “state” became conflated with persistant data, but it’s not the same thing.</p>
<p>Because Phoenix and LiveView apps maintain state it is best to host them in Kubernetes as Stateful Sets. You’ll still need to do some work within your application to make it cluster-aware (the easy part because of libcluster) and for your specific needs choose and implement handing over sessions between pods in the set and nodes in the cluster.</p>
<p>You could also choose to do no handover at all because LiveView applications use “slow poll” which establishes and maintains a connection between the server and the browser. Most if not all load-balancers would ensure that established sessions (confusingly called states at the network level) are preserved meaning that when JS code in the browser sends something back on that established connection it would by default go back to the exact pod and process that established the session. That only breaks when the user reloads the page and the liveview session needs to be established again.</p>
<p>It can be argued that when a user reloads a page it’s a good time to give the load-balancer a chance to pick a new node, pod and process to service that session. Having extremely long-running sessions (as in hours or days) might not what you’re looking for since it may load some nodes more than others. Here’s why. Unless you’ve implemented  a mechanism to feed back how busy each of your nodes are to the load balancing algorithm it’s going to come down to something random. Randomness and round-robin techniques are quite OK when you’re dealing with large volumes of similarly sized sessions. But if you’re too eager to insist that each client is served by the same pod on the same node it got allocated to randomly you’re going against the natural self-levelling order of things and some clients might end up getting served by a pod that’s very busy while other pods are idle.</p>
<p>The choice you app needs to make is based on the cost of initialising a user’s session. (That was also the case in the old monolithic server use-case, except that most session-based implementations ended up having to rebuild the session for every request coming in from the client, which is what made it so impossibly expensive on resources). If that only happens occasionally like at login or reload it might be cheap enough to fill in the user’s state from the database, but if it’s quite a demanding job to do that it might be better to first see if another pod in the system has state loaded for that user and obtain it from there.</p>
<p>From there you need to decide whether to redirect traffic to the original pod or transfer the state to the new pod and resume serving that user from there (i.e. kill the old slow-poll session and establish a new one from the new pod). For the reasons I’ve explained above and because redirection can be so tricky depending on your load-balancer environment whihc may vary between stacks, my choice was to move the session “behind the scenes” using direct calls to the node that used to serve the client.</p>
<p>[Edit]<br>
I should also add, so I will. Unless you do so for well-understood reasons unrelated to performance its always advised to stick to one BEAM instance per physical server.</p>
<p>BEAM is encredibly good at what it does, which is to “run” massive amounts of very lightweight processes interacting with each other as efficient as possible. You’re unlikely to improve on that by running multiple instances on the same physical hardware. If you only have one server and need multiple pods and nodes to help with ilve upgrades and such, sure, do that, but as a rule of thumb only cluster when you can do so across physically separate servers connected by high-speed LAN.</p> 
	            </div>

	            <div class="base-line">
	                <div class="thread-counters">
	                    <span class="thread-count count-likes js-likers-trigger" title="Likes" data-post-id="339442" data-batch-url="/posts/batch_likers">
                        1
                      </span>
                      <!-- <span class="thread-count js-solved-indicator" title="Marked as solution"></span> -->
	                </div>
	                <div class="go-to-post">
	                  <a title="Go to post" alt="Go to post" href="https://forum.elixirforum.com/t/production-phoenix-liveview-on-cloud-agnostic-kubernetes/58988/13">Post #12</a>
	                </div>
	            </div>
              <div id="likers-container-339442" 
                   class="likers-container"
                   data-first-post="false"
                   data-batch-url="/posts/batch_likers">
                   <div class="likers-placeholder" 
                     data-likers-post-id="339442"
                     data-batch-url="/posts/batch_likers">
                  <div class="post-likers"></div>
                </div>
              </div>
	        </div>
			

    </div>

    <div class="triangle-top-right type-standard-post cat-standard-post" title="Post #12"></div>
  </section>
</div>
    <div class="postbit" id="339448" data-post-id="339448">
  <section>
    <div class="post-wrap">


					<div class="post-header">
		        <div class="user-avatar">
		          <img alt="MarthinL" src="/assets/icons/user-9f439610.png" width="120" height="120" />
		        </div>
					
						<div class="user-details">
		          <div class="user-name">
		            <h3>
                  MarthinL
                    <span class="op-star" title="Thread Starter">
                      <img alt="OP" class="op-star-icon" src="/assets/thread-icons/thread-icon-thread-starter-df91e872.png" />
                    </span>
                  </h3>
		          </div>
						
						</div>
					
					</div>

	        <div class="thread-main">
	            <div class="post-body" data-turbo="false">
								<aside class="quote no-group" data-username="eclark" data-post="9" data-topic="58988">
<div class="title">
<div class="quote-controls"></div>
<img alt="" width="24" height="24" src="https://forum.elixirforum.com/letter_avatar_proxy/v4/letter/e/8edcca/48.png" class="avatar"> eclark:</div>
<blockquote>
<p>I haven’t found a good cloud-agnostic way to do that yet</p>
</blockquote>
</aside>
<p>I’d be rather surprised if you do find a one-size-fits-all solution. I also don’t see why you’d want the initial cold HTTP request to go back to the same VM rather than to whichever VM is best suited to servicing the request when it arrives. The previous serving VM might no longer be in the best position.</p>
<p>The issue is this: How was the VM to serve a specific client session selected in the first place? If it was random or round-robin then the only possible advantage of that VM over any other would be that it (might be) cheaper for that VM to resurrect the user session than it would be for a fresh VM to load it from the database. Since that so dependent on application specific conditions there’s no way to assume anything on behalf of an application. The least you could (or should) do is to allow the load balancing layer that made the decision in the first place to make the call again, thereby distributing the fresh cold HTML requests evenly across the available VMs. Then it’s an application decision whether to retrieve the old state from whoever previously served the same customer or rebuild it from persistent storage.</p>
<p>A better approach would be to base the distribution of load on the current load on each VM or more accurately which VM would be able to get to respoding to the request soonest. But as we’ve seen, that equation for stateful applications is not simple because it varies depending on what it would require from each VM to have the required state in memory to start processing with. But if could resolve that complexity somehow and control the load-balancer with that data it would solve the whole problem - for as long as it remains best for a cold request to go back to a specific VM it will be what happens but the moment it’s better for the system as a whole and the individual client’s request to get served by anothe VM that too will happen seamlessly.</p>
<p>I don’t have all the answers yet as to how to control MetalLB or any of the public cloud providers’ load balancer yet (and using something like nginx reverse proxy or even a load balancing app in Phoenix is a step backwards again), so that level of optimisation remains out of reach for me, but that’s the direction in which I believe sensible answers lies. Getting cold requests to return to some randomly chosen VM every time is not only of little practical value but could work against overall application effeciency and user experience.</p> 
	            </div>

	            <div class="base-line">
	                <div class="thread-counters">
	                    <span class="thread-count count-likes js-likers-trigger" title="Likes" data-post-id="339448" data-batch-url="/posts/batch_likers">
                        0
                      </span>
                      <!-- <span class="thread-count js-solved-indicator" title="Marked as solution"></span> -->
	                </div>
	                <div class="go-to-post">
	                  <a title="Go to post" alt="Go to post" href="https://forum.elixirforum.com/t/production-phoenix-liveview-on-cloud-agnostic-kubernetes/58988/14">Post #13</a>
	                </div>
	            </div>
              <div id="likers-container-339448" 
                   class="likers-container"
                   data-first-post="false"
                   data-batch-url="/posts/batch_likers">
                   <div class="likers-placeholder" 
                     data-likers-post-id="339448"
                     data-batch-url="/posts/batch_likers">
                  <div class="post-likers"></div>
                </div>
              </div>
	        </div>
			

    </div>

    <div class="triangle-top-right type-standard-post cat-standard-post" title="Post #13"></div>
  </section>
</div>
    <div class="postbit" id="339470" data-post-id="339470">
  <section>
    <div class="post-wrap">


					<div class="post-header">
		        <div class="user-avatar">
		          <img alt="eclark" src="/assets/icons/user-9f439610.png" width="120" height="120" />
		        </div>
					
						<div class="user-details">
		          <div class="user-name">
		            <h3>
                  eclark
                  </h3>
		          </div>
						
						</div>
					
					</div>

	        <div class="thread-main">
	            <div class="post-body" data-turbo="false">
								<blockquote>
<p>Can you formulate better on what is the actual issue?</p>
</blockquote>
<p>Distribution adds latency. The first request in will compute assigns and the rest of the process state. Say the load balancer sent that to host <code>foo</code>. So <code>foo</code> starts that process. Now if the load balancer sends the web socket start request to <code>bar</code> with distribution <code>bar</code> will have to ask <code>foo</code> for the response, before passing it  on. That adds latency and additional points of failure. <code>bar</code> can be overloaded, have a bad NIC or go down for maintence.</p>
<p>Yes distribution handles some of this, but no it’s not the total solution.</p>
<aside class="quote no-group" data-username="DaAnalyst" data-post="11" data-topic="58988">
<div class="title">
<div class="quote-controls"></div>
<img alt="" width="24" height="24" src="https://forum.elixirforum.com/letter_avatar_proxy/v4/letter/d/278dde/48.png" class="avatar"> DaAnalyst:</div>
<blockquote>
<p>Providing support for an arbitrary identifier (a <code>token_id</code> or a <code>user_id</code> in our case) for stickiness would solve this issue, but K8s doesn’t seem to care much about stateful apps.</p>
</blockquote>
</aside>
<p>You’re totally correct. A great session needs access to a user or something similar. That allows blue/green deployments and other fun parts to be integrated into the HTTP routing. Using header/cookie values to determine the <code>server_shard</code> or <code>consistent_hash_ring_location</code>, and then remembering that shard in a signed cookie will likely be our first pass. If we have a good set of header and cookies to use as the default it should provide a good baseline.</p>
<aside class="quote no-group" data-username="MarthinL" data-post="14" data-topic="58988">
<div class="title">
<div class="quote-controls"></div>
<img alt="" width="24" height="24" src="https://forum.elixirforum.com/letter_avatar_proxy/v4/letter/m/3da27b/48.png" class="avatar"> MarthinL:</div>
<blockquote>
<p>I don’t have all the answers yet as to how to control MetalLB or any of the public cloud providers’ load balancer yet</p>
</blockquote>
</aside>
<p>Controlling only the load balancer at layer four will not be enough since it doesn’t know enough about the service health or the user who created the request. We will need most of a service mesh for stable routing with load and the health of downstream services. From previous experience, it takes many systems working together to make reactive long poll work at scale.</p>
<p>At Batteries Included, we’re using Istio for service mesh and mTLS, Istio ingresss gateway for layer 7 HTTP(s) input, cert-manager + internal CA for ssl, and some custom envoy/istio smarts (Istio takes WASM compiled plugins so its very extensible). We also use MetalLB for lower layers or the cloud provider’s load balancer. That gives us something pretty close to the state of the art (none of the publicly available cloud agnostic solutions can do dsr or direct service return, which is a bummer).</p>
<ul>
<li><a href="https://static.googleusercontent.com/media/research.google.com/en//pubs/archive/44824.pdf" rel="noopener nofollow ugc">https://static.googleusercontent.com/media/research.google.com/en//pubs/archive/44824.pdf</a></li>
<li><a href="https://github.com/facebookincubator/katran" class="inline-onebox" rel="noopener nofollow ugc">GitHub - facebookincubator/katran: A high performance layer 4 load balancer · GitHub</a></li>
<li><a href="https://www.usenix.org/system/files/osdi23-saokar.pdf" rel="noopener nofollow ugc">https://www.usenix.org/system/files/osdi23-saokar.pdf</a></li>
</ul> 
	            </div>

	            <div class="base-line">
	                <div class="thread-counters">
	                    <span class="thread-count count-likes js-likers-trigger" title="Likes" data-post-id="339470" data-batch-url="/posts/batch_likers">
                        1
                      </span>
                      <!-- <span class="thread-count js-solved-indicator" title="Marked as solution"></span> -->
	                </div>
	                <div class="go-to-post">
	                  <a title="Go to post" alt="Go to post" href="https://forum.elixirforum.com/t/production-phoenix-liveview-on-cloud-agnostic-kubernetes/58988/15">Post #14</a>
	                </div>
	            </div>
              <div id="likers-container-339470" 
                   class="likers-container"
                   data-first-post="false"
                   data-batch-url="/posts/batch_likers">
                   <div class="likers-placeholder" 
                     data-likers-post-id="339470"
                     data-batch-url="/posts/batch_likers">
                  <div class="post-likers"></div>
                </div>
              </div>
	        </div>
			

    </div>

    <div class="triangle-top-right type-standard-post cat-standard-post" title="Post #14"></div>
  </section>
</div>
    <div class="postbit" id="339472" data-post-id="339472">
  <section>
    <div class="post-wrap">


					<div class="post-header">
		        <div class="user-avatar">
		          <img alt="MarthinL" src="/assets/icons/user-9f439610.png" width="120" height="120" />
		        </div>
					
						<div class="user-details">
		          <div class="user-name">
		            <h3>
                  MarthinL
                    <span class="op-star" title="Thread Starter">
                      <img alt="OP" class="op-star-icon" src="/assets/thread-icons/thread-icon-thread-starter-df91e872.png" />
                    </span>
                  </h3>
		          </div>
						
						</div>
					
					</div>

	        <div class="thread-main">
	            <div class="post-body" data-turbo="false">
								<aside class="quote no-group" data-username="eclark" data-post="15" data-topic="58988">
<div class="title">
<div class="quote-controls"></div>
<img alt="" width="24" height="24" src="https://forum.elixirforum.com/letter_avatar_proxy/v4/letter/e/8edcca/48.png" class="avatar"> eclark:</div>
<blockquote>
<p>At Batteries Included, we’re using Istio for service mesh and mTLS, Istio ingresss gateway for layer 7 HTTP(s) input, cert-manager + internal CA for ssl, and some custom envoy/istio smarts (Istio takes WASM compiled plugins so its very extensible). We also use MetalLB for lower layers or the cloud provider’s load balancer. That gives us something pretty close to the state of the art (none of the publicly available cloud agnostic solutions can do dsr or direct service return, which is a bummer).</p>
</blockquote>
</aside>
<p>That sounds the oppisite of cloud-agnostic, like you’re being or trying to be the cloud yourself and get a piece of every pie that way?</p>
<p>My own attempts to estract value from the service mesh concept brought me to conclude that there is too much overlap and competing notions between service mesh and the baseline facilities everyone on this forum is used to where remote procedure calls and process identity are integrally parts of the environment. All the examples I could find were in go because that environment is so bare it actually adds value. But in the Erlang/Elixir/Phoenix domain all it added was complexity. That’s why I chose to walk away from that world and approach my distributed solution in a much simpler and effective manner.</p> 
	            </div>

	            <div class="base-line">
	                <div class="thread-counters">
	                    <span class="thread-count count-likes js-likers-trigger" title="Likes" data-post-id="339472" data-batch-url="/posts/batch_likers">
                        0
                      </span>
                      <!-- <span class="thread-count js-solved-indicator" title="Marked as solution"></span> -->
	                </div>
	                <div class="go-to-post">
	                  <a title="Go to post" alt="Go to post" href="https://forum.elixirforum.com/t/production-phoenix-liveview-on-cloud-agnostic-kubernetes/58988/16">Post #15</a>
	                </div>
	            </div>
              <div id="likers-container-339472" 
                   class="likers-container"
                   data-first-post="false"
                   data-batch-url="/posts/batch_likers">
                   <div class="likers-placeholder" 
                     data-likers-post-id="339472"
                     data-batch-url="/posts/batch_likers">
                  <div class="post-likers"></div>
                </div>
              </div>
	        </div>
			

    </div>

    <div class="triangle-top-right type-standard-post cat-standard-post" title="Post #15"></div>
  </section>
</div>
    <div class="postbit" id="339477" data-post-id="339477">
  <section>
    <div class="post-wrap">


					<div class="post-header">
		        <div class="user-avatar">
		          <img alt="eclark" src="/assets/icons/user-9f439610.png" width="120" height="120" />
		        </div>
					
						<div class="user-details">
		          <div class="user-name">
		            <h3>
                  eclark
                  </h3>
		          </div>
						
						</div>
					
					</div>

	        <div class="thread-main">
	            <div class="post-body" data-turbo="false">
								<aside class="quote no-group" data-username="MarthinL" data-post="16" data-topic="58988">
<div class="title">
<div class="quote-controls"></div>
<img alt="" width="24" height="24" src="https://forum.elixirforum.com/letter_avatar_proxy/v4/letter/m/3da27b/48.png" class="avatar"> MarthinL:</div>
<blockquote>
<p>That sounds the oppisite of cloud-agnostic</p>
</blockquote>
</aside>
<p>It runs on any cloud that runs Kubernetes, on already-running Kubernetes clusters, and on a single machine with a docker-like API. It is cloud agnostic.</p>
<aside class="quote no-group" data-username="MarthinL" data-post="16" data-topic="58988">
<div class="title">
<div class="quote-controls"></div>
<img alt="" width="24" height="24" src="https://forum.elixirforum.com/letter_avatar_proxy/v4/letter/m/3da27b/48.png" class="avatar"> MarthinL:</div>
<blockquote>
<p>yourself and get a piece of every pie that way?</p>
</blockquote>
</aside>
<p>You’re assuming ill intent rather than considering that I have a different point of view, which makes for a poor conversation.</p>
<aside class="quote no-group" data-username="MarthinL" data-post="16" data-topic="58988">
<div class="title">
<div class="quote-controls"></div>
<img alt="" width="24" height="24" src="https://forum.elixirforum.com/letter_avatar_proxy/v4/letter/m/3da27b/48.png" class="avatar"> MarthinL:</div>
<blockquote>
<p>My own attempts to estract value from the service mesh concept brought me to conclude that there is too much overlap and competing notions</p>
</blockquote>
</aside>
<p>I have seen different in production and at scale. For a truly scalable and distributed balancing, there’s a need for layer 4 balancing (MetalLb can do this sometimes, but not in the cloud since arp is controlled so you have to use the cloud for this sometimes), a need for layer 7 (session based routing, internal network outages, etc all require a deep understanding of the complete Kubernetes cluster), and application load. All of that adds up to a service mesh being very useful in this case of reactive hosting at any scale.</p>
<p>Then, it’s worth its weight in gold when debugging an outage at cloud scale. These systems at scale aren’t a monolith; they are many different systems used by many other teams. Service mesh allows common ground in metrics and tracing. (Assuming all of this is done well and tuned well and used at appropriate scale)</p>
<p>After that, the cloud or any hosting location with shared fiber or network hardware you didn’t build yourself is a very hostile place these days, and it’s great to have security built in. A well-configured service mesh will use mTLS and SSL cert identities for lots of security goodness (forward secrecy from eavesdropping, crypto attestation, identity validation, etc.). That security, along with the debugging and metrics provided, is worth much when facing true adversaries.</p>
<p>After that service meshes are a great way to get shadow traffic. Meaning they can be used as development tools for testing new versions before deployment, or debugging why performance is changing.</p>
<p>From there, service meshes are helpful for very advanced distributed systems paradigms, such as scatter-gather machine learning and data provenance attestation. Since the full chain of all request, responses are running the same mesh that stuff gets much more manageable.</p>
<p>Each of those is a complex issue that service meshes make tractable. They are faced in production at scale today. Not everyone needs to solve every one of those at the same time and trying to would make your production systems an unstable nightmare. However, when needed and when applied in the correct way, service meshes are a great  tool.</p> 
	            </div>

	            <div class="base-line">
	                <div class="thread-counters">
	                    <span class="thread-count count-likes js-likers-trigger" title="Likes" data-post-id="339477" data-batch-url="/posts/batch_likers">
                        2
                      </span>
                      <!-- <span class="thread-count js-solved-indicator" title="Marked as solution"></span> -->
	                </div>
	                <div class="go-to-post">
	                  <a title="Go to post" alt="Go to post" href="https://forum.elixirforum.com/t/production-phoenix-liveview-on-cloud-agnostic-kubernetes/58988/17">Post #16</a>
	                </div>
	            </div>
              <div id="likers-container-339477" 
                   class="likers-container"
                   data-first-post="false"
                   data-batch-url="/posts/batch_likers">
                   <div class="likers-placeholder" 
                     data-likers-post-id="339477"
                     data-batch-url="/posts/batch_likers">
                  <div class="post-likers"></div>
                </div>
              </div>
	        </div>
			

    </div>

    <div class="triangle-top-right type-standard-post cat-standard-post" title="Post #16"></div>
  </section>
</div>
    <div class="postbit" id="339558" data-post-id="339558">
  <section>
    <div class="post-wrap">


					<div class="post-header">
		        <div class="user-avatar">
		          <img alt="MarthinL" src="/assets/icons/user-9f439610.png" width="120" height="120" />
		        </div>
					
						<div class="user-details">
		          <div class="user-name">
		            <h3>
                  MarthinL
                    <span class="op-star" title="Thread Starter">
                      <img alt="OP" class="op-star-icon" src="/assets/thread-icons/thread-icon-thread-starter-df91e872.png" />
                    </span>
                  </h3>
		          </div>
						
						</div>
					
					</div>

	        <div class="thread-main">
	            <div class="post-body" data-turbo="false">
								<aside class="quote no-group" data-username="eclark" data-post="15" data-topic="58988">
<div class="title">
<div class="quote-controls"></div>
<img alt="" width="24" height="24" src="https://forum.elixirforum.com/letter_avatar_proxy/v4/letter/e/8edcca/48.png" class="avatar"> eclark:</div>
<blockquote>
<p>Distribution adds latency.</p>
</blockquote>
</aside>
<p>Every process step, buffer latch and meter of conductor adds latency. But if using distribution results in increased overall latency you’re either doing it wrong or shouldn’t be using distribution. If you send all your traffic through some choke point to orchestrate the distribution you’re going to struggle for sure. But if you’ve correctly utilised distribution principles in your application the possible extra steps to get the session state to the serving VM will be offset by a far greater reduction in latency because you’re responding to the request from a local server rather than one on the other side of the world.</p>
<p>Distributed processing is a strange animal - when used where something inherent in the problem space definitively determines where each piece of processing must happen the objectives are clearly defined, the metrics are easily undersood and applied and the whole implementation becomes simple and performant.</p>
<p>In general terms, if it’s not patently obvious which cluster or node would be in a position to service the request so much quicker than any other that the extra steps to make that possible is dwarfed by the saved time then distributed processing does not belong in that solution. It’s often a dead giveaway that you’re on the wrong track with distribution the moment you have to rely on round-robin or random allocation of work load to processors.</p>
<p>Too many have caught out by the apparent opportunity to improve performance with distribution but then having to come up with some determinant by which to distribute the load. That rarely yields the expected results and usually additional compleity outweighs the performance gains if there are any.</p>
<p>Don’t conflate distributed processing and load balancing or parallel processing and be highly suspicious of anyone offering one-size-fits-all solutions for distributed processing. The proper way to distribute processing and/or data storage is entirely application specific and therefore cannot be offered as a service by a platform without a means for the application to install its specific domain knowledge into the distribution algorithm.</p>
<p>The service mesh concept is great for its intended audience, but with my single Phoenix app I’m not part of that audience. A lot of service mesh is about service discovery, but once I know how to reach each of my clusters securely I know exactly what services they offer because they are copies of me. Another big part of service mesh as you mentioned is common observability which is awesome when you have services written in different languages and frameworks but when your entire suite of applications runs in Elixir the tools at your disposal through that are more than enough and comes without the overhead and complexities of additional layers of libraries and concepts which are not native to Elixir.</p>
<p>Bottom line, I trust your anchor tenant(s) are happy that your product is aligned with their needs but its value to me would be marginal at best and more likely negative. Everyone’s use case is different so there might be some on this forum who are involved in multi-technology environments and therefore could find your product of value, but in general the founding tennets of LiveView specifically had included the vision of avoiding the cost and complexity of working in different languages by making it viable to write everything, front end and back end, in one language and framework. Those like myself who are using Phoenix and Liveview for that purpose - one cohesive environment for everything are unlikely to step outside those boundaries any time soon. Your target audience are the unfortunate ones whose working environments made it impossible for them to stay inside the single language world. So thanks for helping taking care of them. With any luck, I won’t become one of them any time soon.</p> 
	            </div>

	            <div class="base-line">
	                <div class="thread-counters">
	                    <span class="thread-count count-likes js-likers-trigger" title="Likes" data-post-id="339558" data-batch-url="/posts/batch_likers">
                        0
                      </span>
                      <!-- <span class="thread-count js-solved-indicator" title="Marked as solution"></span> -->
	                </div>
	                <div class="go-to-post">
	                  <a title="Go to post" alt="Go to post" href="https://forum.elixirforum.com/t/production-phoenix-liveview-on-cloud-agnostic-kubernetes/58988/18">Post #17</a>
	                </div>
	            </div>
              <div id="likers-container-339558" 
                   class="likers-container"
                   data-first-post="false"
                   data-batch-url="/posts/batch_likers">
                   <div class="likers-placeholder" 
                     data-likers-post-id="339558"
                     data-batch-url="/posts/batch_likers">
                  <div class="post-likers"></div>
                </div>
              </div>
	        </div>
			

    </div>

    <div class="triangle-top-right type-standard-post cat-standard-post" title="Post #17"></div>
  </section>
</div>
    <div class="postbit" id="339561" data-post-id="339561">
  <section>
    <div class="post-wrap">


					<div class="post-header">
		        <div class="user-avatar">
		          <img alt="LostKobrakai" src="https://forum.elixirforum.com/user_avatar/forum.elixirforum.com/LostKobrakai/120/3072_2.png" width="120" height="120" />
		        </div>
					
						<div class="user-details">
		          <div class="user-name">
		            <h3>
                  LostKobrakai
                  </h3>
		          </div>
						
						</div>
					
					</div>

	        <div class="thread-main">
	            <div class="post-body" data-turbo="false">
								<aside class="quote no-group" data-username="eclark" data-post="15" data-topic="58988">
<div class="title">
<div class="quote-controls"></div>
<img alt="" width="24" height="24" src="https://forum.elixirforum.com/letter_avatar_proxy/v4/letter/e/8edcca/48.png" class="avatar"> eclark:</div>
<blockquote>
<p>Distribution adds latency. The first request in will compute assigns and the rest of the process state. Say the load balancer sent that to host <code>foo</code>. So <code>foo</code> starts that process. Now if the load balancer sends the web socket start request to <code>bar</code> with distribution <code>bar</code> will have to ask <code>foo</code> for the response, before passing it on. That adds latency and additional points of failure. <code>bar</code> can be overloaded, have a bad NIC or go down for maintence.</p>
</blockquote>
</aside>
<p>That’s simply not how LV works. Both the initial static request and the connecting websocket request compute assigns independently. There’s no state shared between them server side. Therefore it doesn’t matter if the websocket connection hits a different node than the initial static connect.</p>
<p>That’s also the reason why you want to use live navigation, because that one doesn’t use an additional http request, but navigates purely over an already established websocket connection.</p>
<p>What you describe does indeed matter though if a client falls back to long polling over websockets, because each individual poll would get routed by the load balancer. In that case the LV state is indeed transferred within the cluster depending on where the load balancer routed the users request to.</p>
<p>More generally I’d wonder if that really is worth optimizing for. Websockets work in many places and that percentage hopefully continues to rise. Also I’d hope your load balancer does not randomly send users to nodes in different regions. Because then it’s not just the LV adding latency, but also your load balancer sending traffic farther away from the user as well. And if the load balancer balances between nodes in the same region then those nodes should be able to talk to each other without much latency.</p> 
	            </div>

	            <div class="base-line">
	                <div class="thread-counters">
	                    <span class="thread-count count-likes js-likers-trigger" title="Likes" data-post-id="339561" data-batch-url="/posts/batch_likers">
                        2
                      </span>
                      <!-- <span class="thread-count js-solved-indicator" title="Marked as solution"></span> -->
	                </div>
	                <div class="go-to-post">
	                  <a title="Go to post" alt="Go to post" href="https://forum.elixirforum.com/t/production-phoenix-liveview-on-cloud-agnostic-kubernetes/58988/19">Post #18</a>
	                </div>
	            </div>
              <div id="likers-container-339561" 
                   class="likers-container"
                   data-first-post="false"
                   data-batch-url="/posts/batch_likers">
                   <div class="likers-placeholder" 
                     data-likers-post-id="339561"
                     data-batch-url="/posts/batch_likers">
                  <div class="post-likers"></div>
                </div>
              </div>
	        </div>
			

    </div>

    <div class="triangle-top-right type-standard-post cat-standard-post" title="Post #18"></div>
  </section>
</div>
    <div class="postbit" id="339569" data-post-id="339569">
  <section>
    <div class="post-wrap">


					<div class="post-header">
		        <div class="user-avatar">
		          <img alt="julismz" src="https://forum.elixirforum.com/user_avatar/forum.elixirforum.com/julismz/120/28368_2.png" width="120" height="120" />
		        </div>
					
						<div class="user-details">
		          <div class="user-name">
		            <h3>
                  julismz
                  </h3>
		          </div>
						
						</div>
					
					</div>

	        <div class="thread-main">
	            <div class="post-body" data-turbo="false">
								<p>Well <a class="mention" href="/u/marthinl" rel="nofollow">@MarthinL</a> first of all, thanks for answering and answering the way you do, with a lot of concept, experience and dedication. I took your message seriously and read it as a paper post <img src="https://forum.elixirforum.com/images/emoji/apple/stuck_out_tongue.png?v=15" title=":stuck_out_tongue:" class="emoji" alt=":stuck_out_tongue:" loading="lazy" width="20" height="20"></p>
<p>So far the issue is that baremetal servers are expensive just for the sake of future proofing, and having 2 K8S nodes on a single server is really, I don’t know how to put it, weird and nonsensical.</p>
<p>So I think I’ll run the app in standalone mode until we start having issues and spikes, and then I’ll add another baremetal, k8S, MetalB, etc.</p>
<p>As for nginx, I was using it as a load balancer… I thought it solved the sticky session issues with liveview, but maybe the K8S headless service takes care of that, I don’t know how, because it doesn’t know the IP of each pod, I should read a bit more about that.</p>
<p>In the current case, I won’t use it, and I trust that Bandit will resolve all requests like a champ. If at some point I want to put an assets subdomain that handles caching js, css, etc, I will have no choice but to forward the traffic to bandit from nginx since it will be using port 80 (possibly it will be handled with cloudflare for the moment).</p>
<p>So, will this message be discarded? At the very least, while the product is working I will do some tests based on everything you told me (except MetalLB that doesn’t work on VPS) in stage environments with some DO VPS to have everything ready and configured when the day of horizontal scaling comes.</p>
<p>And here is an additional question for all who knows and want to answer:</p>
<p>In the described context, It’s better to have only one elixir/erlang node with full resources which take advantage of all the CPU (Intel Xeon-E 2386G - 6c/12t), or it’s better to have at least 2 libcluster nodes?</p>
<p>Could I perhaps find a limitation on the endpoint with Bandit no matter how much processing power the application has?</p>
<p>It is a rather difficult adventure at times, especially for those of us who are not devops, but it has great satisfactions.</p>
<p>Thanks again.</p> 
	            </div>

	            <div class="base-line">
	                <div class="thread-counters">
	                    <span class="thread-count count-likes js-likers-trigger" title="Likes" data-post-id="339569" data-batch-url="/posts/batch_likers">
                        0
                      </span>
                      <!-- <span class="thread-count js-solved-indicator" title="Marked as solution"></span> -->
	                </div>
	                <div class="go-to-post">
	                  <a title="Go to post" alt="Go to post" href="https://forum.elixirforum.com/t/production-phoenix-liveview-on-cloud-agnostic-kubernetes/58988/20">Post #19</a>
	                </div>
	            </div>
              <div id="likers-container-339569" 
                   class="likers-container"
                   data-first-post="false"
                   data-batch-url="/posts/batch_likers">
                   <div class="likers-placeholder" 
                     data-likers-post-id="339569"
                     data-batch-url="/posts/batch_likers">
                  <div class="post-likers"></div>
                </div>
              </div>
	        </div>
			

    </div>

    <div class="triangle-top-right type-standard-post cat-standard-post" title="Post #19"></div>
  </section>
</div>
    <div class="postbit" id="339592" data-post-id="339592">
  <section>
    <div class="post-wrap">


					<div class="post-header">
		        <div class="user-avatar">
		          <img alt="eclark" src="/assets/icons/user-9f439610.png" width="120" height="120" />
		        </div>
					
						<div class="user-details">
		          <div class="user-name">
		            <h3>
                  eclark
                  </h3>
		          </div>
						
						</div>
					
					</div>

	        <div class="thread-main">
	            <div class="post-body" data-turbo="false">
								<aside class="quote no-group" data-username="LostKobrakai" data-post="19" data-topic="58988">
<div class="title">
<div class="quote-controls"></div>
<img alt="" width="24" height="24" src="https://forum.elixirforum.com/user_avatar/forum.elixirforum.com/lostkobrakai/48/3072_2.png" class="avatar"> LostKobrakai:</div>
<blockquote>
<p>Both the initial static request and the connecting websocket request compute assigns independently.</p>
</blockquote>
</aside>
<p>Hence why distribution adds latency. <code>foo</code> gets the first assigns computed. Assuming that your state has some cache locality and starts some running proccesses (in production it matters a lot more than the LV community assumes). So then <code>bar</code> gets the websocket request and has to decide to send messages for the state that’s running on <code>foo</code> or move/respawn those processes.</p>
<p>Reconnects are another example of adding latency if requests aren’t pinned to the same machine because of assign recomputation/process startup time.</p>
<p>If the entirety of your state for a session is small, none of this matters. However, that’s only the case on younger applications.</p> 
	            </div>

	            <div class="base-line">
	                <div class="thread-counters">
	                    <span class="thread-count count-likes js-likers-trigger" title="Likes" data-post-id="339592" data-batch-url="/posts/batch_likers">
                        0
                      </span>
                      <!-- <span class="thread-count js-solved-indicator" title="Marked as solution"></span> -->
	                </div>
	                <div class="go-to-post">
	                  <a title="Go to post" alt="Go to post" href="https://forum.elixirforum.com/t/production-phoenix-liveview-on-cloud-agnostic-kubernetes/58988/21">Post #20</a>
	                </div>
	            </div>
              <div id="likers-container-339592" 
                   class="likers-container"
                   data-first-post="false"
                   data-batch-url="/posts/batch_likers">
                   <div class="likers-placeholder" 
                     data-likers-post-id="339592"
                     data-batch-url="/posts/batch_likers">
                  <div class="post-likers"></div>
                </div>
              </div>
	        </div>
			

    </div>

    <div class="triangle-top-right type-standard-post cat-standard-post" title="Post #20"></div>
  </section>
</div>
</template></turbo-stream><turbo-stream action="replace" target="load-more-container"><template><div id="load-more-container" class="load-more-container">
    <a class="load-more-button" data-turbo-stream="true" href="/topics/58988/load_more?page=3">Load more posts (10 remaining)</a>
</div></template></turbo-stream>