<turbo-stream action="append" target="posts_list"><template>    <div class="postbit" id="378336" data-post-id="378336">
  <section>
    <div class="post-wrap">


					<div class="post-header">
		        <div class="user-avatar">
		          <img alt="itekhi" src="https://forum.elixirforum.com/user_avatar/forum.elixirforum.com/itekhi/120/38403_2.png" width="120" height="120" />
		        </div>
					
						<div class="user-details">
		          <div class="user-name">
		            <h3>
                  itekhi
                  </h3>
		          </div>
						
						</div>
					
					</div>

	        <div class="thread-main">
	            <div class="post-body" data-turbo="false">
								<aside class="quote no-group" data-username="garrison" data-post="59" data-topic="73003">
<div class="title">
<div class="quote-controls"></div>
<img alt="" width="24" height="24" src="https://forum.elixirforum.com/letter_avatar_proxy/v4/letter/g/3bc359/48.png" class="avatar"> garrison:</div>
<blockquote>
<p>it really is storage hardware and APIs specifically that are so unreliable</p>
</blockquote>
</aside>
<p>Oh wow… My jaw dropped when I read your post, I couldn’t believe things are that bad in storage hardware world, but apparently, it’s reality… And now tables have turned that I’m taking Sorc96 place wanting to ask you his questions <img src="https://forum.elixirforum.com/images/emoji/apple/sweat_smile.png?v=15" title=":sweat_smile:" class="emoji" alt=":sweat_smile:" loading="lazy" width="20" height="20"></p>
<aside class="quote no-group" data-username="garrison" data-post="57" data-topic="73003">
<div class="title">
<div class="quote-controls"></div>
<img alt="" width="24" height="24" src="https://forum.elixirforum.com/letter_avatar_proxy/v4/letter/g/3bc359/48.png" class="avatar"> garrison:</div>
<blockquote>
<p>Hopefully the value proposition of this work is starting to become clear lol.</p>
<p>This is what I mean when I say that I want to solve these problems once and never again.</p>
</blockquote>
</aside>
<p>I take my hat off to you. At first I didn’t understand, but now I’m eagerly looking forward to Hobbes <img src="https://forum.elixirforum.com/images/emoji/apple/face_exhaling.png?v=15" title=":face_exhaling:" class="emoji" alt=":face_exhaling:" loading="lazy" width="20" height="20"></p> 
	            </div>

	            <div class="base-line">
	                <div class="thread-counters">
	                    <span class="thread-count count-likes js-likers-trigger" title="Likes" data-post-id="378336" data-batch-url="/posts/batch_likers">
                        0
                      </span>
                      <!-- <span class="thread-count js-solved-indicator" title="Marked as solution"></span> -->
	                </div>
	                <div class="go-to-post">
	                  <a title="Go to post" alt="Go to post" href="https://forum.elixirforum.com/t/hobbes-a-low-level-distributed-database-for-the-elixir-programming-language/73003/62">Post #61</a>
	                </div>
	            </div>
              <div id="likers-container-378336" 
                   class="likers-container"
                   data-first-post="false"
                   data-batch-url="/posts/batch_likers">
                   <div class="likers-placeholder" 
                     data-likers-post-id="378336"
                     data-batch-url="/posts/batch_likers">
                  <div class="post-likers"></div>
                </div>
              </div>
	        </div>
			

    </div>

    <div class="triangle-top-right type-standard-post cat-standard-post" title="Post #61"></div>
  </section>
</div>
    <div class="postbit" id="378340" data-post-id="378340">
  <section>
    <div class="post-wrap">


					<div class="post-header">
		        <div class="user-avatar">
		          <img alt="dimitarvp" src="https://forum.elixirforum.com/user_avatar/forum.elixirforum.com/dimitarvp/120/38664_2.png" width="120" height="120" />
		        </div>
					
						<div class="user-details">
		          <div class="user-name">
		            <h3>
                  dimitarvp
                  </h3>
		          </div>
						
						</div>
					
					</div>

	        <div class="thread-main">
	            <div class="post-body" data-turbo="false">
								<p>Your knowledge on the topic is commendable. But I do wonder how can you solve this in a comprehensive manner, once and for all? Barring writing a driver that uses disk in their raw form (block devices?) then I am not sure you can protect against the kernel and the existing drivers trying to be too smart and too helpful, and crippling proper database performance as a result – which as you alluded to is an old and known problem.</p> 
	            </div>

	            <div class="base-line">
	                <div class="thread-counters">
	                    <span class="thread-count count-likes js-likers-trigger" title="Likes" data-post-id="378340" data-batch-url="/posts/batch_likers">
                        3
                      </span>
                      <!-- <span class="thread-count js-solved-indicator" title="Marked as solution"></span> -->
	                </div>
	                <div class="go-to-post">
	                  <a title="Go to post" alt="Go to post" href="https://forum.elixirforum.com/t/hobbes-a-low-level-distributed-database-for-the-elixir-programming-language/73003/63">Post #62</a>
	                </div>
	            </div>
              <div id="likers-container-378340" 
                   class="likers-container"
                   data-first-post="false"
                   data-batch-url="/posts/batch_likers">
                   <div class="likers-placeholder" 
                     data-likers-post-id="378340"
                     data-batch-url="/posts/batch_likers">
                  <div class="post-likers"></div>
                </div>
              </div>
	        </div>
			

    </div>

    <div class="triangle-top-right type-standard-post cat-standard-post" title="Post #62"></div>
  </section>
</div>
    <div class="postbit" id="378379" data-post-id="378379">
  <section>
    <div class="post-wrap">


					<div class="post-header">
		        <div class="user-avatar">
		          <img alt="garrison" src="/assets/icons/user-9f439610.png" width="120" height="120" />
		        </div>
					
						<div class="user-details">
		          <div class="user-name">
		            <h3>
                  garrison
                    <span class="op-star" title="Thread Starter">
                      <img alt="OP" class="op-star-icon" src="/assets/thread-icons/thread-icon-thread-starter-df91e872.png" />
                    </span>
                  </h3>
		          </div>
						
						</div>
					
					</div>

	        <div class="thread-main">
	            <div class="post-body" data-turbo="false">
								<p>Lol first I think I need to clarify something. When I say I want to solve these problems once I mean <em>for me</em>. I don’t mean that I am, like, the savior of the database universe or something; that would be quite silly.</p>
<p>Actually, it’s the Tigerbeetle devs who have been pushing this boundary forward lately with respect to having a proper storage fault model, and I am very much “copying off of their homework” in my own implementation. But that’s how this stuff works: much of their approach was copied from ZFS (the filesystem), the designers of which also contributed enormously to this idea of systems that don’t incessantly lose data.</p>
<p>What I mean by “solve these problems once” is actually, for me, very specific. I am very interested in sovereignty as I have mentioned, and I want to be able to build systems that store data for some products that I’m working on. Here are a few of the things I want:</p>
<ul>
<li>A scalable relational database with strong correctness guarantees and native support for multitenancy</li>
<li>A distributed blob store/filesystem that integrates with the above and maintains the same guarantees</li>
<li>A full text search engine that integrates with the above and maintains the same guarantees</li>
</ul>
<p>This list is not random, it’s a specific list of things that I actually need for specific functionality in apps that I am actually designing. I want to be able to “own” these things so I am not forever at the mercy of megacorporations that hate us. There are many open source projects available that do some of these things, but none of them are exactly what I want.</p>
<p>The thing is, all of these things (and others) have the same base requirements: a way to consistently replicate persistent data across servers and be somewhat resilient to corruption. So I could write that code three or four times, or I could <em>factor it out</em>, write it once, and then build the other stuff as (mostly) stateless layers on top of it.</p>
<p>Hobbes is <em>that code factored out</em>. Hobbes is a toolkit which will contain all of the parts of building “distributed database-shaped things” which are the <em>same</em> so that I don’t have to re-implement them 10 times. This is actually a <em>practical endeavor</em>. I have clear and specific goals.</p>
<p>However, in sharing this work with the community of course I am trying to also present it in terms of how it might be useful <em>to them</em>, as that’s what people really want to hear (see the first few posts in this thread).</p> 
	            </div>

	            <div class="base-line">
	                <div class="thread-counters">
	                    <span class="thread-count count-likes js-likers-trigger" title="Likes" data-post-id="378379" data-batch-url="/posts/batch_likers">
                        6
                      </span>
                      <!-- <span class="thread-count js-solved-indicator" title="Marked as solution"></span> -->
	                </div>
	                <div class="go-to-post">
	                  <a title="Go to post" alt="Go to post" href="https://forum.elixirforum.com/t/hobbes-a-low-level-distributed-database-for-the-elixir-programming-language/73003/64">Post #63</a>
	                </div>
	            </div>
              <div id="likers-container-378379" 
                   class="likers-container"
                   data-first-post="false"
                   data-batch-url="/posts/batch_likers">
                   <div class="likers-placeholder" 
                     data-likers-post-id="378379"
                     data-batch-url="/posts/batch_likers">
                  <div class="post-likers"></div>
                </div>
              </div>
	        </div>
			

    </div>

    <div class="triangle-top-right type-standard-post cat-standard-post" title="Post #63"></div>
  </section>
</div>
    <div class="postbit" id="378380" data-post-id="378380">
  <section>
    <div class="post-wrap">


					<div class="post-header">
		        <div class="user-avatar">
		          <img alt="garrison" src="/assets/icons/user-9f439610.png" width="120" height="120" />
		        </div>
					
						<div class="user-details">
		          <div class="user-name">
		            <h3>
                  garrison
                    <span class="op-star" title="Thread Starter">
                      <img alt="OP" class="op-star-icon" src="/assets/thread-icons/thread-icon-thread-starter-df91e872.png" />
                    </span>
                  </h3>
		          </div>
						
						</div>
					
					</div>

	        <div class="thread-main">
	            <div class="post-body" data-turbo="false">
								<aside class="quote no-group" data-username="dimitarvp" data-post="63" data-topic="73003">
<div class="title">
<div class="quote-controls"></div>
<img alt="" width="24" height="24" src="https://forum.elixirforum.com/user_avatar/forum.elixirforum.com/dimitarvp/48/38664_2.png" class="avatar"> dimitarvp:</div>
<blockquote>
<p>Barring writing a driver that uses disk in their raw form (block devices?) then I am not sure you can protect against the kernel and the existing drivers trying to be too smart and too helpful, and crippling proper database performance as a result</p>
</blockquote>
</aside>
<p>Well on Linux you kinda can. You just set <code>O_DIRECT</code> and bypass the page cache. That’s literally what it’s for. This is going to become a lot more common now that it is known to be effectively impossible to avoid corruption without that flag. Postgres prominently supports direct I/O now for example.</p>
<p>(Actually it’s funny, this all began with Postgres’s fsyncgate. After that they modified the database to crash on fsync failure, only for that paper to then be published showing <em>that doesn’t work either</em>!)</p>
<p>I am not knowledgeable or resourced enough to do all of the research needed for this on my own. But thankfully Tigerbeetle has pretty solid internals docs and readable code and they care a lot about evangelizing this approach (because it makes their database look good) so I can just learn from them.</p>
<p>If they say ECC memory and <code>O_DIRECT</code> is good enough then I believe them. Obviously no database is going to survive the Sun swallowing the Earth, but if I’m writing a storage engine from scratch (which I literally am right now, see the <code>xks</code> directory) then I may as well use the latest best practices.</p>
<p>BTW I will probably not do the nif thing for a while. I’m just trying to structure everything so that it will be an easy switch. I am not looking forward to dealing with native code because I am not a systems programmer. But I was not a database programmer either so <em>shrug</em>. I’ll make do.</p> 
	            </div>

	            <div class="base-line">
	                <div class="thread-counters">
	                    <span class="thread-count count-likes js-likers-trigger" title="Likes" data-post-id="378380" data-batch-url="/posts/batch_likers">
                        5
                      </span>
                      <!-- <span class="thread-count js-solved-indicator" title="Marked as solution"></span> -->
	                </div>
	                <div class="go-to-post">
	                  <a title="Go to post" alt="Go to post" href="https://forum.elixirforum.com/t/hobbes-a-low-level-distributed-database-for-the-elixir-programming-language/73003/65">Post #64</a>
	                </div>
	            </div>
              <div id="likers-container-378380" 
                   class="likers-container"
                   data-first-post="false"
                   data-batch-url="/posts/batch_likers">
                   <div class="likers-placeholder" 
                     data-likers-post-id="378380"
                     data-batch-url="/posts/batch_likers">
                  <div class="post-likers"></div>
                </div>
              </div>
	        </div>
			

    </div>

    <div class="triangle-top-right type-standard-post cat-standard-post" title="Post #64"></div>
  </section>
</div>
    <div class="postbit" id="378386" data-post-id="378386">
  <section>
    <div class="post-wrap">


					<div class="post-header">
		        <div class="user-avatar">
		          <img alt="garrison" src="/assets/icons/user-9f439610.png" width="120" height="120" />
		        </div>
					
						<div class="user-details">
		          <div class="user-name">
		            <h3>
                  garrison
                    <span class="op-star" title="Thread Starter">
                      <img alt="OP" class="op-star-icon" src="/assets/thread-icons/thread-icon-thread-starter-df91e872.png" />
                    </span>
                  </h3>
		          </div>
						
						</div>
					
					</div>

	        <div class="thread-main">
	            <div class="post-body" data-turbo="false">
								<aside class="quote no-group" data-username="itekhi" data-post="61" data-topic="73003">
<div class="title">
<div class="quote-controls"></div>
<img alt="" width="24" height="24" src="https://forum.elixirforum.com/user_avatar/forum.elixirforum.com/itekhi/48/38403_2.png" class="avatar"> itekhi:</div>
<blockquote>
<p>How large can the dataset be? I don’t get what is the limit to the amount of data that can be stored until query/write performance gets bad.</p>
</blockquote>
</aside>
<p>I won’t know the answer to this for a while as Hobbes is not yet at a place where this can be tested. And even if it were, doing such tests would be pretty (financially) expensive. Scaling to huge numbers is not a priority for me in the short term. What matters to me is that Hobbes is architecturally <em>capable</em> of scaling so that it can grow with me.</p>
<p>With that said, FoundationDB could provide a rough upper bound. FDB is known to tap out around a few hundred terabytes of data (i.e. around a petabyte after replication). This is on the order of maybe a thousand servers. My understanding is that FDB taps out <em>due to keepalives</em>, though, which is a very fixable thing. So I suspect its users (i.e. Apple/Snowflake) simply have not found much use for getting past ~100TB scale in a single cluster.</p>
<p>Apple’s deployment for example is highly multitenant: it stores all CloudKit data, and is probably slowly consuming their (also enormous) Cassandra deployment. You want to have separate clusters for blast radius reasons, and you can move tenants between clusters, so really the cluster size limit is a tenant size limit. And 100TB is a <em>very big tenant</em>, so you can see why there is little need to get larger than that. Also, you want separate clusters so that you can shard geographically for low latency.</p>
<aside class="quote no-group" data-username="itekhi" data-post="61" data-topic="73003">
<div class="title">
<div class="quote-controls"></div>
<img alt="" width="24" height="24" src="https://forum.elixirforum.com/user_avatar/forum.elixirforum.com/itekhi/48/38403_2.png" class="avatar"> itekhi:</div>
<blockquote>
<p>If it’s not much to ask, I understand that there’s also network throughput consideration, that’s why batching exists. But what should you do when massive amount of requests to write small pieces of data come from a lot of sources?</p>
</blockquote>
</aside>
<p>If you have a very big dataset you have to shard it across servers. A primitive approach here is to hash shard data into explicit buckets like e.g. DynamoDB which has 10GB shards controlled by a user-defined (developer-defined, that is) shard key. In that case you just send data to the right server(s) for its shard.</p>
<p>Unfortunately this means that you have to know the distribution of your dataset in advance, which is problematic. To solve this you can do <em>range</em>-based autosharding, where the shards are split/merged automatically and moved around as the distribution changes. This is a more advanced approach.</p>
<p>Spanner, CockroachDB, FoundationDB, and of course Hobbes are examples of the latter approach.</p>
<p>Maintaining atomic transactions across disparate shards gets tricky, though. Older systems just <em>don’t</em>, while newer systems like CockroachDB have atomic transactions but with questionable consistency guarantees. Spanner, FDB, and Hobbes offer very strong consistency guarantees.</p>
<p>The answer to your actual question is deeply technical and completely different among all of the systems I mentioned, but I can try to hand-wave a bit. For Hobbes we essentially take in a bunch of small transactions, group them together into a batch with a single version, perform concurrency control, and then <em>split them up</em> and send them to different servers for storage. This is, like, legit pretty complicated though.</p> 
	            </div>

	            <div class="base-line">
	                <div class="thread-counters">
	                    <span class="thread-count count-likes js-likers-trigger" title="Likes" data-post-id="378386" data-batch-url="/posts/batch_likers">
                        3
                      </span>
                      <!-- <span class="thread-count js-solved-indicator" title="Marked as solution"></span> -->
	                </div>
	                <div class="go-to-post">
	                  <a title="Go to post" alt="Go to post" href="https://forum.elixirforum.com/t/hobbes-a-low-level-distributed-database-for-the-elixir-programming-language/73003/66">Post #65</a>
	                </div>
	            </div>
              <div id="likers-container-378386" 
                   class="likers-container"
                   data-first-post="false"
                   data-batch-url="/posts/batch_likers">
                   <div class="likers-placeholder" 
                     data-likers-post-id="378386"
                     data-batch-url="/posts/batch_likers">
                  <div class="post-likers"></div>
                </div>
              </div>
	        </div>
			

    </div>

    <div class="triangle-top-right type-standard-post cat-standard-post" title="Post #65"></div>
  </section>
</div>
    <div class="postbit" id="378794" data-post-id="378794">
  <section>
    <div class="post-wrap">


					<div class="post-header">
		        <div class="user-avatar">
		          <img alt="jstimps" src="https://forum.elixirforum.com/user_avatar/forum.elixirforum.com/jstimps/120/37213_2.png" width="120" height="120" />
		        </div>
					
						<div class="user-details">
		          <div class="user-name">
		            <h3>
                  jstimps
                  </h3>
		          </div>
						
						</div>
					
					</div>

	        <div class="thread-main">
	            <div class="post-body" data-turbo="false">
								<p>Hi <a class="mention" href="/u/garrison" rel="nofollow">@garrison</a> , I was reading some of Hobbes (specifically the binary enc/dec in Keyset) and realized you may be interested in this erlfdb PR.</p>
<p><a href="https://github.com/foundationdb-beam/erlfdb/pull/57" class="onebox" target="_blank" rel="noopener nofollow ugc">https://github.com/foundationdb-beam/erlfdb/pull/57</a></p>
<p>The implementation from the original erlfdb accumulated an offset int and used that to do ever-increasing match clauses. This turns out to be pretty slow. It makes a measurable difference when doing many erlfdb_tuple operations.</p> 
	            </div>

	            <div class="base-line">
	                <div class="thread-counters">
	                    <span class="thread-count count-likes js-likers-trigger" title="Likes" data-post-id="378794" data-batch-url="/posts/batch_likers">
                        2
                      </span>
                      <!-- <span class="thread-count js-solved-indicator" title="Marked as solution"></span> -->
	                </div>
	                <div class="go-to-post">
	                  <a title="Go to post" alt="Go to post" href="https://forum.elixirforum.com/t/hobbes-a-low-level-distributed-database-for-the-elixir-programming-language/73003/67">Post #66</a>
	                </div>
	            </div>
              <div id="likers-container-378794" 
                   class="likers-container"
                   data-first-post="false"
                   data-batch-url="/posts/batch_likers">
                   <div class="likers-placeholder" 
                     data-likers-post-id="378794"
                     data-batch-url="/posts/batch_likers">
                  <div class="post-likers"></div>
                </div>
              </div>
	        </div>
			

    </div>

    <div class="triangle-top-right type-standard-post cat-standard-post" title="Post #66"></div>
  </section>
</div>
    <div class="postbit" id="378801" data-post-id="378801">
  <section>
    <div class="post-wrap">


					<div class="post-header">
		        <div class="user-avatar">
		          <img alt="garrison" src="/assets/icons/user-9f439610.png" width="120" height="120" />
		        </div>
					
						<div class="user-details">
		          <div class="user-name">
		            <h3>
                  garrison
                    <span class="op-star" title="Thread Starter">
                      <img alt="OP" class="op-star-icon" src="/assets/thread-icons/thread-icon-thread-starter-df91e872.png" />
                    </span>
                  </h3>
		          </div>
						
						</div>
					
					</div>

	        <div class="thread-main">
	            <div class="post-body" data-turbo="false">
								<p>Hah, I thought the offset thing was pretty clever in that you can grab the binary slice without accumulating by character (which would be the naive approach). But it looks like <code>:binary.match()</code> is a nif and there’s no beating that. Thanks for the tip, I’ll keep that in mind.</p>
<p>If you’re reading that code I should note that I only implemented enough of Keyset to be useful internally. The first implementations of a lot of the internals were scaffolded out with hacky manual string encodings and that quickly became an actual problem so I needed some sane form of binary encoding.</p>
<p>I was going to make a point about how unfortunate it is that we can’t prefix the strings with their lengths but I see you already mentioned that in the PR! This is a case where I think the <em>value</em> encodings can make a much better tradeoff by including a header with types/offsets/lengths at the start of each value. That way you don’t have to decode the entire value (much larger than a key) in order to pull out a “column”. Obviously the values have no ordering requirement.</p> 
	            </div>

	            <div class="base-line">
	                <div class="thread-counters">
	                    <span class="thread-count count-likes js-likers-trigger" title="Likes" data-post-id="378801" data-batch-url="/posts/batch_likers">
                        2
                      </span>
                      <!-- <span class="thread-count js-solved-indicator" title="Marked as solution"></span> -->
	                </div>
	                <div class="go-to-post">
	                  <a title="Go to post" alt="Go to post" href="https://forum.elixirforum.com/t/hobbes-a-low-level-distributed-database-for-the-elixir-programming-language/73003/68">Post #67</a>
	                </div>
	            </div>
              <div id="likers-container-378801" 
                   class="likers-container"
                   data-first-post="false"
                   data-batch-url="/posts/batch_likers">
                   <div class="likers-placeholder" 
                     data-likers-post-id="378801"
                     data-batch-url="/posts/batch_likers">
                  <div class="post-likers"></div>
                </div>
              </div>
	        </div>
			

    </div>

    <div class="triangle-top-right type-standard-post cat-standard-post" title="Post #67"></div>
  </section>
</div>
    <div class="postbit" id="379866" data-post-id="379866">
  <section>
    <div class="post-wrap">


					<div class="post-header">
		        <div class="user-avatar">
		          <img alt="garrison" src="/assets/icons/user-9f439610.png" width="120" height="120" />
		        </div>
					
						<div class="user-details">
		          <div class="user-name">
		            <h3>
                  garrison
                    <span class="op-star" title="Thread Starter">
                      <img alt="OP" class="op-star-icon" src="/assets/thread-icons/thread-icon-thread-starter-df91e872.png" />
                    </span>
                  </h3>
		          </div>
						
						</div>
					
					</div>

	        <div class="thread-main">
	            <div class="post-body" data-turbo="false">
								<p>I think it’s time for a little update.</p>
<p>Storage engine development has been going very well. There is, mostly, a storage engine now. It’s named XKS (short for ExKeyStore, which I think is a hilarious pun) and you can find the code in the <code>lib/xks</code> directory. It’s unfinished and obviously still messy, but most of the constituent parts are now there and it’s just a matter of gluing them together.</p>
<p>For those unaware, a “storage engine” is just the part of a database that writes data to disk. Generally a storage engine exposes some level of abstraction (like “key/value store”) and atomic commit functionality. Strangely, storage engines are often completely separate pieces of software from the databases that utilize them. MySQL has InnoDB, MongoDB has WiredTiger, and so on. Some databases do have their own for various reasons: Postgres’s is quite bespoke and quite old (and quite bad), Tigerbeetle’s is <em>very</em> deeply integrated with their data model (it literally only stores fixed-length data, which is unique).</p>
<p>For those of you who remember some of my earlier comments on the topic (before Hobbes was formally announced), you may remember that I originally intended to just use SQLite as a storage engine and dodge this particular yak for now. In fact, that’s exactly what FoundationDB did for over a decade. They have since switched over to RocksDB, for some reason.</p>
<p>So why did I change my mind? Well, firstly at the time I just didn’t know <em>how</em> to write a proper storage engine, but as Hobbes’s development dragged on I ended up studying enough to fix that. But also, I came to realize that there are some areas where deep integration with the storage engine can actually simplify the implementation of Hobbes’s distributed features <em>and</em> improve correctness. And if there are any two things I like, they’re being lazy and not writing bugs. So, uh, yolo.</p>
<p>On the correctness side, rolling the entire storage engine from scratch means Hobbes can maintain <em>zero</em> dependencies (and if there’s anything I hate, it’s dependencies) and perform fully integrated simulation testing of the entire database as a whole. Whereas if I used something like SQLite, I would probably be mocking it out during testing, which is Not Great. Also, XKS is designed to be <em>very</em> resilient to corruption via comically aggressive cryptographic checksumming of the entire file a la ZFS or Tigerbeetle; it’s a modern design.</p>
<p>On the functionality side, FDB’s most unfortunate limitation is that read-only transactions can only be a few seconds long. Most of FDB’s limitations are Good Actually, including the limit on read+write transactions. Long transactions in an optimistic system (or really even a pessimistic system) are poison and to be avoided. But <em>read-only</em> transactions are different as they have no contention under an MVCC model. The reason FDB doesn’t support this is really a skill issue on the part of the SQLite btree: it doesn’t know how to store versioned data. Poor thing.</p>
<p>XKS uses an LSM tree design that pushes database versions directly down into the storage engine. LSMs are a natural fit for this because compaction provides an opportunity to garbage-collect old versions. And if you design this <em>wrong</em> you end up with Postgres’s VACUUM disaster, where, fun fact, they originally designed it so that old tuples would never be garbage collected <em>at all</em> and then apparently found out that that is a bad idea (lol).</p>
<p>The XKS design is quite unique and I have been unable to find another real-world example of an LSM that works this way. I’m sure the idea must exist in research somewhere, but in practice it seems like people just add another version to the keys “in userspace” (see CockroachDB/Pebble). Maybe Spanner’s storage engine does this but we’ll never know because they only gave us a one-paragraph description (WHY).</p>
<p>That’s all for now.</p> 
	            </div>

	            <div class="base-line">
	                <div class="thread-counters">
	                    <span class="thread-count count-likes js-likers-trigger" title="Likes" data-post-id="379866" data-batch-url="/posts/batch_likers">
                        12
                      </span>
                      <!-- <span class="thread-count js-solved-indicator" title="Marked as solution"></span> -->
	                </div>
	                <div class="go-to-post">
	                  <a title="Go to post" alt="Go to post" href="https://forum.elixirforum.com/t/hobbes-a-low-level-distributed-database-for-the-elixir-programming-language/73003/69">Post #68</a>
	                </div>
	            </div>
              <div id="likers-container-379866" 
                   class="likers-container"
                   data-first-post="false"
                   data-batch-url="/posts/batch_likers">
                   <div class="likers-placeholder" 
                     data-likers-post-id="379866"
                     data-batch-url="/posts/batch_likers">
                  <div class="post-likers"></div>
                </div>
              </div>
	        </div>
			

    </div>

    <div class="triangle-top-right type-most-liked cat-most-liked" title="One of the top 3 liked posts in this thread!"></div>
  </section>
</div>
    <div class="postbit" id="379902" data-post-id="379902">
  <section>
    <div class="post-wrap">


					<div class="post-header">
		        <div class="user-avatar">
		          <img alt="garrison" src="/assets/icons/user-9f439610.png" width="120" height="120" />
		        </div>
					
						<div class="user-details">
		          <div class="user-name">
		            <h3>
                  garrison
                    <span class="op-star" title="Thread Starter">
                      <img alt="OP" class="op-star-icon" src="/assets/thread-icons/thread-icon-thread-starter-df91e872.png" />
                    </span>
                  </h3>
		          </div>
						
						</div>
					
					</div>

	        <div class="thread-main">
	            <div class="post-body" data-turbo="false">
								<p>Very thankful for my fuzzers when I have to refactor <a href="https://git.sr.ht/~garrisonc/hobbes/commit/b09aff66ae21567fc010d91331de8407917441ee" rel="noopener nofollow ugc">garbage like this</a>. And no, I do <em>not</em> want to talk about wtf I was thinking when I wrote the first pass. I think I was just trying to get through it.</p>
<p>It’s becoming increasingly clear to me how much code quality matters, and also how much of it is a pure expression of skill and experience. I did not understand how to shape that code properly until I had to write a similar version a few more times elsewhere in the engine, and then it slowly dawned on me. That first attempt was so messy I literally had to sit down and reverse engineer it despite having written it only a few weeks ago.</p>
<p>It’s also becoming increasingly clear to me that very little of the difference can be automatically linted or checked. I have suspected this for a while (and I don’t like linters or formatters very much), but this is a pretty clear example. I doubt anyone who clicks that link will even be able to understand why the new version is better; not out of ignorance, but because it’s <em>enormously context-specific</em>. I doubt a top LLM could make much sense of it either, at least not well enough to make useful recommendations <em>in advance</em>.</p>
<p>I wrote a lot of messy code over the past week trying to get persistence of the remaining data structures to work. Now that the engine can save itself to “disk” (or so it thinks), it’s definitely time to go over everything and clean up before I move on to the fun parts. Though I do find refactoring to be quite relaxing as long as I have good tests, which is of course the hard bit.</p>
<p>I’ve started to get very good at scaffolding out large projects. I suspect junior devs think this ability is black magic, but it’s <a href="https://mitchellh.com/writing/building-large-technical-projects" rel="noopener nofollow ugc">very much a skill</a> that you hone with practice. Eventually you learn how to break the thing up and build it one piece at a time.</p>
<p>I’ve seen people suggest that they find LLMs useful for starting new projects, which I find concerning. If you don’t practice you will never learn how to do it yourself. Maybe fine if models are writing all of the code, but what do you do when you want to build something outside of the training set?</p>
<p>I think I should consider posting more devlog-shaped things. Maybe in a separate thread, or perhaps <a href="https://corporate.fm/blog" rel="noopener nofollow ugc">on the blog</a>. Unclear.</p> 
	            </div>

	            <div class="base-line">
	                <div class="thread-counters">
	                    <span class="thread-count count-likes js-likers-trigger" title="Likes" data-post-id="379902" data-batch-url="/posts/batch_likers">
                        5
                      </span>
                      <!-- <span class="thread-count js-solved-indicator" title="Marked as solution"></span> -->
	                </div>
	                <div class="go-to-post">
	                  <a title="Go to post" alt="Go to post" href="https://forum.elixirforum.com/t/hobbes-a-low-level-distributed-database-for-the-elixir-programming-language/73003/70">Post #69</a>
	                </div>
	            </div>
              <div id="likers-container-379902" 
                   class="likers-container"
                   data-first-post="false"
                   data-batch-url="/posts/batch_likers">
                   <div class="likers-placeholder" 
                     data-likers-post-id="379902"
                     data-batch-url="/posts/batch_likers">
                  <div class="post-likers"></div>
                </div>
              </div>
	        </div>
			

    </div>

    <div class="triangle-top-right type-standard-post cat-standard-post" title="Post #69"></div>
  </section>
</div>
    <div class="postbit" id="379910" data-post-id="379910">
  <section>
    <div class="post-wrap">


					<div class="post-header">
		        <div class="user-avatar">
		          <img alt="dimitarvp" src="https://forum.elixirforum.com/user_avatar/forum.elixirforum.com/dimitarvp/120/38664_2.png" width="120" height="120" />
		        </div>
					
						<div class="user-details">
		          <div class="user-name">
		            <h3>
                  dimitarvp
                  </h3>
		          </div>
						
						</div>
					
					</div>

	        <div class="thread-main">
	            <div class="post-body" data-turbo="false">
								<aside class="quote no-group" data-username="garrison" data-post="70" data-topic="73003">
<div class="title">
<div class="quote-controls"></div>
<img alt="" width="24" height="24" src="https://forum.elixirforum.com/letter_avatar_proxy/v4/letter/g/3bc359/48.png" class="avatar"> garrison:</div>
<blockquote>
<p>I’ve seen people suggest that they find LLMs useful for starting new projects, which I find concerning. If you don’t practice you will never learn how to do it yourself.</p>
</blockquote>
</aside>
<p>LLMs are a great way to overcome the blank canvas terror. Or just make a draft so you can curate it. I don’t view remembering the exact layout of a typical Elixir project as an important skill.</p>
<p>In commercial programming generating and then massaging the generated code is an oft-enough occurrence for LLMs to become objective net positives.</p>
<aside class="quote no-group" data-username="garrison" data-post="70" data-topic="73003">
<div class="title">
<div class="quote-controls"></div>
<img alt="" width="24" height="24" src="https://forum.elixirforum.com/letter_avatar_proxy/v4/letter/g/3bc359/48.png" class="avatar"> garrison:</div>
<blockquote>
<p>Very thankful for my fuzzers when I have to refactor <a href="https://git.sr.ht/~garrisonc/hobbes/commit/b09aff66ae21567fc010d91331de8407917441ee" rel="noopener nofollow ugc">garbage like this</a>.</p>
</blockquote>
</aside>
<p>Took a quick look. It mostly seems aimed at making sure your recursive accumulator-using (and potentially tailcall-optimized) functions more… idiomatic? Easier to work with? Not quite sure but yeah, it’s a touch obscure code indeed. But deconstructing data and making the recursive functions easier to read is already a big win IMO.</p>
<p>How did fuzzers help? Did they expose edge cases?</p> 
	            </div>

	            <div class="base-line">
	                <div class="thread-counters">
	                    <span class="thread-count count-likes js-likers-trigger" title="Likes" data-post-id="379910" data-batch-url="/posts/batch_likers">
                        2
                      </span>
                      <!-- <span class="thread-count js-solved-indicator" title="Marked as solution"></span> -->
	                </div>
	                <div class="go-to-post">
	                  <a title="Go to post" alt="Go to post" href="https://forum.elixirforum.com/t/hobbes-a-low-level-distributed-database-for-the-elixir-programming-language/73003/71">Post #70</a>
	                </div>
	            </div>
              <div id="likers-container-379910" 
                   class="likers-container"
                   data-first-post="false"
                   data-batch-url="/posts/batch_likers">
                   <div class="likers-placeholder" 
                     data-likers-post-id="379910"
                     data-batch-url="/posts/batch_likers">
                  <div class="post-likers"></div>
                </div>
              </div>
	        </div>
			

    </div>

    <div class="triangle-top-right type-standard-post cat-standard-post" title="Post #70"></div>
  </section>
</div>
</template></turbo-stream><turbo-stream action="replace" target="load-more-container"><template><div id="load-more-container" class="load-more-container">
    <a class="load-more-button" data-turbo-stream="true" href="/topics/73003/load_more?page=8">Load more posts (16 remaining)</a>
</div></template></turbo-stream>