caleb-bb
So I’ve got this RAG project called Cake that I’ve been working on a good long while. For this topic, the RAG part is less important than the development process.
I want to use this topic as a showcase but also for discussion. I think I’ve wound up reinventing some wheels here as part of the learning process (which is fine)
Some backstory: I’ve been working on this RAG project since early '25. The first year or so of development was mostly by hand, including the initial iterations of the OTP behaviours. I got on the LLM bandwagon a little later and, at first, it didn’t increase my productivity by much.
What did increase my productivity – dramatically – was writing up an agent contract.
First of all, you’ll notice that there are a lot of merge and push gates. This is intentional. My general principle here is that engineers are sensitive to tedium, but coding agents are not. Moreover, guardrails tend to reduce tedium at the review stage and increase it at the coding stage; therefore, we should go ham with guardrails, because the only time they produce tedium is when agents are doing all the work anyway. So, let’s start with some…
Quality Gates
Credo, meaning mix credo --strict with some extra checks I cribbed from Mike Zornek’s excellent blog. I mention this one first because it enforces better code quality. This makes reviewing less of a chore, which makes the rest of the process go faster. This is both a push and a merge gate.
mix compile --warnings-as-errors, also cribbed from Mike. Again, adding --warnings-as-errors and enforcing this on push would be unbearably operose for some engineers. While working on Cake this evening, my Claude agent failed this check because a private helper landed between two clauses of another function. Per the agent contract, it simply moved the helper. That’s one less thing for me to facepalm over during review. You’ll notice it’s also the sort of groaner that you might commit by e.g. accidentally pasting something into the module after a long day. Not the kind of thing I want to spend time on, so a guardrail takes care of it.
Boundary. because Elixir has no native notion of dependencies besides “def public, defp not public”. Using agentic coding to create a large codebase that is totally hierarchy-naive is a great way to get circular dependencies and other general tomfoolery. The natural solution to that is Boundary. The compiler is totally agnostic about module namespacing but Boundary clues the compiler in. mix compile --warnings-as-errors fails whenever a module tries to do something against the Boundary settings.
mix format --check-formatted because I have better things to do than format code, check if code is formatted, or review unformatted slop. I’m not terribly enthusiastic about monitoring for unused entries in mix.lock either so we also have mix deps.unlock --check-unused.
When I get time I’ll get around to discussing the test suite. For now, you can feel free to poke around the code base and see how the agent contract shapes things.
If you’re a Claude Code user and you want to get a feel for what this does, try pulling down the repo and making modifications with Claude Code. Notice how the agent behaves.
Trending in AI / LLMs
Other Trending Topics
Categories:
Sub Categories:
Forums
Popular Tags
- #ecto
- #liveview
- #troubleshooting
- #learning-elixir
- #library
- #deployment
- #erlang
- #testing
- #genserver
- #mix
- #absinthe
- #remote-other
- #otp
- #plug
- #how-to-question
- #macros
- #postgres
- #elixirconf
- #channels
- #exunit
- #discussion
- #code-sync
- #podcasts
- #javascript
- #onsite
- #dialyzer
- #docker
- #authentication
- #umbrella
- #full-time-contract
- #ai
- #podcasts-by-brainlid
- #ecto-query
- #blog-post
- #elixirconf-us
- #elixir-ls
- #phoenix_html
- #iex
- #graphql
- #genstage
- #websockets
- #supervisor
- #advent-of-code
- #distillery
- #processes
- #elixirconf-eu
- #api
- #forms
- #security
- #metaprogramming










Showing Posts 1 to 10- Show Best Posts
- Show All (oldest first)
- Show All (newest first)
plcholder
Cool project, how much are RAGs useful in todays climate? seems like their usecase was due to early models needing over correction in the prompts, which doesnt seem like a thing for the latest frontier model. Im also on a similar timeline, started taking LLM’s seriously around early 25 and trying to get an edge without burning thousands on the latest claude model
caleb-bb
If the frontier models are not omniscient, then you still need RAG.
plcholder
How So? Ive honestly gotten lazy with my setup since gpt 5.4 and since then after a couple of prompts with the right harnesses I get good results and steering things properly without needing RAG , im daily driving deepseek latest model now because of cost , so I might need it, but so far dsh and v4 flash have felt great
What im trying to get at is RAG seems like a thing of the past for most models now
caleb-bb
RAG is used for things besides writing code.
plcholder
caleb-bb
Many places where specialized knowledge with citations is required.
For example, if you want a legal LLM that can give (somewhat plausible) answers for a paralegal to verify, you want it to research legal documents and generate an answer with – this is important – citations included. This is one reason there are so many startups with legal RAG.
Insurtech is another big one. You want answers about insurance policies, again, with citations to specific documents that you can click/verify. RAG goes there too.
Basically any place where there’s a body of documentation and someone is getting paid to look up documents and answer questions based thereon is at least a candidate for RAG.
plcholder
That’s Interesting, now that you mention it I guess private data being uploaded and structured locally is an important use case for RAG, with legal RAG defining tool calls will also be super simple. I have not played around with RAG for quite a bit, but from your experience with it so far can it super charge a local model say qwen3.8 27b?
I’m also interested in insurtech and legal llm, can you talk more about that market and the demand, sounds super interesting, I can see the use case for legal LLM because writing up documents and court orders is a pain in the ass especially for small firms, but not Insurtech, also what other industries are looking for such tech?
caleb-bb
That’s actually an excellent use case. A lot of places are gonna want local RAG for privacy or regulatory reasons – this will be a huge thing with medicine in the USA because of HIPAA and similar. Smaller models almost by definition have less memorized general knowledge, much less the specialized knowledge that crops up in RAG use cases.
The upshot of that is that local LLMs will be increasingly important. Local models being generally smaller, you’ll want RAG to get accurate answers on things.
plcholder
yeah there are lots of open source chinese models and variants on hugginface, I think the perfect size range is between like 5b - 30b, its been awhile since I played with one, I have a llama & gemma models sitting around on my computer. I didnt know startups were seeing demand in real sectors I just assumed openAI and Anthropic swallowed up all competition, but privacy and token cost is an important variable plus I believe for most mundane tasks local models are really great when tool call and data is explicitly set.
I just wish to know what industry needs such tooling that isnt legal as its heavily crowded and elixir + the beam is an excellent language and runtime for orchestrating these models together
caleb-bb
Medicine and insurance are two big ones.
Also, tech support in general could use RAG because lower-tier tech support reps are basically paid to be RAG applications.