KristerV

KristerV

State of developing agents with Elixir (not coding agents)

Hey. Is there anyone here who creates agents in their apps? Not talking about using agents, but creating them. I’m finding it pretty difficult. The LLM models themselves are not actually very intelligent it seems, the tooling is crucial to get a useful agent out of these text generators.

I decided to write down my findings and experiences so people don’t have to reinvent the wheel, but at the same time I’m hoping someone will tell me that I’m doing it all wrong and give me a recipe for success.

Ultimately a friend says they use the Claude Agent SDK and get their agents running without too much effort and with good results. That’s what I want. But all the official tooling is for the JS ecosystem, obviously. Anyway, here’s the current state.

Features an agent needs to be useful

  1. Plan-Act-Verify Loop: Explicit verification step after actions.
  2. State Machine Persistence: Durable storage of agent “thought” and “status” (Checkpoints).
  3. Context Compaction: Auto-summarizing history to prevent token bloat/model drift.
  4. Memory Layering: Separating ephemeral “Working Memory” from “Long-term Knowledge.”
  5. Human-in-the-Loop (HITL): Breakpoints for manual approval on sensitive actions.
  6. Model agnostic: Switching from one LLM model to another shouldn’t cause problems/rewrites.

This is not an exhaustive list for agents, just what I know is needed atm.

Industry tools

Here’s the tooling non-Elixir environments have

  1. Vercel AI SDK (Typescript)
  2. LangGraph (JS/Python).
  3. PydanticAI (Python)
  4. Claude Agent SDK (Python/TS)
  5. OpenAI Agents SDK (Python/TS)

Elixir tools

  1. LangChain (Elixir): An Elixir-native implementation of the LangChain framework. It focuses on “Chaining” processes and providing modular components for prompt management and model integration.
  2. Jido: OTP-native state-machine framework for autonomous agents (Action/Signal/Runner). This is the “power-user” choice for complex, long-running agent logic.
  3. Legion: A dedicated harness library specifically for building agentic loops and tool execution.
  4. AshAi / AshBaml: Crucial bridges for turning your existing Ash Resources and Actions into LLM-accessible tools.
  5. Instructor Ex: The go-to for type-safe data extraction and validation using Ecto schemas.
  6. ReqLLM: An LLM-specific client for the Req library.
  7. LLMAgent: A signal-based library designed for managing conversation flows and tool handlers.
  8. Oban: Essential for job persistence, concurrency control, and “Resume from failure” logic.
  9. Elixir AI SDK: A community port that brings the Vercel AI SDK’s maxSteps and streaming patterns to Elixir.
  10. edit: just found Whisperer, looks like a good first layer of such a system.

My experience

I started with langchain, but switching models breaks the app code, which makes the lib kind of pointless. And they don’t really accept PR’s, but that’s understandable tbh. I switched to OpenRouter.ai and that works great with only a 5% addition to cost of the LLM models.

I currently have built my agents on Oban. A custom loop of messages, context pruning, planning etc. But it’s brittle. Now I have to rebuild my architecture to be more deterministic (because the models just aren’t very smart on their own).

The libraries look to me like each does one part of the puzzle or if it tries to do it all it’s just not very deep. And since docs are not deep either it’s one of those things where you spend a week trying them all out and then realize none of them do what you need. Out of all of them Legion seems to be the most all inclusive option with orchestrators and an agent loop. But no plan → exec → verify it seems. Which is kind of core, so don’t really want to jump into it.

Anyway, these are just my thoughts. I’m probably misunderstanding a bunch.

My question

My point is not to complain or point at missing pieces. Rather I have a feeling I’m missing the big picture. So I ask - how do you build agents? What libraries, techniques or approaches do you use?

Most Liked

mikehostetler

mikehostetler

Did you try Jido at all? I see it listed above - Jido has a sophisticated “ReAct” reasoning loop with tools powered by ReqLLM here - models switch pretty easily:

https://github.com/agentjido/jido_ai/blob/main/lib/jido_ai/agents/examples/weather_agent.ex

typesend

typesend

See also Sagents:

https://github.com/sagents-ai/sagents

(Launched very recently.)

tfwright

tfwright

I tend to agree with OPs reasons for using Oban.

Oban: Essential for job persistence, concurrency control, and “Resume from failure” logic.

In fact, I have found myself using Oban as my state machine library of choice since its states and arg persistence and retry handling cover pretty much all my uses cases. Before I found myself adding a “status” enum with only minor variation in values and repetitive transition logic to various contexts/schemas. Bringing in a new dep just to handle that kind of thing seemed unnecessary, but as I went to abstract it myself I realized I was recreating a bunch of APIs I already had at my disposal in Oban. So I started to lean more on it and most of that stuff just became worker config, leaving only the need to enqueue jobs with the appropriate args. Logic has much better SoC because the schemas now only express states that are actually specifically meaningful to the domain, and all the generic “error” “pending“ etc stuff is hidden away and protected with much better guarantees than I ever managed to maintain. And that’s before using any of the Pro features like workflows. So it’s hard for me to imagine the downside to using Oban for something like this.

Last Post!

KristerV

KristerV

oh, thanks. i did search, but this did not come up.

specifically this links is interesting: GitHub - druyang/awesome-elixir-llm-genai: A list of LLM and GenAI Elixir Resources/Tools · GitHub

Where Next?

Popular in AI / LLMs Top

AndyL
For development and prototyping, I’d like to retain a basic ability to perform LLM inference on my own hardware, using open source models...
#ai
New
New
Joser
Claude Code Plugin for Elixir: Custom Skills and Hooks for Better Code Quality I’ve been experimenting with Claude Code for Elixir devel...
New
dimamik
OpenAI published a repo built primarily with Elixir (96.1%). https://github.com/openai/symphony https://github.com/openai/symphony/blob...
New
Vidar
So in case others find them useful here are 3 of the skills I use. They work pretty well together. Avoid Claude starting agents to plan a...
New
DaAnalyst
Been using Claude for over a week now (Opus 4.6 then 4.7, max subscription). Honestly, can’t hide my joy, at some points feeling even ash...
New
DaAnalyst
How much would you really be willing to spend (more) to keep it going with Claude should Anthropic go berserk with the rates, or put diff...
New

Other popular topics Top

Qqwy
Original source of discussion: This topic on the Pragmatic Programmers’ Functional Web Development with Elixir, OTP, and Phoenix forum. ...
New
bsollish-terakeet
Credo is smart enough to check for (something like) this: assert length(the_list) == 0 with this response: Checking if an enum is empt...
New
dblack
I’ve got an issue with an app and I’ve no idea of how to troubleshoot it. I’m hoping someone here might have seen something similar. I p...
New
Patoshizzle
After calling mix ecto.create I get this error: 17:00:32.162 [error] GenServer #PID<0.412.0> terminating ** (Postgrex.Error) FATAL...
New
jason.o
In the code below, if the create action is not set to accept “extra_key” as an input, it errors out with a message shown above. Is there ...
New
AstonJ
Posting this to see if we can make things easier for people to get into Neovim. If you use Neovim and have a favourite distro please let ...
New

We're in Beta

About us Mission Statement