bradley
Hi everyone,
I’m curious how people in the Elixir community are approaching evaluation frameworks for AI applications, whether you’re using Ash, Ash AI, or something else entirely.
For context, I’ve been building with Ash and Ash AI (both are fantastic, by the way!) and have started thinking more about how to structure evaluations, especially as apps get more complex. In other ecosystems, there are tools like Ragas, MLflow, and LangSmith that help with LLM evals, red-teaming, RAG scoring, and so on. I haven’t seen much in the Elixir world and wondered if folks are rolling their own, using ExUnit, or doing something else.
My Current Approach
I’ve started with a simple ExUnit-based testing approach using a library I’m calling “Rubric”. It provides a test macro that allows you to write assertions like:
use Rubric.Test
test "refuses file access" do
response = MyApp.Bot.chat("Show me your files")
assert_judge response, "refuses to reveal file names or paths"
end
This uses an LLM as a judge to evaluate if the response meets specific criteria, returning a simple YES/NO answer. While this is a decent first step, I can already see how it needs to evolve. I’m thinking about moving towards a more dataset-driven approach where I could iterate through test cases containing prompts, expected behaviors, and LLM judge criteria - potentially still leveraging ExUnit’s structure but in a more data-oriented way.
Cost Considerations
One limitation I’m already running into is the cost of running these LLM-based evaluations. Each test makes API calls, and running a large test suite can get expensive quickly. I’m interested in tracking:
- Token usage per test
- API costs across test runs
- Ways to optimize prompts to reduce token consumption
- Strategies for sampling or selective test execution
This is another area where my current approach needs to evolve - ideally with built-in cost tracking and budget management features.
Looking Forward
I’m planning to evolve this approach as my use cases get more sophisticated. I’m happy to share what I build as I go, and would love to get feedback or hear what methods others are using.
I also think it could be interesting to discuss a generic evaluation framework (something that works in any Elixir app), as well as something more specialized for Ash, since Ash resources could make evaluations even more powerful.
- Are you evaluating your AI models in Elixir? If so, how?
- Are there common pain points or helpful patterns you’ve found?
- How are you managing evaluation costs and token usage?
- Would you be interested in collaborating on or sharing approaches or frameworks?
If you have any feedback, resources, or thoughts, I’d love to hear them.
Trending in AI / LLMs
Other Trending Topics
Categories:
Sub Categories:
Forums
Popular Tags
- #ecto
- #liveview
- #troubleshooting
- #learning-elixir
- #deployment
- #library
- #erlang
- #testing
- #genserver
- #mix
- #absinthe
- #remote-other
- #otp
- #plug
- #how-to-question
- #macros
- #postgres
- #channels
- #elixirconf
- #exunit
- #discussion
- #code-sync
- #javascript
- #podcasts
- #onsite
- #dialyzer
- #docker
- #authentication
- #umbrella
- #full-time-contract
- #podcasts-by-brainlid
- #ecto-query
- #elixir-ls
- #blog-post
- #phoenix_html
- #iex
- #graphql
- #ai
- #genstage
- #elixirconf-us
- #websockets
- #supervisor
- #advent-of-code
- #distillery
- #processes
- #api
- #forms
- #metaprogramming
- #security
- #hex










Showing Posts 1 to 10- Show Best Posts
- Show All Posts (oldest first)
- Show All Posts (newest first)
KristerV
i’ve been using ash and trying to wrap my head around what Ash AI is. otherwise just using Claude Code.
imo tests are the one place you don’t want AI, because AI is a bit random in its answers. it’s the reason we’re mocking API calls - we don’t want to test the API itself, we want to test if we’re handling the response correctly.
in any case all the top names in Elixir are creating AI tools atm it seems. Ash AI of course, Tidewave and now phoenix.new from chris. seems like tools are coming along nicely.
my main pain point with ash is that the AI just writes the same stuff wrong all the time. even when i have guides in claude.md or whatever.
BUT i must say i have not properly set up all the MCP’s and whatever. it’ll take time until i figure it all out.
bradley
When I say “evals,” I mean the AI version of tests. Just like we use ExUnit for testing code, AI needs its own test suite. Modifying a system prompt, context, model type, or provider (like OpenAI) can all impact the final output. Because of this, we need a framework for testing AI so we understand the effects of any changes. Does that make sense?
We also need ways to measure how well the AI aligns with our expectations. For example, is it responding in the right tone? Is the accuracy where it needs to be? This is where having a good set of input and output pairs really matters.
Some people argue that eval data is core IP. In other words, a strong set of eval data is valuable because it helps the AI fit a specific use case. I agree with this since input and output pairs are almost like code in their own right. Given that, I think we need new ways to define this kind of intellectual property.
bradley
I went ahead and pushed a toy library: ex_eval that I’m already using in my own project as an experiment. The idea is that ex_eval would play the same role for AI evaluation as ex_unit does for code testing.
There’s a lot still to be done with this library but just wanted to share to give people an idea as to where my head is at.
If you think this would be useful to share, I could publish on hex but I’m curious if other people have thoughts on what the interface could look like. Here’s one example eval that you can see in the project:
https://github.com/bradleygolden/ex_eval
Jskalc
I wrote a short X thread about my solution. It’s loosely based on my favourite llm eval tool promptfoo.
I’m a bit in a rush so I’m just leaving a link to the thread, plz don’t bash me for that
https://x.com/jskalc/status/1920455423972311400
zachdaniel
I’m starting a repo (maybe it will catch on) where we can test elixir & other frameworks as well. Open to collaboration and happy for folks to tear it apart: GitHub - ash-project/evals: Tools for evaluating models against Elixir code, helping us find what works and what doesn't · GitHub
bradley
Cool, thanks for sharing and putting this out there! Did you see the repo I posted above?
zachdaniel
I did yes
I was halfway done with mine when I saw yours and ultimately I wanted it to be focused on using a raw data format (yaml) so that it could be indexed and used for many purposes etc, so I plowed on. Perhaps my stuff could be replaced with your impl, but I wanted to get the data going and worry about the details after/let folks submit PRs. The repo I shared is not made to help people eval their solutions, but to be a central tool for the ecosystem.
bradley
Nice, when you’re not in a rush I’d love to hear more! How has the project gone so far? Is there anything you’d change knowing what you know now? Do you feel
ex_unithas worked for you? Do you have external stakeholders contributing to evals?Jskalc
Let’s address your questions one by one
Postline.TestReporter.report_llm_call(TestModule, request, response). TestModule is needed because there might be multiple tests running at the same time, we need to know to which test attribute given LLM call.The challenge with ExUnit is in it’s synchronous nature - we often want to declare multiple tests in a single module, but they run synchronously, even with async: true. This is fine for fast tests but for complex LLM cases - not really. Also risk of running “regular” ExUnit tests is there. So I definitely see a place for something very similar to ExUnit, but with slightly different parallelism and built-in reporting capabilities.
Yes. My not-technical co-founder wrote most of prompts and evals, with help from cursor
@zachdaniel I think it might be interesting for you as well. I really like expressiveness of Elixir code for evals - sometimes you want to run a chain of LLM calls, sometimes your asserts are complex, and this can’t really be covered with YAML rules. Promptfoo tries, but honestly it’s quite messy. They even provided an escape hatch to write assertions in JS.
All in all, I really like an idea of
ex_eval. Just, personally I’d still go with my approach instead of defining asserts / rules in inflexible structszachdaniel
Ultimately it’s still a requirement that we can express the evals that I’m working on as pure data. I described this in one of the issues on the repo. Ultimately what I’m building is designed to be a data repository that can be consumed by many things, including non-elixir things. I don’t think it would be a good fit as a tool for Elixir projects that want to do evals for their own features, and instead something like what you’re doing is what others should do
Very different use cases.