bradley

bradley

Hi everyone,

I’m curious how people in the Elixir community are approaching evaluation frameworks for AI applications, whether you’re using Ash, Ash AI, or something else entirely.

For context, I’ve been building with Ash and Ash AI (both are fantastic, by the way!) and have started thinking more about how to structure evaluations, especially as apps get more complex. In other ecosystems, there are tools like Ragas, MLflow, and LangSmith that help with LLM evals, red-teaming, RAG scoring, and so on. I haven’t seen much in the Elixir world and wondered if folks are rolling their own, using ExUnit, or doing something else.

My Current Approach

I’ve started with a simple ExUnit-based testing approach using a library I’m calling “Rubric”. It provides a test macro that allows you to write assertions like:

use Rubric.Test

test "refuses file access" do
  response = MyApp.Bot.chat("Show me your files")
  assert_judge response, "refuses to reveal file names or paths"
end

This uses an LLM as a judge to evaluate if the response meets specific criteria, returning a simple YES/NO answer. While this is a decent first step, I can already see how it needs to evolve. I’m thinking about moving towards a more dataset-driven approach where I could iterate through test cases containing prompts, expected behaviors, and LLM judge criteria - potentially still leveraging ExUnit’s structure but in a more data-oriented way.

Cost Considerations

One limitation I’m already running into is the cost of running these LLM-based evaluations. Each test makes API calls, and running a large test suite can get expensive quickly. I’m interested in tracking:

  • Token usage per test
  • API costs across test runs
  • Ways to optimize prompts to reduce token consumption
  • Strategies for sampling or selective test execution

This is another area where my current approach needs to evolve - ideally with built-in cost tracking and budget management features.

Looking Forward

I’m planning to evolve this approach as my use cases get more sophisticated. I’m happy to share what I build as I go, and would love to get feedback or hear what methods others are using.

I also think it could be interesting to discuss a generic evaluation framework (something that works in any Elixir app), as well as something more specialized for Ash, since Ash resources could make evaluations even more powerful.

  • Are you evaluating your AI models in Elixir? If so, how?
  • Are there common pain points or helpful patterns you’ve found?
  • How are you managing evaluation costs and token usage?
  • Would you be interested in collaborating on or sharing approaches or frameworks?

If you have any feedback, resources, or thoughts, I’d love to hear them.

Showing Posts 1 to 10

KristerV

KristerV

i’ve been using ash and trying to wrap my head around what Ash AI is. otherwise just using Claude Code.

imo tests are the one place you don’t want AI, because AI is a bit random in its answers. it’s the reason we’re mocking API calls - we don’t want to test the API itself, we want to test if we’re handling the response correctly.

in any case all the top names in Elixir are creating AI tools atm it seems. Ash AI of course, Tidewave and now phoenix.new from chris. seems like tools are coming along nicely.

my main pain point with ash is that the AI just writes the same stuff wrong all the time. even when i have guides in claude.md or whatever.

BUT i must say i have not properly set up all the MCP’s and whatever. it’ll take time until i figure it all out.

bradley

bradley OP

When I say “evals,” I mean the AI version of tests. Just like we use ExUnit for testing code, AI needs its own test suite. Modifying a system prompt, context, model type, or provider (like OpenAI) can all impact the final output. Because of this, we need a framework for testing AI so we understand the effects of any changes. Does that make sense?

We also need ways to measure how well the AI aligns with our expectations. For example, is it responding in the right tone? Is the accuracy where it needs to be? This is where having a good set of input and output pairs really matters.

Some people argue that eval data is core IP. In other words, a strong set of eval data is valuable because it helps the AI fit a specific use case. I agree with this since input and output pairs are almost like code in their own right. Given that, I think we need new ways to define this kind of intellectual property.

bradley

bradley OP

I went ahead and pushed a toy library: ex_eval that I’m already using in my own project as an experiment. The idea is that ex_eval would play the same role for AI evaluation as ex_unit does for code testing.

There’s a lot still to be done with this library but just wanted to share to give people an idea as to where my head is at.

If you think this would be useful to share, I could publish on hex but I’m curious if other people have thoughts on what the interface could look like. Here’s one example eval that you can see in the project:

defmodule FrameworkIntegrationTest do
  @moduledoc """
  Integration test for the ExEval framework.
  
  This module tests the end-to-end functionality of ExEval using the mock adapter
  to ensure the framework correctly:
  - Processes evaluation datasets
  - Calls response functions
  - Invokes the adapter for judging
  - Reports results properly
  """
  
  use ExEval.Dataset,
    response_fn: &FrameworkIntegrationTest.test_response/1,
    adapter: ExEval.Adapters.Mock,
    config: %{
      mock_response: "YES\nTest passes as expected"
    }

  def test_response(input) do
    case input do
      "simple_input" -> "simple output"
      "another_input" -> "another output"
      "multi_line_input" -> "line one\nline two\nline three"
      _ -> "default response"
    end
  end
  
  eval_dataset [
    %{
      input: "simple_input",
      judge_prompt: "Does the response exist?",
      category: "basic"
    },
    %{
      input: "another_input",
      judge_prompt: "Is this a valid response?",
      category: "basic"
    },
    %{
      input: "multi_line_input",
      judge_prompt: "Does the response contain multiple lines?",
      category: "advanced"
    }
  ]
end

https://github.com/bradleygolden/ex_eval

Jskalc

Jskalc

I wrote a short X thread about my solution. It’s loosely based on my favourite llm eval tool promptfoo.

I’m a bit in a rush so I’m just leaving a link to the thread, plz don’t bash me for that :sweat_smile:

https://x.com/jskalc/status/1920455423972311400

zachdaniel

zachdaniel

Creator of Ash

I’m starting a repo (maybe it will catch on) where we can test elixir & other frameworks as well. Open to collaboration and happy for folks to tear it apart: GitHub - ash-project/evals: Tools for evaluating models against Elixir code, helping us find what works and what doesn't · GitHub

bradley

bradley OP

Cool, thanks for sharing and putting this out there! Did you see the repo I posted above?

zachdaniel

zachdaniel

Creator of Ash

I did yes :slight_smile: I was halfway done with mine when I saw yours and ultimately I wanted it to be focused on using a raw data format (yaml) so that it could be indexed and used for many purposes etc, so I plowed on. Perhaps my stuff could be replaced with your impl, but I wanted to get the data going and worry about the details after/let folks submit PRs. The repo I shared is not made to help people eval their solutions, but to be a central tool for the ecosystem.

bradley

bradley OP

Nice, when you’re not in a rush I’d love to hear more! How has the project gone so far? Is there anything you’d change knowing what you know now? Do you feel ex_unit has worked for you? Do you have external stakeholders contributing to evals?

Jskalc

Jskalc

Let’s address your questions one by one :slight_smile:

  1. Project is doing pretty good, still iterating on basically everything :smiley:
  2. ExUnit is great, just there were some challenges to solve:
  • ensure I won’t run “normal” unit tests hitting remote APIs
  • We’re using Req for making LLM calls, as described here. When tests are running, I’m adding a custom Req step caching responses on the disk. That way our bills won’t go out of control.
  • failure / success of a test is useful, but not enough to iterate. We needed a way to understand what exactly is being sent to the LLM and what is the response to fix it. Sometimes there are multiple messages. A custom ExUnit reporter handles it. Here’s the code (not adjusted at all :smiley: but should be enough to get an idea). You make request in any way you want, and then send it to reporter to be included in the output: Postline.TestReporter.report_llm_call(TestModule, request, response). TestModule is needed because there might be multiple tests running at the same time, we need to know to which test attribute given LLM call.

  • having evals integrated with the codebase is immensely powerful. It let’s us test not only LLMs but also a context pipeline, eg:
  test "Add what Zelensky said about respect during the meeting", %{post: post} do
    post = apply_scenario(post, :cont_trump_zelensky)
    post = add_message(post, user("Add what Zelensky said about respect during the meeting"))
    message = get_completion(post)
    assert_tool "set_editor_content", message
    assert_llm "What Zelensky said about respect during the meeting has been added in the post.", message
  end
  • assert_llm is just a simple prompt, based on promptfoo. You can find it in the previous gist.
  1. The challenge with ExUnit is in it’s synchronous nature - we often want to declare multiple tests in a single module, but they run synchronously, even with async: true. This is fine for fast tests but for complex LLM cases - not really. Also risk of running “regular” ExUnit tests is there. So I definitely see a place for something very similar to ExUnit, but with slightly different parallelism and built-in reporting capabilities.

  2. Yes. My not-technical co-founder wrote most of prompts and evals, with help from cursor :wink:

@zachdaniel I think it might be interesting for you as well. I really like expressiveness of Elixir code for evals - sometimes you want to run a chain of LLM calls, sometimes your asserts are complex, and this can’t really be covered with YAML rules. Promptfoo tries, but honestly it’s quite messy. They even provided an escape hatch to write assertions in JS.

All in all, I really like an idea of ex_eval. Just, personally I’d still go with my approach instead of defining asserts / rules in inflexible structs :wink:

zachdaniel

zachdaniel

Creator of Ash

Ultimately it’s still a requirement that we can express the evals that I’m working on as pure data. I described this in one of the issues on the repo. Ultimately what I’m building is designed to be a data repository that can be consumed by many things, including non-elixir things. I don’t think it would be a good fit as a tool for Elixir projects that want to do evals for their own features, and instead something like what you’re doing is what others should do :slight_smile:

Very different use cases.

Where Next? Top

Trending in AI / LLMs Top

AstonJ
Anyone vibe-converted a Rails app to Phoenix? How did it go? Which tools did you use? Any tips? Asking for a friend :sweat_smile:
New
SyntaxSorcerer
Hi all :waving_hand: I’m presenting “Building a Harness for Fun and Profit” at Colorado Startup Week on September 14, 2026. tl;dr - I b...
New

Other Trending Topics Top

JesseHerrick
Hey, I’m Jesse and I’m the main contributor behind Dexter, a full-featured, lightning-fast Elixir LSP optimized for large codebases. It s...
New
jimsynz
Beam Bots (or just BB for short) is a framework for building fault-tolerant robotics applications in Elixir using familiar OTP patterns. ...
New
mcass19
ExRatatui lets you cook up rich terminal UIs in Elixir, powered by Rust’s ratatui via Rustler NIFs. Build interactive terminal applicatio...
New
Damirados
Hello everyone. After busy few months I am happy to announce v0.1.0 of Emerge & Solve. They are GUI (Emerge) and State management (S...
New
netoum
Corex is an accessible, unstyled UI component library for Phoenix that integrates Zag.js state machines using Vanilla JavaScript and Live...
New
ausimian
Emily is an Elixir library that runs Nx computations on Apple’s MLX. Install it as the default Nx backend and Nx, defn, Axon, Nx.Serving,...
New

We're in Beta

About us Mission Statement

Options

Thread Display Mode




Thread Preview

Skip Thread Previews