preciz

preciz

I trained a model last year that we still use in prod and I committed the code.
The model has >99% accuracy.

If I want to train the same model on the same data now it doesn’t learn, I might have forgot something but I can’t figure out what is the problem. The data also seems fine.

I made a GitHub repo that contains the data and the livebook.
https://github.com/preciz/not_learning?tab=readme-ov-file

The model is a binary classifier and no matter what I change it always ends up at 0.5 accuracy and the loss is increasing during training continuously.

Run in Livebook

Could somebody take a look at it and help me realize why the model doesn’t learn?

Showing Posts 1 to 6

billylanchantin

billylanchantin

Sorry, this is a drive-by comment. I’m not able review the particulars.

You say “The data also seems fine.” I my experience, that’s usually the problem even though I didn’t spot it at first. I suggest two things to sanity check that it’s not the data:

  1. Try to train another, simpler model on the same data. Can a different model find a signal in that data?
  2. Try to train your model on a contrived example. Can that model find a signal when it’s definitely present?

If you find that the answers to both those questions are “yes”, then truly something odd is happening. But in my own work, I find that at least one of those answers is usually “no”.

bdarla

bdarla

I replicated the issue. What seems weird (to me, not really an expert), is that I observe too many NaN values in the tensors in the output of the training step.

Is there any chance that the training data/labels are in a wrong format or that a kind of missing normalisation step is required?

preciz

preciz OP

Sean Moriarity wrote me that I should try axon 0.4.
With axon 0.4.1 the model is learning, the loss is going down during training.

Here is the diff:

<   {:axon, "~> 0.6"},
<   {:nx, "~> 0.7"},
<   {:exla, "~> 0.7"},
---
>   {:axon, "~> 0.4.1"},
>   {:nx, "~> 0.4.2"},
>   {:exla, "~> 0.4.2"},

< optimizer = Polaris.Optimizers.adamw(learning_rate: 1.0e-3)
---
> optimizer = Axon.Optimizers.adamw(1.0e-3)

https://github.com/preciz/not_learning/blob/master/model_learning.livemd

preciz

preciz OP

But now I would be happiest if newest Axon version would also learn. Any ideas on that?

NduatiK

NduatiK

There seems to be something wrong with binary_cross_entropy. Try categorical_cross_entropy for your loss function, it should be equivalent for two outputs:

loss =
  &Axon.Losses.categorical_cross_entropy(
    &1,
    &2,
    reduction: :mean
  )

If the change works on your end, could you report the issue to the Axon Github?

preciz

preciz OP

Awesome, with categorical_cross_entropy and softmax output activation the loss is going down and accuracy is >0.99.
Thank you, I will open an issue on Axon repo then.

— All posts loaded —

Where Next? Top

Trending in Questions Top

katta
I having some trouble figuring out if I have set myself too strict of standards for my production server. Currently I can handle 75% of r...
New
brecabral
Documentation While reading the Scoped Routes section, I noticed that the documentation currently refers to a problem without explainin...
New
achenet
Hello, I’m trying to build a basic Phoenix web-app, and I’d like to use Tailwind. However, when I launch mix phx.server, I get an error...
New
kpanic
Hi everyone, I am toying with the idea of building a “match maker” for giving personal help to people that wants to start coding. I sta...
New
asweet-confluent
I recently noticed that Elixir’s Logger defaults its primary log level to :debug when no :logger, :level application configuration is pre...
New
Cxx-mlr
I’m working on a small exercise involving update_in/3, and I came up with this solution: data = %{ name: "Periodic Table", category:...
New
ChrisAmelia
I’ve got trouble wrapping my head around the order in which functions are called in this snippet (from Phoenix’s authentication): toke...
New

Other Trending Topics Top

GenericJam
Edit: 2026 May 15 - This post is archived. Mob is alive!! Main docs: mob v0.7.11 — Documentation A bit of explanation for the slightly c...
New
JesseHerrick
Hey, I’m Jesse and I’m the main contributor behind Dexter, a full-featured, lightning-fast Elixir LSP optimized for large codebases. It s...
New
mudasobwa
I am happy to introduce the very α version of the new programming language compiled to BEAM. Welcome Cure. It has literally three kille...
New
garrison
Hobbes is a low-level distributed database for the Elixir programming language. Hobbes provides a simple, safe, and scalable storage lay...
New
mhanberg
Hi everyone! The first release candidate for the Expert language server project is now available! We’ve published a press release detai...
New
budgie
A little off-topic, but I feel like people here have a good head on their shoulders. I used to be quite good at making software. Was luc...
New

We're in Beta

About us Mission Statement

Options

Thread Display Mode




Thread Preview

Skip Thread Previews