aime
Hi, I’m trying to load Whisper models into Elixir using Nx, EXLA, and Bumblebee.
I’ve encountered the following issue:
When using the whisper-large-v3 model to transcribe audio files under 200 KB, the process consumes over 15 GB of RAM.
This is the script I created:
defmodule WhisperLarge do
alias Bumblebee
require Logger
def run do
Nx.global_default_backend(EXLA.Backend)
{:ok, model} = Bumblebee.load_model({:hf, "openai/whisper-large-v3"})
{:ok, featurizer} = Bumblebee.load_featurizer({:hf, "openai/whisper-large-v3"})
{:ok, tokenizer} = Bumblebee.load_tokenizer({:hf, "openai/whisper-large-v3"})
{:ok, generation_config} = Bumblebee.load_generation_config({:hf, "openai/whisper-large-v3"})
generation_config = Bumblebee.configure(generation_config, max_new_tokens: 448)
serving =
Bumblebee.Audio.speech_to_text_whisper(
model,
featurizer,
tokenizer,
generation_config,
compile: [batch_size: 1],
chunk_num_seconds: 3,
timestamps: :segments,
language: "es",
stream: false
)
result = Nx.Serving.run(serving, {:file, "file.wav"})
text =
result.chunks
|> Enum.map(& &1.text)
|> Enum.join(" ")
IO.puts("Transcripción: #{text}")
end
end
Dependencies:
{:bumblebee, "~> 0.6.0"},
{:nx, "~> 0.9.0"},
{:exla, "~> 0.9.0"}
I’ve also tested with other models like whisper-tiny, which don’t consume nearly as much memory — but they are not as accurate for my use case.
Someone on the Elixir Slack suggested using whisper.cpp or Python, which does use significantly less memory.
However, I was hoping to accomplish the full transcription process entirely in Elixir.
I’d really appreciate any advice or suggestions on reducing memory usage with large models in Nx/Bumblebee.
Trending in Questions
I’m working on a project that simulates the bumbl example in the programming phoenix book. It acts almost like an email client. We have a...
New
Hello,
I know there is an approach for handling lists that allows for optimized traversal, but I can’t recall the specific method (somet...
New
Hi everyone,
I am toying with the idea of building a “match maker” for giving personal help to people that wants to start coding.
I sta...
New
Documentation
While reading the Scoped Routes section, I noticed that the documentation currently refers to a problem without explainin...
New
So my question is quite simple and i have found no conclusive answer on forum, google or AI.
Should we use :erlang.float for Integer to ...
New
I recently noticed that Elixir’s Logger defaults its primary log level to :debug when no :logger, :level application configuration is pre...
New
apply_graft/2 doesn’t rewrite an add_many sub-workflow’s deps on an add step. Grafted jobs cancel with “upstream job was deleted”
Version...
New
Other Trending Topics
Hey, I’m Jesse and I’m the main contributor behind Dexter, a full-featured, lightning-fast Elixir LSP optimized for large codebases. It s...
New
I am happy to introduce the very α version of the new programming language compiled to BEAM.
Welcome Cure.
It has literally three kille...
New
Hi there! We created Gust: A task orchestrator inspired by Airflow.
For those who have never heard about Aiflow, it’s a Python-based wor...
New
Hi everyone!
The first release candidate for the Expert language server project is now available!
We’ve published a press release detai...
New
Beam Bots (or just BB for short) is a framework for building fault-tolerant robotics applications in Elixir using familiar OTP patterns. ...
New
Xamal is a deployment tool for Elixir apps that deploys native releases to bare metal servers over SSH. It’s a port of GitHub - basecamp/...
New
Categories:
Sub Categories:
Forums
Popular Tags
- #ecto
- #liveview
- #troubleshooting
- #learning-elixir
- #library
- #deployment
- #erlang
- #testing
- #genserver
- #mix
- #absinthe
- #remote-other
- #otp
- #plug
- #how-to-question
- #macros
- #postgres
- #elixirconf
- #channels
- #exunit
- #discussion
- #code-sync
- #podcasts
- #javascript
- #onsite
- #dialyzer
- #docker
- #authentication
- #umbrella
- #full-time-contract
- #podcasts-by-brainlid
- #ecto-query
- #elixirconf-us
- #ai
- #blog-post
- #elixir-ls
- #phoenix_html
- #iex
- #graphql
- #genstage
- #websockets
- #supervisor
- #advent-of-code
- #distillery
- #processes
- #api
- #forms
- #metaprogramming
- #hex
- #security











Showing Posts 1 to 1- Show Best Posts
- Show All (oldest first)
- Show All (newest first)
mgwidmann
it may explain what you’re seeing, but I just learned this: Bumblebee: Slow load_model in GenServer, slow Nx.Serving.run in exs file - #12 by jonatanklosko