kevinschweikert
Hey folks,
i am trying to get Membrane and Whisper working together. For now, the goal is, to use Mebrane so stream a file into Whisper. I worked on the following Livebook but cannot figure out, why Whisper is not correctly transcribing.
When i pipe the output of the converter into portaudio it is very noisy. When i set the converter to s16le the PortAudio audio sounds fine! But when i use :f32le and write it to a PCM file and import it into Audacity, it sounds fine again, despite the PortAudio Sink is not.
Either way, Whisper is not translating correctly. Can you help me out, where the issue could be?
Mix.install(
[
{:membrane_core, "~> 0.12.7"},
{:membrane_file_plugin, "~> 0.15.0"},
{:membrane_mp4_plugin, "~> 0.27.0"},
{:membrane_fake_plugin, "~> 0.10.0"},
{:membrane_raw_audio_format, "~> 0.11.0"},
{:membrane_ffmpeg_swresample_plugin,
git: "https://github.com/membraneframework/membrane_ffmpeg_swresample_plugin",
branch: "master"},
{:membrane_aac_plugin, "~> 0.15.0"},
{:membrane_aac_fdk_plugin, "~> 0.15.1"},
{:membrane_portaudio_plugin, "~> 0.16.1"},
{:kino_bumblebee, "~> 0.3.0"},
{:exla, "~> 0.5.1"}
],
config: [nx: [default_backend: EXLA.Backend], membrane_core: [logger: [verbose: false]]]
)
#dependecies:
brew install fdk-aac
brew install portaudio
model = "tiny"
{:ok, model_info} = Bumblebee.load_model({:hf, "openai/whisper-#{model}"})
{:ok, featurizer} = Bumblebee.load_featurizer({:hf, "openai/whisper-#{model}"})
{:ok, tokenizer} = Bumblebee.load_tokenizer({:hf, "openai/whisper-#{model}"})
{:ok, generation_config} = Bumblebee.load_generation_config({:hf, "openai/whisper-#{model}"})
generation_config = Bumblebee.configure(generation_config, max_new_tokens: 100)
generation_config = %{
generation_config
| forced_token_ids: [
# tell whisper that spoken text is german
{1, Bumblebee.Tokenizer.token_to_id(tokenizer, "<|de|>")},
{2, Bumblebee.Tokenizer.token_to_id(tokenizer, "<|transcribe|>")}
]
}
serving =
Bumblebee.Audio.speech_to_text(model_info, featurizer, tokenizer, generation_config,
compile: [batch_size: 1],
defn_options: [compiler: EXLA]
)
Nx.Serving.start_link(serving: serving, name: WhisperServing)
defmodule WhisperSink do
use Membrane.Sink
require Logger
def_input_pad(:input, demand_unit: :buffers, accepted_format: _any)
@impl true
def handle_playing(_ctx, state) do
{[demand: :input], state}
end
@impl true
def handle_write_list(:input, buffers, _ctx, state) do
for buffer <- buffers do
data =
buffer.payload
|> Nx.from_binary(:f32)
Logger.info(inspect(Nx.Serving.batched_run(WhisperServing, data)))
end
{[demand: {:input, length(buffers)}], state}
end
end
defmodule Pipeline do
use Membrane.Pipeline
@in_file "/path/to/video.mp4"
@impl true
def handle_init(_context, _opts) do
structure = [
child(:file_source, %Membrane.File.Source{location: @in_file, seekable?: true})
|> child(:demuxer, %Membrane.MP4.Demuxer.ISOM{optimize_for_non_fast_start?: true})
|> via_out(Pad.ref(:output, 2))
|> child(:depayloader_audio, Membrane.MP4.Depayloader.AAC)
|> child(:audio_parser, %Membrane.AAC.Parser{
in_encapsulation: :none,
out_encapsulation: :ADTS
})
|> child(:aac_decoder, Membrane.AAC.FDK.Decoder)
|> child(:converter, %Membrane.FFmpeg.SWResample.Converter{
output_stream_format: %Membrane.RawAudio{
sample_format: :f32le,
sample_rate: 16_000,
channels: 1
}
})
# |> child(:speaker, Membrane.PortAudio.Sink),
|> child(:whisper, WhisperSink),
get_child(:demuxer)
|> via_out(Pad.ref(:output, 1))
|> child(:fake, Membrane.Fake.Sink.Buffers)
]
{[spec: structure, playback: :playing], %{}}
end
@impl true
def handle_element_end_of_stream(:file_sink, _pad, _ctx_, _state) do
IO.inspect("terminate")
{[terminate: :normal], nil}
end
@impl true
def handle_element_end_of_stream(_child, _pad, _ctx, _state) do
{[], nil}
end
def handle_child_notification(notification, _element, _context, state) do
# dbg(notification)
{[], state}
end
end
{:ok, _, _} = Pipeline.start_link()
Trending in Questions
I having some trouble figuring out if I have set myself too strict of standards for my production server. Currently I can handle 75% of r...
New
Hello,
I’m trying to build a basic Phoenix web-app, and I’d like to use Tailwind.
However, when I launch mix phx.server, I get an error...
New
I’m working on a small exercise involving update_in/3, and I came up with this solution:
data = %{
name: "Periodic Table",
category:...
New
I’ve got trouble wrapping my head around the order in which functions are called in this snippet (from Phoenix’s authentication):
toke...
New
Is there any way to avoid the Hologram compiler running when using iex? It seems like the front-end code could potentially be disregarded...
New
** (ArgumentError) expected :max_attempts to be a positive integer, got: {:@, [line: 10, column: 19], [{:max_attempts, [line: 10, column:...
New
Hello folks!
So at work, we are seeing some situations where we have to define some “fixed” strings that are used across the codebase in...
New
Other Trending Topics
Edit: 2026 May 15 - This post is archived.
Mob is alive!!
Main docs: mob v0.7.11 — Documentation
A bit of explanation for the slightly c...
New
I am happy to introduce the very α version of the new programming language compiled to BEAM.
Welcome Cure.
It has literally three kille...
New
Hobbes is a low-level distributed database for the Elixir programming language.
Hobbes provides a simple, safe, and scalable storage lay...
New
A little off-topic, but I feel like people here have a good head on their shoulders.
I used to be quite good at making software. Was luc...
New
Hey. Is there anyone here who creates agents in their apps? Not talking about using agents, but creating them. I’m finding it pretty diff...
New
ExRatatui lets you cook up rich terminal UIs in Elixir, powered by Rust’s ratatui via Rustler NIFs. Build interactive terminal applicatio...
New
Categories:
Sub Categories:
Forums
Popular Tags
- #ecto
- #liveview
- #troubleshooting
- #learning-elixir
- #library
- #deployment
- #erlang
- #testing
- #genserver
- #mix
- #absinthe
- #remote-other
- #otp
- #plug
- #how-to-question
- #macros
- #postgres
- #elixirconf
- #channels
- #exunit
- #discussion
- #code-sync
- #podcasts
- #javascript
- #onsite
- #dialyzer
- #docker
- #authentication
- #umbrella
- #full-time-contract
- #podcasts-by-brainlid
- #ai
- #ecto-query
- #elixirconf-us
- #blog-post
- #elixir-ls
- #phoenix_html
- #iex
- #graphql
- #genstage
- #websockets
- #supervisor
- #advent-of-code
- #distillery
- #processes
- #elixirconf-eu
- #api
- #forms
- #metaprogramming
- #hex











Showing Posts 1 to 4- Show Best Posts
- Show All (oldest first)
- Show All (newest first)
mat-hek
Hi there, I’ll look into the converter & port audio issue. Regarding speech-to-text, have you tried this example?
kevinschweikert
Hi @mat-hek,
oh wow, thanks for the LiveBook example! I did not find it. After some trial and error and the membrane_transcription repo from @lawik i found out, that i obviously need to create a buffer, otherwise the length of the
Membrane.Buffer.payloadis to small for Whisper. The example would have saved me some timeI currently got this to work with a Whisper Serving startet in the Supervision tree and used manual demand. Otherwise the buffer would be too big for the Whisper model, because it can only handle a maximum chunk of 30s.
Again, thank you for your help!
mat-hek
Yeah, we’ll make those livebooks more discoverable. If the payload is too small/big, you can just concat/split it in the WhisperSink as well
However, to use automatic demands, you’d also have to move the logic from the task to the element or 
awaiton the task - otherwise, you can be flooded with dataBTW, you can use functions from Membrane.RawAudio — Membrane: raw audio format v0.12.3 instead of hardcoding sample rate & friends. Happy hacking
mat-hek
PortAudio bug fixed in Release v0.16.2 · membraneframework/membrane_portaudio_plugin · GitHub