gorghoa
Hi!
We are using Bumblebee.Audio with whisper. It works pretty well and we are enjoying it.
We would like to enhance some domain specific words transcription though, we saw that Whisper can have an extra textual prompt for this (see Whisper prompting guide ).
We did not find a way to leverage this prompting facilities with Bumblebee.Audio. Is this correct?
If whisper prompting is not yet available within bumblebee, what would it take to implement it? With some guiding, maybe we could help?
Thanks!
Rodrigue
Trending in Questions
I having some trouble figuring out if I have set myself too strict of standards for my production server. Currently I can handle 75% of r...
New
Documentation
While reading the Scoped Routes section, I noticed that the documentation currently refers to a problem without explainin...
New
Hello,
I’m trying to build a basic Phoenix web-app, and I’d like to use Tailwind.
However, when I launch mix phx.server, I get an error...
New
Hi everyone,
I am toying with the idea of building a “match maker” for giving personal help to people that wants to start coding.
I sta...
New
So my question is quite simple and i have found no conclusive answer on forum, google or AI.
Should we use :erlang.float for Integer to ...
New
I recently noticed that Elixir’s Logger defaults its primary log level to :debug when no :logger, :level application configuration is pre...
New
So i have been using ash framework for a while and i love it. However currently the issue im having with ash framework is the error handl...
New
Other Trending Topics
Edit: 2026 May 15 - This post is archived.
Mob is alive!!
Main docs: mob v0.7.11 — Documentation
A bit of explanation for the slightly c...
New
Hey, I’m Jesse and I’m the main contributor behind Dexter, a full-featured, lightning-fast Elixir LSP optimized for large codebases. It s...
New
I am happy to introduce the very α version of the new programming language compiled to BEAM.
Welcome Cure.
It has literally three kille...
New
Hobbes is a low-level distributed database for the Elixir programming language.
Hobbes provides a simple, safe, and scalable storage lay...
New
Hi everyone!
The first release candidate for the Expert language server project is now available!
We’ve published a press release detai...
New
A little off-topic, but I feel like people here have a good head on their shoulders.
I used to be quite good at making software. Was luc...
New
Categories:
Sub Categories:
Forums
Popular Tags
- #ecto
- #liveview
- #troubleshooting
- #learning-elixir
- #library
- #deployment
- #erlang
- #testing
- #genserver
- #mix
- #absinthe
- #remote-other
- #otp
- #plug
- #how-to-question
- #macros
- #postgres
- #elixirconf
- #channels
- #exunit
- #discussion
- #code-sync
- #podcasts
- #javascript
- #onsite
- #dialyzer
- #docker
- #authentication
- #umbrella
- #full-time-contract
- #podcasts-by-brainlid
- #ai
- #ecto-query
- #elixirconf-us
- #blog-post
- #elixir-ls
- #phoenix_html
- #iex
- #graphql
- #genstage
- #websockets
- #supervisor
- #advent-of-code
- #distillery
- #processes
- #elixirconf-eu
- #api
- #forms
- #metaprogramming
- #hex










Showing Posts 1 to 3- Show Best Posts
- Show All (oldest first)
- Show All (newest first)
garrison
I didn’t dig very far, but it kinda looks like those tokens just get thrown into the decoder at the start like an LLM. I don’t see any explicit way to do this in the
Bumblebee.AudioAPI as it seems to offer only a high-level stream interface, but it probably wouldn’t be too hard to add.jstimps
Have you seen reliable results from prompting Whisper with the official python sdk? We have tried the method described in the linked blog post and found that the results are extremely sensitive to the contents of the prompt, to the point where changing a single
,to a.yielded very different results.As such we weren’t able to find a reliable way to actually effect the output. Curious to hear your experience!
garrison
This is just what LLMs are like when they’re not instruction-tuned (time flies). Now you see why ChatGPT was so popular
I know multimodal instruction-tuned LLMs exist. I don’t know anything about using them for speech-to-text, but I imagine it has been done.