herisson
Hi everyone, I recently took on a personal challenge of running Meta’s Segment Anything Model in Elixir. Since it’s not supported by Bumblebee, I started using Ortex to run ONNX models.
Everything is mostly smooth, but my final masks are a bit distorted. Since it’s my first time using Nx/Ortex/ONNX, I’m struggling to figure out the problem. Could it be an issue with the ONNX model itself?
Here is the outputs I’m getting:
I’ve put my code into a gist so you can check it out.
If you have some time to take a look, I would appreciate it. I’ve run out of ideas ![]()
Thanks!
Trending in Questions
I’m working on a project that simulates the bumbl example in the programming phoenix book. It acts almost like an email client. We have a...
New
Hello,
I know there is an approach for handling lists that allows for optimized traversal, but I can’t recall the specific method (somet...
New
Hi everyone,
I am toying with the idea of building a “match maker” for giving personal help to people that wants to start coding.
I sta...
New
Documentation
While reading the Scoped Routes section, I noticed that the documentation currently refers to a problem without explainin...
New
So my question is quite simple and i have found no conclusive answer on forum, google or AI.
Should we use :erlang.float for Integer to ...
New
I recently noticed that Elixir’s Logger defaults its primary log level to :debug when no :logger, :level application configuration is pre...
New
Hi, I’ve just set up an application with ash_authentication. There is only magic link strategy for now, so there is no confirmation add o...
New
Other Trending Topics
Hey, I’m Jesse and I’m the main contributor behind Dexter, a full-featured, lightning-fast Elixir LSP optimized for large codebases. It s...
New
I am happy to introduce the very α version of the new programming language compiled to BEAM.
Welcome Cure.
It has literally three kille...
New
Hi there! We created Gust: A task orchestrator inspired by Airflow.
For those who have never heard about Aiflow, it’s a Python-based wor...
New
Hi everyone!
The first release candidate for the Expert language server project is now available!
We’ve published a press release detai...
New
Beam Bots (or just BB for short) is a framework for building fault-tolerant robotics applications in Elixir using familiar OTP patterns. ...
New
Xamal is a deployment tool for Elixir apps that deploys native releases to bare metal servers over SSH. It’s a port of GitHub - basecamp/...
New
Categories:
Sub Categories:
Forums
Popular Tags
- #ecto
- #liveview
- #troubleshooting
- #learning-elixir
- #library
- #deployment
- #erlang
- #testing
- #genserver
- #mix
- #absinthe
- #remote-other
- #otp
- #plug
- #how-to-question
- #macros
- #postgres
- #elixirconf
- #channels
- #exunit
- #discussion
- #code-sync
- #podcasts
- #javascript
- #onsite
- #dialyzer
- #docker
- #authentication
- #umbrella
- #full-time-contract
- #podcasts-by-brainlid
- #ecto-query
- #elixirconf-us
- #ai
- #blog-post
- #elixir-ls
- #phoenix_html
- #iex
- #graphql
- #genstage
- #websockets
- #supervisor
- #advent-of-code
- #distillery
- #processes
- #api
- #forms
- #metaprogramming
- #hex
- #security











Showing Posts 12 to 3- Show Best Posts
- Show All (oldest first)
- Show All (newest first)
Anko
hmm sorry I thought i had one but I sent the wrong link, still looking for the right one
update: vietanhdev/segment-anything-2-onnx-models at main
herisson
yes there was indeed an error on the gist sorry. However the livebook that @kip made is fully functional.
I’ll try to make it work with SAM 2 but I have yet to fin ONNX models of it.
Anko
I tried to run the livebook from here
https://raw.githubusercontent.com/elixir-image/image/main/livebook/segment_anything.livemd
It gives me
** (ArgumentError) cannot broadcast tensor of dimensions {3, 768, 1024} to {1, 3, 1024, 1024} with axes [1, 2, 3]
(nx 0.7.3) lib/nx/shape.ex:253: Nx.Shape.broadcast!/4
I had to change the resizing line to
to make it end up as square. Other options are here Image — image v0.48.1
Would love to have this work with SAM 2
herisson
That’s amazing, thank you!
Indeed, the image library makes things much simpler and cleaner. I didn’t know these badges existed, awesome!
Happy to help! I also learned a lot to get the code working.
kip
I’ve made a modified version of your Livebook to use Image for the image pipeline. I think it simplifies some of the code.
I’ve also made it so the encoder and decoder are downloaded from hugging face using
reqso the Livebook now works standalone. So here’s the badge!Thanks for taking the time to help me through my lack of understanding. Its been a good learning experience.
kip
Really appreciate it, thanks very much. Back to experimental mode!
herisson
That’s my bad, with all my different trials I got mixed up in my code versions…
I’ve uploaded my complete working code and models on hugging face which you can find here : ginkgoo/SegmentAnythingModel-Elixir-Ortex at main
That should do the work !
kip
Seems I’m still a bit stuck - and hoping your patience hasn’t run out
Using the encoder link you pointed me at, I’m seeing the following error:
And the error is:
Which I think is saying that the shape of the image data is
{1024, 1024, 3}but the model expects{3, 1024, 1024}.Given you already have a working model, may I still ask if you are able to upload the encoder and decoder you are using so I can at least eliminate that as an issue in my experiment?
kip
@herisson, much appreciated - I will give that a go and see where I get to. Many thanks.
herisson
Yes, their examples aren’t very comprehensive in terms of explanations. From what I understand, SAM is a “two-stage model.”
The first stage takes an image and transforms it into something the next model can understand (image embeddings). This stage, which I refer to as the encoder or vision encoder, is common in many image processing tasks, not unique to SAM. That’s why it’s not included in their export script.
The second stage, which you exported, takes the image embeddings and other inputs to produce the mask (referred to as the decoder).
I can’t upload the decoder to a repository right now as I’m not at home, but you can find the vision encoder here. Since you’ve exported the decoder, you should be set with these two models!
Hope it makes things clearer