Vidar

Vidar

I’m putting this in its own thread as that makes more sense.

The little prehistory here is that I have a RPI5 system with Hailo 8 working, but the non AI/ video part is not quite where we want it. The Rockchip 3588 processor, with SBCs often compared to the RPI5, seems like a more capable beast. It has twice the number of processor cores, but more importantly it has a built in NPU and very capable VPU. It also have other specialized hardware which might be of interest such as GPU, jpeg encoding and decoding, 48MPx ISP with functions like [auto focus, HDR, RAW conversion, lens corrections], 8K hdmi output, 4K hdmi input, more audio functions than I’ll ever need, various IO, and then some. So I got a Radxa 5T version, which also comes in an industrial version, with the full intent of running Nerves on it in anger.

Getting this up and booting with Nerves was quite straight forward. There are two different kernel options:

  1. The manufacturers with most hardware working fine but with a custom older kernel (6.1) and many proprietary blobs for drivers. (The NPU driver itself is actually open source). Issues I ran into with this one was getting WiFi working (the Radxa 5T SBC I got has a newer chip than 6.1), and getting a GPU accelerated browser for a kiosk mode using the Mali610 GPU. I backported the WiFi so that was fine, but getting a GPU accelerated browser was not so easy.

  2. I then tested the other alternative, an open source 6.18 kernel with patches for the 3588 made by Collabora with Mesa3D graphics, Rocket driver and Teflon TFlite delegate. I believe many of these or maybe all (?) are now in mainline Linux. Anyway, getting Nerves booted was troublefree, and I also got a GPU accelerated Chromium running in kiosk mode fairly easily. Headphone sound is a bottomless mystery to me though so that is not working yet.

    The problem with this is the NPU and VPU rely on the open source drivers and kernel to work. Teflon is young and has basically just implemented the operations needed to run their test Mobilenet (if I remember correctly). I’ve written more in another thread about trying to get these open source alternatives to work with Yolo 8, and in the end I basically explored the NPU registers and poked them directly. (The LUT was surprising as both Rocket and the official documentation had the number of LUT values wrong. Further, rather than being an actual LUT table it is more like a look up interpolated graph). In the end I found that the Rocket NPU initialization does not seem to support all that is needed for all operations on the NPU, and poking registers directly was hardly a good solution. So not really a way forward.

At that point I thought I could take the original Rockchip open source NPU driver and patch that into something that would work with the open source 6.18 kernel. But it seems I can’t do that without also having to change the 6.18 kernel (which is adopted to Rocket I presume). That actually makes sense as Rockchip themselves also had to fork the their kernel from the mainline to get their combination working. But if I changed the 6.18 kernel then other parts depending on it might accidentally break now or in the future. And all this was just for the NPU… I also wanted the VPU working, and the open source alternative is not full operation there either. A better plan was needed.

So, back to square one. Instead of trying to wrestle the Rockchip NPU and VPU drivers to fit with a 6.18 open source kernel the better plan was to keep the older Rockchip 6.1 kernel with the hardware working. And then somehow massage the Mali610 GPU driver blob into providing browsers with GPU acceleration.

Internetting I discovered that the key issue had been narrowed down to some missing communication between the browser and the renderer. And someone had even made a fix! But after applying the fix the speed was just 20 fps in the browser. The hook fix did make the GPU work, but it solved it by copying frames from the GPU to the CPU and then to memory for the renderer. Hence the low speed.

The better solution would be a zero frame copying by using references instead. So far that seems to work. The browser is now GPU accelerated at around 60 fps. Video playback in the browser should be accelerated by the specialized VPU though, so that is next on the list.

So far so promising.

Showing Posts 28 to 19

Vidar

Vidar OP

So the pipeline tests, as described above, gets about 45 fps with Yolo v8s. Latency 65ms, whereof 58ms is NPU. Doing 3 NPU cores on each frame should bring down the latency but it just marginally less at 50ms, but then with just 20fps overall. Something is off… After checking it turns out the model itself has to be onnx converted with a setting for using the cores in parallel. I’ll leave that for another day. (Did the check, no latency improvement. It is optimized for throughput not latency it seems, so round robin with 1 frame per core is the way to go).

Orange Pi 5 Plus. Maybe we are overthinking it. The device tree is supposed to cater to hardware differences, and the 3588 and the various libraries for the accelerators are all the same. Thus, I figured it might be worth trying a simple mix: The Radxa 5T setup but with the Orange Pi 5 Plus device tree instead. (The 40 pin IO pin might be different. Worth double checking first). So, here goes nothing, the Radxa 5T setup with a Orange Pi 5 Plus device tree instead:

https://github.com/BadBeta/nerves_experimental_radxa_orange_pi_5_plus_frankenbuild

Vidar

Vidar OP

Doing a full throughput test of the library now. Any issues that require kernel modifications should show up, but looking good so far.

The hardware of the Radxa 5T and the Orange can’t be very different? Wifi chipset and such? Any Orange custom kernel would have to provide the same functionality so they might be very similar. (Or if luck have it copied from each other). Some diff digging might solve it.

There are many products using the RK3588 so it would be a lot more interesting if this wasn’t limited to Radxa products. I kind of like the FriendlyElec and Lion SBC variants for their better housing, not too mention the industrial versions, tablets and information panels.

lawik

lawik

Nerves Core Team

Ah, you are probably right. I might try something later but I imagine there are enough differences to be trouble.

Vidar

Vidar OP

I could not get rid of all Radxa U-boot and this is all on Radxa’s proprietary BSP 6.1 kernel. Well, maybe the Orange is so similar it just runs? Crossing fingers!

The library might require a few changes in the kernel to make the zero copy composing possible, so there will likely be some more patches at that point.

Speaking of patches, the BSP 6.1 patch at the repo is a leftover from the NPU register exploration. Should not be needed anymore, so bad cleaning of me. I removed that too.

lawik

lawik

Nerves Core Team

I removed the references. Will see how it goes :slight_smile:

I am currently just building to see if it will boot. I imagine there might be a bunch of tweaks needed for the orange pi so might grab billal’s system if it doesn’t work.

Separate library is probably the right move, keep the system lean and strictly a compile-time concern.

Vidar

Vidar OP

Ah, there is some leftover reference from the larger kiosk build. The fix is likely just removal. I’ll try to check for any other leftovers too before fixing.

I did try to build that repo in isolation to check, but I guess the cunning computer grabbed that one from elsewhere.

You got hold of a Rock 5T or adapting to the Orange equivalent?

The pipeline composing will be a separate library I thought. The current repo is thus fairly naked booting hardware, but with this library on top it would Elixir access to all the hardware bits I find interesting.

Edit: I told Claude to go fetch the issue and check for more of similar nature. He disposed of that one and another similar line. So the error should be gone. Hopefully didn’t mess up anything else in the process.

lawik

lawik

Nerves Core Team

Cloned and attempted to build: Issue #1 logged :slight_smile:

lawik

lawik

Nerves Core Team

It makes sense to use what’s available and Zig seems a nice touch for wrapping it up well.

Yeah. I think different people and applications will want different things. Some will just want a processed but full-quality frame and work with that from Elixir, some want a video stream. Some want to involve the NPU at some stage but in what way, what model, what fps, I think there are devils all over those details.

I mean put together what you need and that’ll give anyone else stuff to build from :slight_smile:

Vidar

Vidar OP

My plan is sadly more boring and basic. I just use Zigler to wrap all commands of the existing C libraries from Rockchip to use these hardware accelerators.

These libraries are already meant to work together in a hardware pipeline from Rockchip. Each of the pipe elements is a hardware accelerator and the handoffs between them are done as zero copy using the processors purpose made DMA buffer. The RGA draws the Yolo rectangles directly onto the DMA buffer.

The only odd duckling out is the initial use of the ISP before the NPU. That seems less pipe integrated. The ISP takes a Bayer RGB input, and outputs a normal picture frame, while the NPU wants an interleaved data structure. So there are two options to solve that: 1. Modify the input of the AI models to take the ISP output as direct input. No conversion needed in the pipeline then, but a big hassle and likely causing the odd error. Or 2, which I have chosen, to use the RGA accelerator to do the conversion. The RGA should have no issue keeping up with the max 60fps flow rate (a limit of the 2x30 fps ISPs), but it does add a bit of latency.

When the NPU wants a new frame that has to be resized to model input before storing as a copy in the NPU ring buffer. The RGA does that too. The NPU setup is 4 frame positions for 3 NPU cores, so one is always being loaded so a new frame is ready when a core needs it. (That architecture takes 3 times more memory , so maybe making an option to use 3 cores for 1 frame. That apparently has less overall throughput though).

I’ve used Membrane some before but I wouldn’t say I’m really familiar with it. It is more like keyboard pecking, observe and hope for the best. Rinse and repeat until tests are happy. Over time some more systematic wisdom might accumulate. I plan to pipe the video output into Membrane to stream the output but still work to be done implementing the above.

References is the way here for sure, and the libraries and their DMA buffers are doing zero copy all along the hardware pipeline with exception for the sideline with NPU grab with resized copy. To get the main pipeline output from the DMA buffer at the end I currently also do a copy. I like to think there is a better solution to be found.

The NPU on the RK3588 is not supported by the XLA, so by extension I figured likely not by EXLA or NX either? That was my initial thinking anyway. And then I found that this 3588 has hardware accelerators sticking out everywhere so I thought I’d try to make it possible to compose them, line them up and use them from Elixir.

lawik

lawik

Nerves Core Team

Wow, you are tooling this thing up!

Would that example pipeline just be passing memory pointers or controlling a NIF? Or are you actually commanding the devices from Elixir?

For pipelines like this I think Membrane is an answer but you can always build a Membrane pipeline on top of a good base library but you can’t always build the tightest possible thing if you only have Membrane elements.

I think the Nx ecosystem has a fair number of references for how they do pointers and on-device stuff with NIFs and such to avoid copying back and forth. It has been a topic in ADBC + Explorer, in Pythonx + Nx and probably in EXLA/EMLX.

But essentially the Elixir code ends up passing around some kind of opaque reference that represents a device and memory location that is owned by a NIF typically.

The telling things when to pull a frame and where to throw it to can be pretty good in Elixir from what I’ve seen and having access to that from Elixir is ideal. But being able to tell it to just go at full clip is probably a good thing to have as well.

Where Next? Top

Trending in Discussions Top

cblavier
Hey there, It’s been more than a year since we started using LiveView as our main UI library and building a whole library of UI componen...
New
mudasobwa
I am happy to introduce the very α version of the new programming language compiled to BEAM. Welcome Cure. It has literally three kille...
New
budgie
A little off-topic, but I feel like people here have a good head on their shoulders. I used to be quite good at making software. Was luc...
New
axelson
Hi there! :wave: @frigidcode and I (but mostly him) have been running an Elixir Book club, we’re almost done with Designing Elixir Syste...
New
achempion
I’ve been using Emacs as my main code editor for more than a two years. It’s a custom build version although I’ve tried doom emacs and sp...
New
budgie
I love Elixir. It’s one of 2 programming languages I’ve ever fallen in love with. But I don’t use it anymore. Serverless was the promis...
New
jtormey
Lately I’ve been thinking about how to organize components as a LiveView application grows. One of the pain points I’ve found (for myself...
New

Other Trending Topics Top

GenericJam
Edit: 2026 May 15 - This post is archived. Mob is alive!! Main docs: mob v0.7.11 — Documentation A bit of explanation for the slightly c...
New
garrison
Hobbes is a low-level distributed database for the Elixir programming language. Hobbes provides a simple, safe, and scalable storage lay...
New
KristerV
Hey. Is there anyone here who creates agents in their apps? Not talking about using agents, but creating them. I’m finding it pretty diff...
New
mcass19
ExRatatui lets you cook up rich terminal UIs in Elixir, powered by Rust’s ratatui via Rustler NIFs. Build interactive terminal applicatio...
New
webofbits
With AI doing more of the implementation work, I’ve been wondering how much coding I should deliberately keep doing myself. My main conc...
#ai
New
georgeguimaraes
Just published claude-code-elixir, a plugin marketplace for Claude Code with Elixir support. These are the plugins I’ve been using for my...
New

We're in Beta

About us Mission Statement

Options

Thread Display Mode




Thread Preview

Skip Thread Previews