quda

quda

Is there a library/framework in Elixir for massive distributed processing of large data sets (BigData) ?
Actually, I need something like Apache Hadoop.
I found some mentions that GenStage+Flow could do it, but I cannot figure it and cannot grasp the “massive distributed parallel processing” in these frameworks. How to spread it on tens on nodes/cluster to perform “massive” MapReduce to flows of +50TB of data?
Are there other Elixir solutions for this ?

PS: For the time being we are doing this big data processing using a (ancient) custom stack developed on PhP/Hadoop, but as we need to migrate the entire back-end to Elixir I have to look for an Elixir (easy to implement, performant, sustainable) solution for this back-end tool.

Showing Posts 1 to 6

bdarla

bdarla

You could find some relevant material in Concurrent Data Processing in Elixir (PragProg).
In addition to Gestate and Flow, the book discusses Broadway. According to the book, Broadway “offers a convenient way to build data-ingestion pipelines that consume events from external message brokers, like RabbitMQ, Apache Kafka, Amazon SQS, and more”.
I have to admit that although I have not read the book myself yet, but it is certainly relevant to your interests.

LostKobrakai

LostKobrakai

GenStage, Flow, Broadway can be building blocks for something like that, but they‘re certainly not a solution for that problem without a lot of additional work put into the parts not provided by them. Especially the distribution part does not exist in them.

quda

quda OP

Quite indeed. This could be a project per se that requires lots of time/resources for study, planning, design, testing etc. and not a side task to complete “by the end of the month”.

Now I have to convince the customer about this. :grimacing:

MrDoops

MrDoops

You can also look at GitHub - elixir-explorer/explorer: Series (one-dimensional) and dataframes (two-dimensional) for fast and elegant data exploration in Elixir · GitHub for fast columnar / OLAP workloads - some sort of Broadway ETL setup that grabs a CSV, imports to Explorer, runs your OLAP workloads, then loads the aggregated/transformed results is a very common sort of pipeline Explorer would be good at.

I’d recommend also looking at solutions such as https://flink.apache.org/ or https://materialize.com/ depending on your use case.

Flow + Broadway might get you far enough though if PHP + Hadoop is working already. I have a sneaking suspicion the Elixir ecosystem, maybe via the Nx projects, will end up with some Apache BEAM / distributed dataflow type tooling. At the very least something similar to Explorer taking elixir expressions that execute via Rust NIFs but in streaming dataflow contexts like Apache BEAM (the other BEAM - not our BEAM, but which is also cool).

SirWerto

SirWerto

Maybe you can give a look to Vessel. It is a interface with Hadoop written in Elixir. I don’t know the status of the package but could be helpful if you are keeping Hadoop in your stack.

quda

quda OP

Vessel looks quite promising for our scope, but sadly it seems abandoned and un-released: “Vessel is currently in a pre-v1 state”. And not a word about distributed processing.

— All posts loaded —

Where Next? Top

Trending in Questions Top

RSP87
I’m working on a project that simulates the bumbl example in the programming phoenix book. It acts almost like an email client. We have a...
New
nseaSeb
Hello, I know there is an approach for handling lists that allows for optimized traversal, but I can’t recall the specific method (somet...
New
RemyXRenard
I’m seeing that a list inside a Kino.DataTable will be interpreted as a charlist, even if the Kino.configure() is set to charlists: :as_l...
New
velrest
So my question is quite simple and i have found no conclusive answer on forum, google or AI. Should we use :erlang.float for Integer to ...
New
brecabral
Documentation While reading the Scoped Routes section, I noticed that the documentation currently refers to a problem without explainin...
New
samoloth
Hi, I’ve just set up an application with ash_authentication. There is only magic link strategy for now, so there is no confirmation add o...
New
FlyingNoodle
If a change or preparation module uses Ash.Changeset.get_argument/2 or Ash.Query.get_argument/2 (or any of the other get_argument functio...
New

Other Trending Topics Top

mudasobwa
I am happy to introduce the very α version of the new programming language compiled to BEAM. Welcome Cure. It has literally three kille...
New
marciok
Hi there! We created Gust: A task orchestrator inspired by Airflow. For those who have never heard about Aiflow, it’s a Python-based wor...
New
jimsynz
Beam Bots (or just BB for short) is a framework for building fault-tolerant robotics applications in Elixir using familiar OTP patterns. ...
New
Dmk
Xamal is a deployment tool for Elixir apps that deploys native releases to bare metal servers over SSH. It’s a port of GitHub - basecamp/...
New
netoum
Corex is an accessible, unstyled UI component library for Phoenix that integrates Zag.js state machines using Vanilla JavaScript and Live...
New
webofbits
With AI doing more of the implementation work, I’ve been wondering how much coding I should deliberately keep doing myself. My main conc...
#ai
New

We're in Beta

About us Mission Statement

Options

Thread Display Mode




Thread Preview

Skip Thread Previews