victorolinasc

victorolinasc

Hi everybody!

I have an expensive CPU action that I want to serve through a public internet endpoint. It doesn’t matter what it does, just that it usually takes some time (seconds) and that it costs CPU (say a few percent of CPU usage). For a real use-case you can think of an expensive hashing algorithm like PBKDF or any other on a signin endpoint.

What I want to avoid is being hit by a flood of requests and suffer from resource starvation where everything would become unresponsive. Like the signin case, a public endpoint without authentication makes it a perfect target for attacks that want to bring the server down. Think that all network solutions are in place like rate-limiting, WAFs and so on.

I thought about 2 strategies in the BEAM:

1- Using a pool of processes. I’ve implemented this with poolboy and works fine but is hard to tune it. I have to benchmark the cost of the function in a production server and reach a pool size that will be a good enough “sharing” of resources. I am thinking about switching to wpool and having a bigger than needed pool with a callback module that would check CPU usage before dispatching to the pool. This seems to me very unreliable and prone to error… If we get several concurrent requests and the CPU usage from all of these is under control, then they would all start at the same time and the CPU would spike anyway…

2 - Using a slave node for doing just this operation and fight with the emulator flags to have it use other logical cores. Suppose I have 8 cpus available, I could dedicate 2 or 3 to the slave node and the rest to the main system. Though, with this strategy, it would still be vulnerable to resource starvation and make all calls to it fail which is not my intention here. I’d rather have timeouts than a denial of service.

So, I’d like to know if there are any other algorithms or strategies to deal with this. I appreciate your time :slight_smile:

First 10 of 19 Posts Switch mode

benwilson512

benwilson512

Author of Craft GraphQL APIs in Elixir with Absinthe

Have you thought about measures to prevent clients from misbehaving? For example with public signin pages, rate limits or tools like ReCaptcha can limit the impact an individual client can have.

victorolinasc

victorolinasc OP

I have thought about. Rate-limiting is applied but a DDoS would penetrate anyway. reCaptcha is also applied but, well, Google went down these days and brought reCaptcha with it… which made me think about this.

In anycase, I wanted to know if there are other papers on the subject or strategies to consider.

vfsoraki

vfsoraki

How much time can a user wait for his/her response? I mean if it takes 10 seconds normally for a process to finish, can the user wait like 10 minutes for it? Is it crucial for the system to do this calculation in a synchronized manner for the system to do other jobs?

I mean, if a user can wait a long time, then maybe you can queue these jobs and process them with a job processor.

This way, if the system is flooded, only the queue grows, CPU is not maxed out. You can monitor the queue length to see if your system is doing fine, or it is lagging behind. If you notice the system is not handling the load, then maybe you can scale your job processors for a limited time to clear the queue or scale your system to handle more load. You can also notify your users that you are under load. Maybe you cal also analyze your queue and remove some malicious jobs, if there are any.

You can do more fine-tuning too. Having multiple queues (high/low priorities), different job processors for them and somewhat easily horizontal scaling are to name a few.

There are a number of queue job processing libraries available for Elixir, a simple search reveals them.

victorolinasc

victorolinasc OP

This is similar to the process pool approach I believe. There the queue is the process mailbox and here is a persistent queue or an external system queue. In any case it is hard to tune it properly in different environments.

I wonder if there is some algorithm to auto-tune it according to available resources.

vfsoraki

vfsoraki

There is a subtle difference. The process mailbox is external, so it is more fault tolerant and more analysable.

I don’t know if there is an algorithm, but have a look at Horizon package that Laravel provides for its queue system. It has some kind of configuration that may help you or give you ideas.

derek-zhou

derek-zhou

If the job takes seconds to complete in the fast path, then you can consider to use a job queue system like:
https://github.com/sorentwo/oban

Let the general public to use a normal queue and paying customers to use a higher priority queue. This can scale up to many nodes.

victorolinasc

victorolinasc OP

Using an external queue system (be it RabbitMQ, Oban, Kafka or whatever) makes it a lot harder to have synchronous responses to reply to the end user. I could pass the PID of the request as an arg and send a reply later and so on or persist the result somewhere and check it later on, but I think this is a bit overwhelming and it doesn’t consider the environment I am running the requests as the meter to accept/deny more work.

What I want is closer to rate-limiting but should consider CPU and not requests per minute or something like that. I think I will try a custom rule from plug_attack

dom

dom

Is the work done in pure Erlang/Elixir code, a NIF, a dirty NIF, an external process?

victorolinasc

victorolinasc OP

The work involves crypto and that delegates to NIFs that are built-in the runtime (if I understand it properly).

chasers

chasers

You can easily do this with LiveView.

Where Next? Top

Trending in Questions Top

stjefim
Hello! Suppose you are building workflow (order / task / payment) processing system with the following requirements: Each workflow con...
New
jonnycharles
I’m in search of an Elixir library that offers PDF generation capabilities similar to Ruby’s Prawn. While there have been discussions abo...
New
spammy
I’m looking to build a personal workflow to quickly deploy web applications written in elixir/phoenix, for local consumption (ie not on t...
New
dli
Before I dive in myself, did anyone successfully sprinkle Hologram into their existing LiveView app? Looking for hints regarding: Addi...
New
roeland
Kia ora, We have been using elixir-google-api to connect to Google Drive. However, with the updates to Tesla due to CVEs this is now bro...
New
bottlenecked
Hi all, I wanted to ask how the community is dealing with post-release steps. Today we have Ecto migrations, which make sure that the db...
New
rahultumpala
Hello, I have an Elixir backend that implements a custom protocol over TCP. I want to load test the backend and assess the performance o...
New

Other Trending Topics Top

JesseHerrick
Hey, I’m Jesse and I’m the main contributor behind Dexter, a full-featured, lightning-fast Elixir LSP optimized for large codebases. It s...
New
jimsynz
Beam Bots (or just BB for short) is a framework for building fault-tolerant robotics applications in Elixir using familiar OTP patterns. ...
New
mcass19
ExRatatui lets you cook up rich terminal UIs in Elixir, powered by Rust’s ratatui via Rustler NIFs. Build interactive terminal applicatio...
New
Damirados
Hello everyone. After busy few months I am happy to announce v0.1.0 of Emerge & Solve. They are GUI (Emerge) and State management (S...
New
netoum
Corex is an accessible, unstyled UI component library for Phoenix that integrates Zag.js state machines using Vanilla JavaScript and Live...
New
ausimian
Emily is an Elixir library that runs Nx computations on Apple’s MLX. Install it as the default Nx backend and Nx, defn, Axon, Nx.Serving,...
New

We're in Beta

About us Mission Statement