coen.bakker

coen.bakker

Blink is a library for fast bulk data insertion into PostgreSQL databases using the COPY command. It provides a clean, declarative syntax for defining seeders.

Features:

  • Uses PostgreSQL’s COPY for fast bulk inserts
  • Tables inserted in declaration order to respect foreign key constraints
  • Access data from previously defined tables when building subsequent tables
  • Store auxiliary context data that won’t be inserted into the database
  • Load data from CSV/JSON files with Blink.from_csv/2 and Blink.from_json/2
  • :transform option for type conversion when loading from files
  • Integrates with ExMachina nicely
  • Rollback on errors
  • Adapter pattern for supporting other databases

Example:

defmodule MyApp.Seeder do
  use Blink

  def call do
    new()
    |> add_table(:users)
    |> add_table(:posts)
    |> insert(MyApp.Repo)
  end

  def table(_store, :users) do
    [
      %{id: 1, name: "Alice", email: "alice@example.com"},
      %{id: 2, name: "Bob", email: "bob@example.com"}
    ]
  end

  def table(store, :posts) do
    users = store.tables.users
    # Build posts referencing users...
  end
end

Links:

https://github.com/nerds-and-company/blink

Showing Posts 1 to 6

Asd

Asd

Good library.

I’ve read the code and found a couple of fairly obvious bugs (like non-escaped strings in generated CSV) and limitations (like reading everything in memory), so I made a PR with fixes.

I am also providing fairly cheap consultancy services if you want to have this kind of review and contribution in your private projects.

coen.bakker

coen.bakker OP

Great. Ty.

I was aware of the memory issue and had a fix in mind similar to the one in your PR. I’ll have a closer look when I have time.

coen.bakker

coen.bakker OP

v0.5.0 Released

Version 0.5.0 is now available. This release marks a big step toward 1.0.0 — it covers all the major changes I had planned. Now the focus shifts to gathering feedback, fixing bugs, and addressing any remaining breaking changes before 1.0.0 (though I don’t have any in mind).

The headline feature is stream support, which enables memory-efficient seeding of large datasets.

Both table/2 clauses return streams in the example below, but returning lists still works as before.

defmodule Blog.Seeder do
  use Blink

  def call do
    new()
    |> with_table("users")
    |> with_table("posts")
    |> run(Blog.Repo, timeout: :infinity)
  end

  def table(_seeder, "users") do
    Stream.map(1..200_000, fn i ->
      %{
        id: i,
        name: "User #{i}",
        email: "user#{i}@example.com",
        ...
        inserted_at: ~U[2024-01-01 00:00:00Z],
        updated_at: ~U[2024-01-01 00:00:00Z]
      }
    end)
  end
  
  def table(seeder, "posts") do
    users_stream = seeder.tables["users"]

    Stream.flat_map(users_stream, fn user ->
      for i <- 1..20 do
        %{
          id: (user.id - 1) * 20 + i,
          title: "Post #{i} by #{user.name}",
          body: "This is the content of post #{i}",
          user_id: user.id,
          ...
          inserted_at: ~U[2024-01-01 00:00:00Z],
          updated_at: ~U[2024-01-01 00:00:00Z]
        }
      end
    end)
  end
end

Other highlights

  • JSONB support — nested maps are automatically JSON-encoded during insertion
  • Configurable timeout — :timeout option for long-running transactions
  • Configurable batch size — :batch_size option controls stream chunking (default: 10,000 rows)
  • Performance improvement — CSV encoding executes significantly faster
  • Bug fix — CSV escaping now correctly handles pipes, quotes, newlines, and backslashes

Breaking changes

  • Blink.Store → Blink.Seeder
  • insert/3 → run/3
  • add_table/2 → with_table/2
  • add_context/2 → with_context/2
  • Return values simplified to :ok (raises on failure)
  • Adapter call/4 callback now receives table_name as a string

Full changelog: v0.5.0 release

Asd

Asd

You missed a couple of other important things from my PR:

  1. Doing

    try do
      adapter.call(...)
    rescue
      UndefinedFunctionError ->
        raise "Module #{inspect adapter} must implement call/4"
    end
    

    is a strange approach. Removing the try completely would result in the more readable and meaningful exception.

    Plus, it is a buggy approach. Take for example a situation then the call function itself calls an undefined function. This try clause would hide this error, making debugging a nightmare

  2. You new approach opens and parses a CSV file twice in stream mode. First one to get the headers and second one to stream the data. This is not an issue when there is a one huge file, but it is an issue when there are a lot of small files. Opening a file is an operation which is more expensive than reading from a file

coen.bakker

coen.bakker OP

  1. Yes, I see. That does make sense. I removed the try-rescue block.
  2. Also changed this now.

Together they make up the changes of version 0.5.1 (see changelog)

Thank you.

Currently, exploring concurrent database connections for faster seeding, while keeping the API clean.

coen.bakker

coen.bakker OP

v0.6.0 Released

Version 0.6.0 is now available. This version brings parallel COPY operations, enabling significantly faster bulk inserts when seeding data.

defmodule Blog.Seeder do
  use Blink

  def call do
    new()
    |> with_table("categories")
    |> with_table("events")
    |> run(Blog.Repo, max_concurrency: 8)
  end

  def table(_seeder, "categories"), do: # ...
  def table(_seeder, "events"), do: # ...
end

Highlights:

  • Parallel COPY operations — the new :max_concurrency option (default: 6) allows batches to be inserted using multiple database connections in parallel
  • Per-table options — configure :batch_size and :max_concurrency per table via with_table/4
  • New guide — added Configuring Options guide

Full changelog: CHANGELOG.md

— All posts loaded —

Where Next? Top

Trending in Announcing Top

type1fool
WebAuthnLiveComponent WebAuthnComponents See this post about renaming the package. Passwordless authentication for Phoenix LiveView app...
New
GenericJam
Edit: 2026 May 15 - This post is archived. Mob is alive!! Main docs: mob v0.7.11 — Documentation A bit of explanation for the slightly c...
New
woylie
I released Doggo, a collection of unstyled Phoenix components. https://github.com/woylie/doggo Features Unstyled Phoenix components....
New
JesseHerrick
Hey, I’m Jesse and I’m the main contributor behind Dexter, a full-featured, lightning-fast Elixir LSP optimized for large codebases. It s...
New
ahamez
Hi everyone, I’ve been working on this protobuf library for 3 years. We use it in the company I work for, EasyMile, to communicate with ...
New
garrison
Hobbes is a low-level distributed database for the Elixir programming language. Hobbes provides a simple, safe, and scalable storage lay...
New
kip
I’ll shortly be launching Text, a nascent text analysis library. Current functionality In this early version (not ready for prime time) ...
New

Other Trending Topics Top

mudasobwa
I am happy to introduce the very α version of the new programming language compiled to BEAM. Welcome Cure. It has literally three kille...
New
mhanberg
Hi everyone! The first release candidate for the Expert language server project is now available! We’ve published a press release detai...
New
budgie
A little off-topic, but I feel like people here have a good head on their shoulders. I used to be quite good at making software. Was luc...
New
webofbits
With AI doing more of the implementation work, I’ve been wondering how much coding I should deliberately keep doing myself. My main conc...
#ai
New
bartblast
Hey folks, I just published a post about Hologram’s funding and where the project goes next - the short version: Curiosum as Main Spons...
New
budgie
I love Elixir. It’s one of 2 programming languages I’ve ever fallen in love with. But I don’t use it anymore. Serverless was the promis...
New

We're in Beta

About us Mission Statement

Options

Thread Display Mode




Thread Preview

Skip Thread Previews