coen.bakker

coen.bakker

Is this `seeds.exs` idiomatic Elixir?

One of the requirements of the Book Club exercise on elixirland.dev is to seed data.

I am referring to this requirement in the exercise.

## Requirements

...

  ### **Seeding**
  * Running mix ecto.setup creates the database tables but also seeds the database
  * Seeding inserts 4,000 books that each have 10 pages
  * Some seeded books have an active page, but not all
  * Seeding is fast

How happy or unhappy would you be if a co-worker would create a pull request containing this seeds.exs? Is it written in idiomatic Elixir? Use of comments, etc.

import Ecto.Query, only: [from: 2]
alias BookClub.Repo
alias BookClub.Books.{Book, Page}

# Setup
# Sets log level to :info for performance
initial_log_level = Logger.level()
Logger.configure(level: :info)
IO.puts("Start database seeding")
start_time = System.os_time(:millisecond)

# Constants
n_books = 4000
n_pages_per_book = 10
inserted_at = NaiveDateTime.utc_now(:second)
batch_size = 800

# Insert 4,000 books
books =
  for _ <- 1..n_books do
    %{
      title: XlFaker.generate_title(),
      inserted_at: inserted_at,
      updated_at: inserted_at
    }
  end

books
|> Enum.chunk_every(batch_size)
|> Enum.each(&Repo.insert_all(Book, &1))

# Batch insert 10 pages per book
# Gives some books an active page
book_ids =
  from(b in Book, select: b.id)
  |> Repo.all()

pages =
  book_ids
  |> Enum.flat_map(fn book_id ->
    pages =
      for i <- 1..n_pages_per_book do
        %{
          book_id: book_id,
          content: XlFaker.generate_page(),
          number: i,
          status: :inactive,
          inserted_at: inserted_at,
          updated_at: inserted_at
        }
      end

    List.update_at(
      pages,
      :rand.uniform(n_pages_per_book) - 1,
      &Map.put(&1, :status, Enum.random([:active, :inactive]))
    )
  end)

pages
|> Enum.chunk_every(batch_size)
|> Enum.each(&Repo.insert_all(Page, &1))

# Teardown
end_time = System.os_time(:millisecond)
IO.puts("Finish database seeding")
IO.puts("Seeded #{n_books} books in #{end_time - start_time}ms")
Logger.configure(level: initial_log_level)

First Post! Switch mode

LostKobrakai

LostKobrakai

I don’t think the log level should be overwritten. Either the system is configured to info level and info is logged or it isn’t. I’d also consider changing the IO.puts to use logger.

The book ids could be returned from the Repo.insert_all calls with the :returning option + a flat_map.

And as with all code generating data: This “works”, but it doesn’t really do much in terms of making sure it continues to work over time and in the face of changes to the system it seeds. There’s a reason tests commonly use factories instead of code like this. Though this is one piece of code, tests are usually quite many, so it’s less of a problem.

Most Liked

fuelen

fuelen

I’d be okay if this script is just a temporary solution. The code itself is good enough. My concerns are not about style of Elixir code but more about approach in general. It works for simple cases, but as the project grows, I’d not put

* Seeding is fast

to the requirements.
What is more important for me is correctness of data. The approach of inserting raw data will quickly become a mess on several dozen tables. I’m a proponent of using business functions in seed scripts. I’m okay to wait a bit longer. Inserting seed data can be sped up by parallelization (Task.async_stream or flow can help). Because of that, having progress bars in seeds is also a good thing (owl can help).

The number of inserted data must be configurable, so it is possible to specify tiny numbers to run seeds as a part of the test suit. Just create 1 record for each record type to ensure that the script doesn’t fail. Most likely, you won’t be happy when the script stops working when you need it the most :slight_smile:

On one of the projects, we have modules for creating seeds defined in lib as a regular .ex file because we want these modules to be available in release. We run them when we launch a new server for testing. This happens quite often thanks to CI/CD. So, removing seeds.exs completely is OK, if someone is struggling with this just because the file is shipped with the default phoenix setup.

fuelen

fuelen

One tiny optimization could be applied with placeholders. Probably, it won’t be noticeable, but still good as an exercise.

dimitarvp

dimitarvp

  1. Seeds should be IMO idempotent which means that you should have corresponding Repo.delete or Repo.delete_all before inserts. This of course assumes that you can in fact identify the data that must be seeded and remove it at all which in many cases is not possible i.e. I worked in places where certain tenants / companies / users must always exist and some even had hardcoded IDs. You can still achieve idempotency by simply not deleting and not inserting anything if you detect that these special records that must be always exist are already there.

  2. I would not rely on the current datetime; I’d hardcode now as a fixed datetime in the past. To me determinism trumps almost all other concerns. However that too is not a hard rule because there are businesses where certain records are only valid if they’ve been created no more than f.ex. 6 months ago. So use your best judgement but still hardcode as much as you can if you can get away with it. Again – determinism.

Disagreed, seeds are practically their own universe and the only thing they share with the main app is the storage (in most cases a relational DB). It’s quite OK to override various runtime properties like logging in there. I even stuffed OTel tracing in one project’s seeds because they were taking mysteriously long (spoiler alert: one table was too big and the seeds scanned for records + did various levels of locking; we fixed it after).

Last Post!

coen.bakker

coen.bakker OP

To avoid loose ends: This code snippet doesn’t work as intended because the async stream never actually runs.

Fixed version: xlp-book-club-API/example/priv/repo/seeds.exs at main · elixirland/xlp-book-club-API · GitHub

Where Next?

Trending in Questions Top

stjefim
Hello! Suppose you are building workflow (order / task / payment) processing system with the following requirements: Each workflow con...
New
jonnycharles
I’m in search of an Elixir library that offers PDF generation capabilities similar to Ruby’s Prawn. While there have been discussions abo...
New
spammy
I’m looking to build a personal workflow to quickly deploy web applications written in elixir/phoenix, for local consumption (ie not on t...
New
silverdr
Using Phoenix.LiveView.TagEngine as an EEx.Engine is deprecated! To compile HEEx, use Phoenix.LiveView.TagEngine.compile/2 instead. Sta...
New
dli
Before I dive in myself, did anyone successfully sprinkle Hologram into their existing LiveView app? Looking for hints regarding: Addi...
New
bottlenecked
Hi all, I wanted to ask how the community is dealing with post-release steps. Today we have Ecto migrations, which make sure that the db...
New
michallepicki
I am using Oban and occasionally, shortly after a deployment, a handful of jobs can fail because of dependency on other parts of the syst...
New

Other Trending Topics Top

JesseHerrick
Hey, I’m Jesse and I’m the main contributor behind Dexter, a full-featured, lightning-fast Elixir LSP optimized for large codebases. It s...
New
jimsynz
Beam Bots (or just BB for short) is a framework for building fault-tolerant robotics applications in Elixir using familiar OTP patterns. ...
New
Damirados
Hello everyone. After busy few months I am happy to announce v0.1.0 of Emerge &amp; Solve. They are GUI (Emerge) and State management (S...
New
ausimian
Emily is an Elixir library that runs Nx computations on Apple’s MLX. Install it as the default Nx backend and Nx, defn, Axon, Nx.Serving,...
New
type1fool
I just stumbled on a newly redesigned elixir-lang.org. :tada: It looks like @Software_Mansion did the work, and I think it is generally a...
New
akoutmos
@hugobarauna and I (Alex Koutmos) have been hard at work on writing a book on Nerves that takes you from simply blinking LEDs to building...
New

We're in Beta

About us Mission Statement