pironim

pironim

I’m trying to find good way to migrate data for Phoenix application.

Phoenix has regular migrations and it looks like rails migration.
My previous rails project was big and we often write data migrations(to update some records or delete)
rule of thumb - separate regular(schema) migrations from data migrations
for rails we have several libraries like GitHub - ilyakatz/data-migrate: Migrate and update data alongside your database structure. · GitHub
and you can call it like:

rake db:migrate:with_data 

I find GitHub - samsamai/ecto_immigrant: Data migrations for your ecto-backed elixir application · GitHub but it looks outdated and there is no options.

Maybe somebody know good library to add migrations and have ability to run data migration after specific schema migration?

there different opinions how to run data migrations
link them to release and run after schema migrations
, mix schema and data migrations
run migrations after schema migrations.

schema depends on your app

Showing Posts 1 to 10

LostKobrakai

LostKobrakai

Can you explain the benefit from that? To me this sounds like it’s just waiting for some inconsistency between those two to happen.

pironim

pironim OP

In two years you can get many migrations with outdated code and you need to fix them from time to time(module removed or renamed). It’s hard to do If you mix schema changes and data changes

with separate file for data migration you can get some control
e.g. run only schema migrations on local machine (dev, test env)
disable specific data migration and skip evaluation.

it’s all about control, testability and managebility.

Usually data migration have short life time

  1. create it
  2. write tests
  3. release to all envs
  4. comment or remove it somehow
LostKobrakai

LostKobrakai

Ok, I don’t have such problems because I migrate data using plain sql, therefore I can keep them around without problems.

shanesveller

shanesveller

I’d highly encourage writing these kinds of transformations as pure modules and functions that wrap straightforward Ecto calls, i.e. Repo.update_all. This approach lends itself well to testing, repeatable execution, and will look familiar to developers as they’re essentially specialized context modules. Tying them to the mindset of migrations that can be “checked off” as completed and don’t need to happen again doesn’t often mesh with my experience in reality. More often I see situations where it takes several passes through the data with the same intent to fully clean it in-place before a stricter DDL design can be applied. You can still drop the modules and their tests from your codebase once the need is conclusively finished, but you won’t need any acrobatics to repeat the effort again.

That’s also why I generally recommend that traditional Ecto migrations only contain DDL statements and no updates or insertions if at all avoidable.

LostKobrakai

LostKobrakai

I even skip anything MyApp.Repo, because it means :my_app would need to be started to be able to run migrations. If I really don’t know the sql I do MyApp.Repo.to_sql/2 and copy it in the migration, where it’s run using execute(sql) or execute(up, down).

dimitarvp

dimitarvp

True, but for one-off, deterministic and simple conversions from an old column to a new column Ecto migratory updating code works okay even months in the future.

Agreed with everything else.

pironim

pironim OP

I understand your point.
It’s mostly linked to migrations structure and testability

On previous rails project we use similar approach class which contain everything and it’s easy to test it.
Data migration was kind of humble object to run Migrator class with all logic

let me show rails example with data migration file

# cat db/data/20190814115720_remove_redundant_records.rb
class RemoveRedundantRecords < ActiveRecord::Migration
  def self.up
    DataMigrationJob.perform_sync Migrators::OldNotificationsDestroyer.name, 0
  end

  def self.down
  end
end

You can test Migrator like regular class and run it during db:migrate

Also to turn off or remove migration it should be executed everywhere and after certain period e.g. 1 year it can be disabled/removed.

dimitarvp

dimitarvp

One thing that saved the sanity of me and my colleagues both in Ruby on Rails and Elixir/Phoenix projects was to periodically squish all migrations into a singular .sql file with the contemporary schema and then just “restart” migrations (i.e. have zero of them after the squish).

At one point you can absolutely feel how maintaining the old migration scripts and making sure every new onboarded dev goes through them without errors is just not worth it.

But I still don’t recommend the squish / pruning to take place before 6 months (or 50-100 migrations).

shanesveller

shanesveller

A really nice compromise on this idea is Ecto.Migrator.with_repo/2 which was introduced in the 3.x series. You won’t have your full OTP app running, but you can start a skeleton Ecto Repo with minimal connections for the duration of your DB interactions, which don’t have to be traditional migrations. You can readily use Repo calls or context modules, as long as the connection count is set appropriately for your needs. We use this in Distillery custom commands, seeds, etc. to avoid having the whole supervision tree start up just to do some INSERTs.

Something fairly similar is built-in as mix ecto.dump and mix ecto.load, and doesn’t require you to fully discard the previous migration code content. Instead it just loads a structure snapshot and then marks previous migrations up to that point as applied, and they can still be replayed from-scratch using the normal ecto.migrate to ensure continuity for environments that aren’t able to use this approach, like your Production database. I don’t personally endorse fully squashing or discarding the historical content ourselves when you can get all of the same speed improvements using built-in functionality and still have a great escape hatch when needed.

shanesveller

shanesveller

This is part of the disconnect, perhaps, because I’m advocating for treating data-munging tasks as operationally and conceptually separate from schema evolution, and considering that a virtue rather than an inconvenience.

Where Next? Top

Trending in Questions Top

katta
I having some trouble figuring out if I have set myself too strict of standards for my production server. Currently I can handle 75% of r...
New
brecabral
Documentation While reading the Scoped Routes section, I noticed that the documentation currently refers to a problem without explainin...
New
achenet
Hello, I’m trying to build a basic Phoenix web-app, and I’d like to use Tailwind. However, when I launch mix phx.server, I get an error...
New
kpanic
Hi everyone, I am toying with the idea of building a “match maker” for giving personal help to people that wants to start coding. I sta...
New
asweet-confluent
I recently noticed that Elixir’s Logger defaults its primary log level to :debug when no :logger, :level application configuration is pre...
New
Cxx-mlr
I’m working on a small exercise involving update_in/3, and I came up with this solution: data = %{ name: "Periodic Table", category:...
New
ChrisAmelia
I’ve got trouble wrapping my head around the order in which functions are called in this snippet (from Phoenix’s authentication): toke...
New

Other Trending Topics Top

mjason
I’ve been building Longx, an open-source, self-hosted workspace for AI agents. The backend uses Elixir, Phoenix, Ash, and SQLite; the web...
New
lucaspauli
Hello all :wave: I’m happy to share LiveAnimate, a library for declarative UI animations in Phoenix LiveView. The idea: add one &lt;.m...
New
GenericJam
Edit: 2026 May 15 - This post is archived. Mob is alive!! Main docs: mob v0.7.11 — Documentation A bit of explanation for the slightly c...
New
JesseHerrick
Hey, I’m Jesse and I’m the main contributor behind Dexter, a full-featured, lightning-fast Elixir LSP optimized for large codebases. It s...
New
mudasobwa
I am happy to introduce the very α version of the new programming language compiled to BEAM. Welcome Cure. It has literally three kille...
New
garrison
Hobbes is a low-level distributed database for the Elixir programming language. Hobbes provides a simple, safe, and scalable storage lay...
New

We're in Beta

About us Mission Statement

Options

Thread Display Mode




Thread Preview

Skip Thread Previews