Fl4m3Ph03n1x

Fl4m3Ph03n1x

Background

I have to read a CSV and currently this is happening at compile time as the function runs in a module attribute:

# Imagine this csv file has 3 columns "sport, country, league"
@csv_sports_data 
  :my_app
  |> :code.priv_dir()
  |> Path.join("awesome_csv.csv")
  |> File.stream!()
  |> CSV.decode!(headers: false, separator: ?;)
  |> Stream.map(&List.to_tuple/1)
  |> Enum.uniq()

So now, because this runs at compile time (iirc) I have a variable with the data I need in tuple format. So far so good.

Problem

The problem comes when I need to do the same thing, multiple times, with small variations:

# The duplication, IT BURNS !!!

# Imagine this csv file has 3 columns "sport, country, league"
@csv_sports_data 
  :my_app
  |> :code.priv_dir()
  |> Path.join("awesome_csv.csv")
  |> File.stream!()
  |> CSV.decode!(headers: false, separator: ?;)
  |> Stream.map(&List.to_tuple/1)
  |> Enum.uniq()

@sports 
    :my_app
    |> :code.priv_dir()
    |> Path.join("awesome_csv.csv")
    |> File.stream!()
    |> CSV.decode!(headers: false, separator: ?;)
    |> Stream.map(&List.to_tuple/1)
    |> Stream.uniq()
    # Always trim data from pesky users!
    |> Stream.map(
      fn {sport, country, league} ->
        {String.trim(sport), String.trim(country), String.trim(league)} 
      end)
    |> Stream.map(fn {sport, _country, _league} -> sport end)
    #No empty sports!
    |> Enum.filter(fn sport -> sport != "" end) 

  @countries 
    :my_app
    |> :code.priv_dir()
    |> Path.join("awesome_csv.csv")
    |> File.stream!()
    |> CSV.decode!(headers: false, separator: ?;)
    |> Stream.map(&List.to_tuple/1)
    |> Stream.uniq()
    # Always trim data from pesky users!
    |> Stream.map(
      fn {sport, country, league} ->
         # Always trim data from pesky users!
        {String.trim(sport), String.trim(country), String.trim(league)} 
      end)
    |> Stream.map(fn {_sport, country, _league} -> country end) 
    # We allow empty countries to make the example interesting

As you can see, I have a lot of duplicated code. At the very least I could place

:my_app
    |> :code.priv_dir()
    |> Path.join("awesome_csv.csv")
    |> File.stream!()
    |> CSV.decode!(headers: false, separator: ?;)
    |> Stream.map(&List.to_tuple/1)
    |> Stream.uniq()

Into a function or variable and then re-use it in @sports and countries. The trimming function is also another candidate. And then there are the little differences for @sports and @countries where I select only the values I want.

Things I tried

So, my first try was to use the @csv_sports_data inside the @sports and @countries attributes. Obviously this didn’t work, as I can’t use something that was not yet compiled into an attribute that is itself being generated at compile time.

# This wont work
@sports 
   @csv_sports_data
    # Always trim data from pesky users!
    |> Stream.map(
      fn {sport, country, league} ->
        {String.trim(sport), String.trim(country), String.trim(league)} 
      end)
    |> Stream.map(fn {sport, _country, _league} -> sport end)
    #No empty sports!
    |> Enum.filter(fn sport -> sport != "" end) 

My second try was to consider Macros. According to my understanding, I could create a Macro that reads the CSV file at compile time and then have @sports and @countries use it. However, I personally am a believer of the saying:

“The first rule about Macros - don’t use Macros”

And I feel the usage of a Macro for this specific situation would be quite overkill. So I would like to avoid it.

And then there is also the trim function:

Stream.map(
      fn {sport, country, league} ->
         # Always trim data from pesky users!
        {String.trim(sport), String.trim(country), String.trim(league)} 
      end)

Which I cannot place inside a def or defp for the sake of reuse.

What now?

Surely I am missing something. Perhaps the solution I was given to work with the CSV is flawed, or perhaps I am forgetting some mechanism that would reduce the amount of duplicated code I have.

  • How can I remove all the duplication?

Showing Posts 1 to 10

michallepicki

michallepicki

You can operate on values (not module attributes) in a module body, do some computations and only then assign them to attributes:

csv_sports_data = :my_app
  |> :code.priv_dir()
  |> Path.join("awesome_csv.csv")
  |> File.stream!()
  |> CSV.decode!(headers: false, separator: ?;)
  |> Stream.map(&List.to_tuple/1)
  |> Stream.uniq()
  |> Enum.map(fn {sport, country, league} ->
    {String.trim(sport), String.trim(country), String.trim(league)}
  end)

  @csv_sports_data csv_sports_data

And re-use the already computed value to declare other module attributes:

  @sports csv_sports_data
  |> Enum.map(fn {sport, _country, _league} -> sport end)
  |> Enum.filter(fn sport -> sport != "" end)

  @countries csv_sports_data
  |> Enum.map(fn {_sport, country, _league} -> country end)
michallepicki

michallepicki

You can also move logic to other module that will become a dependency so it will get compiled earlier, where you can split your logic in functions however you like, for example:

defmodule SportsCsvReader do
  def read_sports_data() do
    :my_app
    |> :code.priv_dir()
    |> Path.join("awesome_csv.csv")
    |> File.stream!()
    |> CSV.decode!(headers: false, separator: ?;)
    |> Stream.map(&List.to_tuple/1)
    |> Stream.uniq()
    |> Enum.map(fn {sport, country, league} ->
      {String.trim(sport), String.trim(country), String.trim(league)}
    end)
  end

  def extract_sports(sports_data) do
    sports_data
    |> Enum.map(fn {sport, _country, _league} -> sport end)
    |> Enum.filter(fn sport -> sport != "" end)
  end

  def extract_countries(sports_data) do
    sports_data
    |> Enum.map(fn {_sport, country, _league} -> country end)
  end
end

and then you’ll be able to use it directly in your other module:

  csv_sports_data = SportsCsvReader.read_sports_data()
  @csv_sports_data csv_sports_data
  @sports SportsCsvReader.extract_sports(csv_sports_data)
  @countries SportsCsvReader.extract_countries(csv_sports_data)

edit: or re-use this logic in any other module

eksperimental

eksperimental

I think what you are trying to achieve defining attributes is what you should be doing but defining macros.

Using a macro to optimize things that can be calculated at compiled time is a perfectly use for a macro.

You just need to figure out what is available at compile time and what is not, reuse and move into functions the rest.

eksperimental

eksperimental

in addition if calculating countries and sports is an expensive operation, build a really simple caching system with ETS.
I don’t think you can get any faster than this :wink:

eksperimental

eksperimental

exactly! but i would convert read_sports_data/0 into a macro

LostKobrakai

LostKobrakai

That‘s not really needed. Macros would only make things more complex, as at no point AST has to be modified.

Fl4m3Ph03n1x

Fl4m3Ph03n1x OP

Does this work for Elixir 1.5?

Currently this code is not working:

  csv_data =
    :my_app
    |> :code.priv_dir()
    |> Path.join("awesome_csv.csv")
    |> File.stream!
    |> CSV.decode!(headers: false, separator: ?;)
    |> Stream.map(&List.to_tuple/1)
    |> Stream.uniq
    # Always trim data from pesky users!
    |> Enum.map(fn {sport, country, league} ->
      {String.trim(sport), String.trim(country), String.trim(league)}
    end)

  @compiled_csv_data csv_data

  @sports
    @compiled_csv_data
    |> Stream.map(fn {sport, _country, _league} -> sport end)
    |> Enum.filter(fn sport -> sport != "" end)

  # I suck at sports, so I need to know which ones I can play!
  def hard_sport?(sport) when sport in @sports, do: false

Error

== Compilation error in file lib/my_app/csv_reader.ex ==
** (ArgumentError) invalid args for operator “in”, it expects a compile-time list or compile-time range on the right side when used in guard expressions, got: nil

I believe this happens because of csv_data, which is not executed at compile time (I think).

lud

lud

You must remove the = here. This is a common mistake I do all the time :smiley:

I agree with @LostKobrakai , you don’t need macros here as you are not creating code by generating AST.

You should create a helper module specialized in reading your CSVs with all the required variations and parameters.
Those functions will be available at runtime, obviously. Then, in your main module you would require this helper module, so you can also call the functions at compile time.

michallepicki

michallepicki

Yes, as mentioned by @lud you need to fix the module attribute declaration. The error message is not intuitive, though!

michallepicki

michallepicki

I believe you don’t need to require a module to use its functions at compile time :slight_smile:

Where Next? Top

Trending in Questions Top

RSP87
I’m working on a project that simulates the bumbl example in the programming phoenix book. It acts almost like an email client. We have a...
New
kszambelanczyk
Hello! Could someone please give me a help/sample code, how to delete a file from s3 using waffle/waffle_ecto from Phoenix app. I creat...
New
RemyXRenard
I’m seeing that a list inside a Kino.DataTable will be interpreted as a charlist, even if the Kino.configure() is set to charlists: :as_l...
New
velrest
So my question is quite simple and i have found no conclusive answer on forum, google or AI. Should we use :erlang.float for Integer to ...
New
samoloth
Hi, I’ve just set up an application with ash_authentication. There is only magic link strategy for now, so there is no confirmation add o...
New
FlyingNoodle
If a change or preparation module uses Ash.Changeset.get_argument/2 or Ash.Query.get_argument/2 (or any of the other get_argument functio...
New
nseaSeb
Hello, I know there is an approach for handling lists that allows for optimized traversal, but I can’t recall the specific method (somet...
New

Other Trending Topics Top

mudasobwa
I am happy to introduce the very α version of the new programming language compiled to BEAM. Welcome Cure. It has literally three kille...
New
marciok
Hi there! We created Gust: A task orchestrator inspired by Airflow. For those who have never heard about Aiflow, it’s a Python-based wor...
New
jimsynz
Beam Bots (or just BB for short) is a framework for building fault-tolerant robotics applications in Elixir using familiar OTP patterns. ...
New
Dmk
Xamal is a deployment tool for Elixir apps that deploys native releases to bare metal servers over SSH. It’s a port of GitHub - basecamp/...
New
Damirados
Hello everyone. After busy few months I am happy to announce v0.1.0 of Emerge & Solve. They are GUI (Emerge) and State management (S...
New
netoum
Corex is an accessible, unstyled UI component library for Phoenix that integrates Zag.js state machines using Vanilla JavaScript and Live...
New

We're in Beta

About us Mission Statement

Options

Thread Display Mode




Thread Preview

Skip Thread Previews