voughtdq

voughtdq

As an exercise, I decided to convert a CSV, similar in format to the one located here, to a map. Please notice that the file is subject to some non-free license, just in case you planned to use the data commercially.

The idea is that I convert the three categories (provided as columns in the CSV file) into a map of maps of lists.

%{"Group1" => %{"Classification1" => ["Domain1", "Domain2", "Domain3"]},
              %{"Classification2" => ["Domain4", "Domain5", "Domain6"]},
              %{"Classification3" => []}, 
%{"Group2" => %{"Classification4" => ["Domain7"]}}

You’ll notice that a classification can have a missing domain (in other words, the tuple would be {grouping, classification, ""}.

I am using NimbleCSV to get the file and convert it to the map.

defmodule TaxonomyMap do
  @doc """
  Open the target file, parse it, and create a list of tuples of
  {grouping, classification, specialization} for each row.
  """
  def get_taxonomies(file) do
    file
    |> File.stream!(read_ahead: 1000)
    |> NimbleCSV.RFC4180.parse_stream
    |> Stream.map(fn [_, grouping, classification, specialization, _, _] ->
      {grouping, classification, specialization}
    end)
  end

  defp maybe_blank_list(value) do
    case value do
      "" -> []
      _ -> [value]
    end
  end

  defp classification_map(classification, specialization) do
    %{classification => maybe_blank_list(specialization)}
  end

  defp nil?(acc, key_or_keys) do
    get_in(acc, key_or_keys) |> is_nil
  end

  def run(taxonomies) do
    taxonomies
    |> Enum.reduce(%{}, fn({g, c, s}, acc) ->
      cond do
        nil?(acc, [g]) ->
          # The grouping is not in the map
          # Add the grouping, classification, and specialization for this row
          put_in(acc, [g], classification_map(c, s))
        nil?(acc, [g, c]) ->
          # The classification is not in the grouping
          # Add the classification and specialization to the grouping
          put_in(acc, [g], Map.merge(get_in(acc, [g]), classification_map(c, s)))
        !nil?(acc, [g, c]) ->
          # The classification and grouping both exist
          # Add the specialization to the grouping
          put_in(acc, [g, c], get_in(acc, [g, c]) ++ maybe_blank_list(s))
      end
    end)
  end

The way you’d use this is by running TaxonomyMap.get_taxonomies("taxonomy.csv") |> TaxonomyMap.run.

Is there anything that can or should be improved? I’d love to hear your thoughts on how I can improve for clarity or to make my code more “elixiric”.

Showing Posts 1 to 3

bbense

bbense

It all looks reasonably good to me. The one place that might be more “elixiry” is the cond switch inside the reduce.

I found that part somewhat hard to reason about without flipping back and forth in the file. You’re basically doing a test and then a transformation. It’s more ‘elixiry’ to simply write the transformations as pattern matching function heads or case statements and let the computer sort out which one to use.

|> Enum.reduce(%{}, fn({g, c, s}, acc) -> inject(acc, [g, c], s ) end 

def inject( acc, [group, class], spec ) do
       case get_in(acc, [group, class] ) do 
           nil   -> inject_class( acc, [group, class], spec)
           found -> put_in(acc, [group, class], found ++ maybe_blank_lists(spec))
       end 
end  

def inject_class(acc, [group, class], spec ) do 
     case get_in(acc, [group] ) do 
           nil   -> put_in(acc, [group], classification_map(class, spec))
           found -> put_in(acc, [group], Map.merge( found, classification_map(class, spec)))
     end 
end 

I’m not entirely happy with that, but hope it shows the idea. I feel like there is probably an additional refactoring involving pulling out the case statements from inject that might make the code even more straightforward.

voughtdq

voughtdq OP

Noice! Thanks for the pattern. This is exactly what I needed to hear.

I can’t wait to get back to my workstation to refactor this.

geofflangenderfer

geofflangenderfer

here are the headers and the first row of the input csv:

Code,Grouping,Classification,Specialization,Definition,Notes,Display Name,Section
193200000X,Group,Multi-Specialty,,"A business group of one or more individual practitioners, who practice with different areas of specialization.",[7/1/2003: new],Multi-Specialty Group,Individual

and here’s the diff with the suggested change:

diff --git a/a.ex b/a.ex
index d9614c6..d9a3bc2 100644
--- a/a.ex
+++ b/a.ex
@@ -30,20 +30,22 @@ defmodule TaxonomyMap do
   def run(taxonomies) do
     taxonomies
     |> Enum.reduce(%{}, fn({g, c, s}, acc) ->
-      cond do
-        nil?(acc, [g]) ->
-          # The grouping is not in the map
-          # Add the grouping, classification, and specialization for this row
-          put_in(acc, [g], classification_map(c, s))
-        nil?(acc, [g, c]) ->
-          # The classification is not in the grouping
-          # Add the classification and specialization to the grouping
-          put_in(acc, [g], Map.merge(get_in(acc, [g]), classification_map(c, s)))
-        !nil?(acc, [g, c]) ->
-          # The classification and grouping both exist
-          # Add the specialization to the grouping
-          put_in(acc, [g, c], get_in(acc, [g, c]) ++ maybe_blank_list(s))
-      end
+      inject(acc, [g, c], s )
     end)
   end
+
+  def inject( acc, [group, class], spec ) do
+    case get_in(acc, [group, class] ) do
+      nil   -> inject_class( acc, [group, class], spec)
+      found -> put_in(acc, [group, class], found ++ maybe_blank_lists(spec))
+    end
+  end
+
+  def inject_class(acc, [group, class], spec ) do
+    case get_in(acc, [group] ) do
+      nil   -> put_in(acc, [group], classification_map(class, spec))
+      found -> put_in(acc, [group], Map.merge( found, classification_map(class, spec)))
+    end
+  end
+
 end
— All posts loaded —

Where Next? Top

Trending in Questions Top

RSP87
I’m working on a project that simulates the bumbl example in the programming phoenix book. It acts almost like an email client. We have a...
New
nseaSeb
Hello, I know there is an approach for handling lists that allows for optimized traversal, but I can’t recall the specific method (somet...
New
brecabral
Documentation While reading the Scoped Routes section, I noticed that the documentation currently refers to a problem without explainin...
New
kpanic
Hi everyone, I am toying with the idea of building a “match maker” for giving personal help to people that wants to start coding. I sta...
New
velrest
So my question is quite simple and i have found no conclusive answer on forum, google or AI. Should we use :erlang.float for Integer to ...
New
asweet-confluent
I recently noticed that Elixir’s Logger defaults its primary log level to :debug when no :logger, :level application configuration is pre...
New
apz
I’m new to elixir and just tried to install the elixirLS extension for VScode(ium) and it is throwing some errors that I would like help ...
New

Other Trending Topics Top

webofbits
Aludel 0.8.0 is now available on GitHub and Hex. This release adds typed_judge assertions for evaluations that need deterministic, struc...
New
GenericJam
Edit: 2026 May 15 - This post is archived. Mob is alive!! Main docs: mob v0.7.11 — Documentation A bit of explanation for the slightly c...
New
JesseHerrick
Hey, I’m Jesse and I’m the main contributor behind Dexter, a full-featured, lightning-fast Elixir LSP optimized for large codebases. It s...
New
mudasobwa
I am happy to introduce the very α version of the new programming language compiled to BEAM. Welcome Cure. It has literally three kille...
New
garrison
Hobbes is a low-level distributed database for the Elixir programming language. Hobbes provides a simple, safe, and scalable storage lay...
New
marciok
Hi there! We created Gust: A task orchestrator inspired by Airflow. For those who have never heard about Aiflow, it’s a Python-based wor...
New

We're in Beta

About us Mission Statement

Options

Thread Display Mode




Thread Preview

Skip Thread Previews