tajihiro

tajihiro

I have data following structure with csv file.

| id | name | lang |
| 1 | Bob  | Elixir |
| 1 | Bob  | Java |
| 1 | Bob  | Ruby |
| 2 | Tom  | Elixir |
| 2 | Tom  | C# |
| 3 | Dave | Python |
| 3 | Dave | Java |
| 3 | Dave | Ruby |
| 4 | Jane | Python |

I would like to group by name and make vertical lang data to horizontal like below.

| name | lang1 | lang2 | lang3 |
| Bob | Elixir | Java| Ruby |
| Tom | Elixir | C# |  |  
| Dave | Python | Java | Ruby |
| Jane | Python |  |  |

How can I make it by best way?

Most Liked

LostKobrakai

LostKobrakai

Parse the csv using e.g. nimble_csv, group with Enum.group_by + Enum.map to make a single row for the grouped values and encode to csv again.

LostKobrakai

LostKobrakai

CSV.decode returns a list of {:ok, data} tuples, so you need adjust your code to extract the data out of those tuples.

tajihiro

tajihiro OP

Finally I ended up with this.

defmodule FormatTsv do
  def convert do
    headers = [:id, :name, :lang]
    out_file = "out.tsv"
    result = "data.tsv"
      |> File.stream!
      |> CSV.decode(separator: ?\t, headers: headers)
      |> Enum.map(&(elem(&1, 1)))
      |> Enum.group_by(fn(x)-> x[:id] end)
      |> Enum.map(fn {id, values} -> 
                    "#{id}\t#{Enum.at(values, 0)[:name]}\t#{Enum.at(values, 0)[:lang]}\t#{Enum.at(values, 1)[:lang]}\t#{Enum.at(values, 2)[:lang]}\n"
                  end)
    File.write(out_file, result)
  end
end

Thanks.

Last Post!

ericgray

ericgray

Saw this the other day and just getting some time to respond. The structure of the data makes it a little bit of a pain to parse. I wrote some code that works assuming the ids in the tsv file are ordered. If your looking for this type of result sometimes it’s better to transform the data into another format. For this purpose I thought a struct would be the appropriate data structure to use.

defstruct [headers: [], data_rows: [], column_count: 0]

headers is a list that will accumulate all the headers.
data_rows will contain a list of maps for each row in the file.
column_count will keep track of how many header columns are needed.

So thinking in terms of rows of data the desired result is to take this input

| id | name | lang |
| 1 | Bob  | Elixir |
| 1 | Bob  | Java |
| 1 | Bob  | Ruby |
| 2 | Tom  | Elixir |
| 2 | Tom  | C# |
| 3 | Dave | Python |
| 3 | Dave | Java |
| 3 | Dave | Ruby |
| 4 | Jane | Python |

and reduce them to single rows of maps

%{id: 1, name: "Bob", languages: ["Elixir", "Java", "Ruby"]},
%{id: 2, name: "Tom, "languages: ["C#", "Elixir"]},
%{id: 3, name: "Dave,"languages: ["Ruby", "Java", "Python"]},
%{id: 4, name: "Jane", languages: ["Python"]}

This is what I came up with hope it helps.

defmodule FormatTsv do
  defstruct [headers: [], data_rows: [], column_count: 0]
  NimbleCSV.define(TSV, separator: " | ")
  @file_input "priv/data.tsv"
  @file_output "priv/output.tsv"

  def convert do
    @file_input
    |> File.stream!()
    |> TSV.parse_stream()
    |> Stream.map(fn [id, name, lang] -> new_data_row(id, name, lang) end)
    |> transform_data()
    |> write_to_file()
  end

  def transform_data(data_rows) do
    data_rows
    |> group_data_rows()
    |> add_last_row()
    |> add_headers()
    |> reverse_rows()
  end

  def add_last_row({row, struct}) do
    Map.put(struct, :data_rows, [row|struct.data_rows])
  end

  def add_headers(struct) do
    Map.put(struct, :headers, headers(struct.column_count))
  end

  def reverse_rows(struct) do
    Map.put(struct, :data_rows, Enum.reverse(struct.data_rows))
  end

  def group_data_rows(rows) do
    Enum.reduce(rows, {%{}, %FormatTsv{}}, fn(row, acc) ->
      {map, struct} = acc
      group_data(row, map, struct)
    end)
  end

  def group_data(row, map, struct) when map == %{} do
    {row, struct}
  end

  def group_data(row, map, struct) do
    case row.id == map.id do
      true ->
        [language] = row.languages
        {add_language(map, language), struct}

      false ->
        {row, add_data_row(struct, map)}
    end
  end

  def add_language(map, language) do
    Map.put(map, :languages, [language|map.languages])
  end

  def add_data_row(struct, map) do
    Map.put(struct, :data_rows, [map|struct.data_rows])
    |> update_column_count(map)
  end

  def update_column_count(struct, map) do
    header_count = Enum.count(map.languages)
    column_count = struct.column_count

    case header_count > column_count do
      true ->
        Map.put(struct, :column_count, header_count)

      false ->
        struct
    end
  end

  def headers(num) do
    ["id", "name"] ++ Enum.map(1..num, &("lang#{&1}"))
  end

  def header_row(struct) do
    "| " <> Enum.map_join(struct.headers, " | ", &(&1)) <> " |\n"
  end

  def data_row_to_string(row, lang_count, column_count) when lang_count < column_count do
    "| #{row.id} | #{row.name} |" <> " " <> Enum.join(row.languages, " | ") <> String.duplicate(" |", column_count - lang_count) <> " |\n"
  end

  def data_row_to_string(row, _lang_count, _column_count) do
    "| #{row.id} | #{row.name} |" <> " " <> Enum.join(row.languages, " | ") <> " |\n"
  end

  def write_to_file(struct) do
    File.write(@file_output, header_row(struct), [:append])

    for row <- struct.data_rows do
      lang_count = Enum.count(row.languages)
      column_count = struct.column_count
      data_row = data_row_to_string(row, lang_count, column_count)
      File.write(@file_output, data_row, [:append])
    end

  end

  def new_data_row(id, name, language) do
    %{}
    |> put_id(id)
    |> put_name(name)
    |> put_language(language)
  end

  def put_id(map, id) do
    id = String.trim_leading(id, "| ")
    Map.put(map, :id, id)
  end

  def put_name(map, name) do
    name = String.trim_trailing(name)
    Map.put(map, :name, name)
  end

  def put_language(map, language) do
    language = String.trim_trailing(language, " |")
    Map.put(map, :languages, [language])
  end
end

Where Next? Top

Trending in Questions Top

katta
I having some trouble figuring out if I have set myself too strict of standards for my production server. Currently I can handle 75% of r...
New
achenet
Hello, I’m trying to build a basic Phoenix web-app, and I’d like to use Tailwind. However, when I launch mix phx.server, I get an error...
New
kpanic
Hi everyone, I am toying with the idea of building a “match maker” for giving personal help to people that wants to start coding. I sta...
New
Cxx-mlr
I’m working on a small exercise involving update_in/3, and I came up with this solution: data = %{ name: "Periodic Table", category:...
New
ChrisAmelia
I’ve got trouble wrapping my head around the order in which functions are called in this snippet (from Phoenix’s authentication): toke...
New
dillonoconnor
Is there any way to avoid the Hologram compiler running when using iex? It seems like the front-end code could potentially be disregarded...
New
thiagogsr
** (ArgumentError) expected :max_attempts to be a positive integer, got: {:@, [line: 10, column: 19], [{:max_attempts, [line: 10, column:...
New

Other Trending Topics Top

GenericJam
Edit: 2026 May 15 - This post is archived. Mob is alive!! Main docs: mob v0.7.11 — Documentation A bit of explanation for the slightly c...
New
mudasobwa
I am happy to introduce the very α version of the new programming language compiled to BEAM. Welcome Cure. It has literally three kille...
New
garrison
Hobbes is a low-level distributed database for the Elixir programming language. Hobbes provides a simple, safe, and scalable storage lay...
New
budgie
A little off-topic, but I feel like people here have a good head on their shoulders. I used to be quite good at making software. Was luc...
New
KristerV
Hey. Is there anyone here who creates agents in their apps? Not talking about using agents, but creating them. I’m finding it pretty diff...
New
mcass19
ExRatatui lets you cook up rich terminal UIs in Elixir, powered by Rust’s ratatui via Rustler NIFs. Build interactive terminal applicatio...
New

We're in Beta

About us Mission Statement

Options

Thread Display Mode




Thread Preview

Skip Thread Previews