dorgan

dorgan

Now that the next Elixir version will add Code.quoted_to_algebra/2 and Code.string_to_quoted_with_comments/2, we’re able to take some source code, parse it, change it and turn it back to formatted text. There are a couple gotchas if you change the ast in certain ways: since quoted_to_algebra requires the ast and comments to be given as separate arguments, we need to reconcile the line numbers of ast nodes and comments if we want the comments to be correctly placed.

So I wrote Sourceror, an (experimental) library that provides utilities to perform manipulations of the source code. I’m still working on more docs and examples(and tests), but this is an example of a function that expands multi alias syntax(ie: Foo.{Bar, Baz}) into their own lines:
https://github.com/doorgan/formatter/blob/main/lib/examples/multi_alias.ex#L32-L50

Or a function to add a dependency to mix.exs ala npm install:
https://github.com/doorgan/formatter/blob/main/lib/examples/deps_add.ex#L34-L76

Since the new functions are only available in Elixir master, Sourceror depends on Elixir 1.13.0-dev and can only be installed via git dependency.

Showing Posts 1 to 10

a8t

a8t

Random, pretty particular question. I was curious about writing something that could alphabetize my Mix dependencies. Would this be the right tool to build that with?

dorgan

dorgan OP

Yes! What you have to be mindful of with this kind of manipulations in contrast with macros is that you need to consider how line numbers move around. You need to both reorder the dependencies, and correct the line numbers, otherwise comments may be misplaced. The reason is that Code.quoted_to_algebra requires the ast and comments as separate arguments and mixes them by their line numbers.

One way to achieve what you want is by doing this:

"""
defp deps do
  [
    {:a, "~> 1.0"},
    {:z, "~> 1.0"},
    {:g, "~> 1.0"},
    # Comment for r
    {:r, "~> 1.0"},
    {:y, "~> 1.0"},
    # Comment for :u
    {:u, "~> 1.0"},
    {:e, "~> 1.0"},
    {:s, "~> 1.0"},
    {:v, "~> 1.0"},
    {:c, "~> 1.0"},
    {:b, "~> 1.0"},
  ]
end
"""
|> Sourceror.parse_string()
|> Sourceror.postwalk(fn
  {:defp, meta, [{:deps, _, _} = fun, body]}, state ->
    [{{_, _, [:do]}, block_ast}] = body
    {:__block__, block_meta, [deps]} = block_ast

    lines = Enum.map(deps, fn {:__block__, meta, _} -> meta[:line] end)

    deps =
      Enum.sort_by(deps, fn {:__block__, _, [{{_, _, [name]}, _}]} ->
        Atom.to_string(name)
      end)

    deps =
      Enum.zip([lines, deps])
      |> Enum.map(fn {old_line, dep} ->
        {_, tuple_meta, [{left, right}]} = dep
        line_correction = old_line - tuple_meta[:line]

        tuple_meta = Sourceror.correct_lines(tuple_meta, line_correction)
        left = Macro.update_meta(left, &Sourceror.correct_lines(&1, line_correction))
        right = Macro.update_meta(right, &Sourceror.correct_lines(&1, line_correction))

        {:__block__, tuple_meta, [{left, right}]}
      end)

    quoted = {:defp, meta, [fun, [do: {:__block__, block_meta, [deps]}]]}
    state = Map.update!(state, :line_correction, & &1)
    {quoted, state}

  quoted, state ->
    {quoted, state}
end)
|> Sourceror.to_string()
|> IO.puts()

# =>
defp deps do
  [
    {:a, "~> 1.0"},
    {:b, "~> 1.0"},
    {:c, "~> 1.0"},
    {:e, "~> 1.0"},
    {:g, "~> 1.0"},
    # Comment for r
    {:r, "~> 1.0"},
    {:s, "~> 1.0"},
    # Comment for :u
    {:u, "~> 1.0"},
    {:v, "~> 1.0"},
    {:y, "~> 1.0"},
    {:z, "~> 1.0"}
  ]
end

The other thing to note is that this is not your regular AST, Sourceror uses the literal_encoder: &{:ok, {:__block__, &2, [&1]}} option for Code.string_to_quoted_with_comments/2 under the hood, so you need to expect literals to be wrapped in blocks, so for example {:a, "~> 1.0"} will become:

{:__block__, [line: 1], [
  {{:__block__, [line: 1], [:a]},
   {:__block__, [line: 1, delimiter: "\""], ["~> 1.0"]}}
]}

This is explained a bit in the Formatting considerations section of the new functions :slight_smile:

Of course I will try to expand Sourceror as we find more complex use cases that could be simplified :slight_smile:

AndyL

AndyL

@dorgan this is exciting work - thank you!

Can you talk a little bit about who is the target audience, and your vision for possible use-cases?

Could this be used by something like elixir-ls to implement refactoring operations? (eg rename variable, rename function, extract function, inline function, rename module)

What types of contributions and testing would be most useful to you?

dorgan

dorgan OP

The target audience is primarily tool authors, like elixir-ls or credo.

Yes, those are the kind of use cases I had in mind :slight_smile: The Sourceror.to_string/2 function has an option to set the indentation level of the resulting code for that particular use case. I will probably add functions to know how many lines an ast node uses, so one could replace a line range instead of the whole file.

This started while exploring ways to allow credo to autofix some of the issues it finds, the multi alias expansion example derived from that.

Mostly finding what people find most cumbersome or confusing to do, I think the most important thing right now is to start experimenting. There may be some bugs in Code.quoted_to_algebra/2 too, some experiments in that front would be nice as well so we can add more regression tests to core Elixir :slight_smile:

dorgan

dorgan OP

Sourceror is now available on hex.pm and supports Elixir versions down to 1.10 :slight_smile:

AndyL

AndyL

@dorgan - thanks for the support down to 1.10!

Here’s a question about Sourceror and Elixir types and doctests…

I expect that a Sourceror transformation could modify a function signature or a return type…

Would Sourceror also have the ability to transform typespecs? Or inline documentation? (eg for doctests)

dorgan

dorgan OP

It should be possible, it’s information you have in the AST
For example:

iex(4)> Sourceror.parse_string(~S"""
...(4)> @spec foo(String.t()) :: :ok
...(4)> def foo(a \\ 5), do: a + 10
...(4)> """)
{:__block__, [trailing_comments: [], leading_comments: []],
 [
   {:@,
    [
      trailing_comments: [],
      leading_comments: [],
      end_of_expression: [newlines: 1, line: 1],
      line: 1
    ],
    [
      {:spec, [trailing_comments: [], leading_comments: [], line: 1],
       [
         {:"::", [trailing_comments: [], leading_comments: [], line: 1],
          [
            {:foo,
             [
               trailing_comments: [],
               leading_comments: [],
               closing: [line: 1],
               line: 1
             ],
             [
               {{:., [trailing_comments: [], leading_comments: [], line: 1],
                 [
                   {:__aliases__,
                    [trailing_comments: [], leading_comments: [], line: 1],
                    [:String]},
                   :t
                 ]},
                [
                  trailing_comments: [],
                  leading_comments: [],
                  closing: [line: 1],
                  line: 1
                ], []}
             ]},
            {:__block__, [trailing_comments: [], leading_comments: [], line: 1],
             [:ok]}
          ]}
       ]}
    ]},
   {:def, [trailing_comments: [], leading_comments: [], line: 2],
    [
      {:foo,
       [
         trailing_comments: [],
         leading_comments: [],
         closing: [line: 2],
         line: 2
       ],
       [
         {:\\, [trailing_comments: [], leading_comments: [], line: 2],
          [
            {:a, [trailing_comments: [], leading_comments: [], line: 2], nil},
            {:__block__,
             [trailing_comments: [], leading_comments: [], token: "5", line: 2],
             [5]}
          ]}
       ]},
      [
        {{:__block__,
          [
            trailing_comments: [],
            leading_comments: [],
            format: :keyword,
            line: 2
          ], [:do]},
         {:+, [trailing_comments: [], leading_comments: [], line: 2],
          [
            {:a, [trailing_comments: [], leading_comments: [], line: 2], nil},
            {:__block__,
             [trailing_comments: [], leading_comments: [], token: "10", line: 2],
             '\n'}
          ]}}
      ]
    ]}
 ]}

Typespecs are just module attributes, same for doc annotations, so you should be able to get the module attributes associated to a function by walking the tree with Macro.postwalk or Sourceror.postwalk if you’re also doing some transformation. You need to make some assumptions, though, for instance, considering a module attribute that comes before a function definition as being an annotation for such function. For doctests, since you already have access to the docstrings(they’re module attributes), it would be a matter of parsing the docstring looking for any doctest and then run Sourceror functions on it.

I think the real difficulty comes from the fact that you need to change a node that is not a children of the function, and that comes before the function. My first thought is that by postwalking, if you go up and find a function, and want to update it’s typespec, there could be a way to tell the postwalker that an update should be performed in a sibling node, maybe passing a callback and making it reduce over the parent’s children(I do something similar in the multi-alias expansion example). All of this while also applying line corrections.

I believe it is possible because we have all the data we need, but it’s something I need to put some thought on to be able to do it in a reliable and relatively straightforward way, and it would definitely be something worth adding to the library :slight_smile:

dorgan

dorgan OP

Sourceror v0.4.0 was released.

It fixes a couple bugs, makes several improvements to comments line corrections, and adds more functions to work with ranges and positions, like Sourceror.get_range/1 and Sourceror.compare_positions/2.

I also converted the multi alias expansion example into a proper document you can import to Livebook, with step-by-step explanations of the process: sourceror/notebooks/expand_multi_alias.livemd at main · doorgan/sourceror · GitHub

You can see the full changelog here.

egze

egze

Was thinking in the same direction, but writing a Credo check. A script to auto-sort the deps would be nice (although my editor currently does it also without issues)

dorgan

dorgan OP

Sourceror 0.6.0 is out :slight_smile:

This release introduces some breaking changes, as the way comments are handled by the library has been fundamentally changed. In essence, instead of requiring the user to calculate how line numbers should be shifted, Sourceror tries to “fix” the line numbers in a way that makes sense for the Elixir formatter when you call Sourceror.to_string/2 or Sourceror.extract_comments/2.

To illustrate this, the dependency sorting example can now be reduced to this traversal:

Macro.postwalk(fn
  {:defp, meta, [{:deps, _, _} = fun, body]} ->
    [{{_, _, [:do]}, block_ast}] = body
    {:__block__, block_meta, [deps]} = block_ast

    deps =
      Enum.sort_by(deps, fn {:__block__, _, [{{_, _, [name]}, _}]} ->
        Atom.to_string(name)
      end)

    {:defp, meta, [fun, [do: {:__block__, block_meta, [deps]}]]}

  quoted ->
    quoted
end)

Note that, because the line number correction hack is no longer needed, traversals over Sourceror’s AST is the same as traversals over regular AST, just don’t discard comments metadata :slight_smile:

The multi alias expansion livebook was further simplified thanks to this change.

Also, now Sourceror.get_range/1 returns the actual start and end positions for any node(provided it has line and column metadata), making it suitable for things like “replace the contents between these two positions with this new content”. Column offsets are counted as UTF-8 offsets(like the regular Elixir AST), so tools that want to support the Language Server Protocol need to convert them to UTF-16 offsets.

Changelog:

1. Enhancements

  • [Sourceror] - to_string no longer requires line number corrections to produce properly formatted code.
  • [Sourceror] - Added prewalk/2 and prewalk/3.
  • [Sourceror] - parse_string won’t warn on unnecesary quotes.
  • [Sourceror.TraversalState] - Sourceror.PostwalkState was renamed to Sourceror.TraversalState to make it more generic for other kinds of traversals.

2. Removals

  • [Sourceror] - get_line_span was removed in favor of using get_range and calculating the difference from the range start and end lines.
  • [Sourceror.TraversalState] - line_correction field was removed as it is no longer needed.

3. Bug fixes

  • [Sourceror] - get_range now properly returns ranges that map a node to it’s actual start and end positions in the original source code.

Where Next? Top

Trending in Announcing Top

wojtekmach
Hey everyone! Req is an HTTP client for Elixir that I’ve been working on for quite some time. There is already a lot of HTTP clients out...
New
handnot2
Samly can be used to enable SAML 2.0 Single Sign On in a Plug/Phoenix application. This library uses Erlang esaml to provide plug enabl...
New
woylie
Flop is an Elixir library that applies filtering, ordering and pagination parameters to your Ecto queries. offset-based pagination with...
New
MRdotB
I needed to reuse React components from my Chrome extension in my Phoenix/LiveView backend. I noticed that for Svelte/Vue, there are live...
New
garrison
Hobbes is a low-level distributed database for the Elixir programming language. Hobbes provides a simple, safe, and scalable storage lay...
New
marciok
Hi there! We created Gust: A task orchestrator inspired by Airflow. For those who have never heard about Aiflow, it’s a Python-based wor...
New
fuelen
Hi all! I want to present a small library which provides a mix task for generating an Entity-Relationship Diagram for Ecto schemas. You...
New

Other Trending Topics Top

mudasobwa
I am happy to introduce the very α version of the new programming language compiled to BEAM. Welcome Cure. It has literally three kille...
New
mudasobwa
I am seeing a lot of aplications of Argumentum ad Vericundiam in software discussions. They do link some piece of writing and point us to...
New
AstonJ
This showed up on my feed.. anyone heard of it? Just hype? Ox Alpha is a reasoning model designed for coding, sustained ag...
New
sergio
It’s not that it’s vocabulary is too advanced. It’s something worse. I get lost trying to follow even a paragraph written by Claude. It’...
New
sorenone
Today we’re releasing Oban for Python. Not an Oban client in Python. Not a pythonx wrapper embedded in Elixir. Nope, it’s a fully operati...
New
akoutmos
@hugobarauna, Dr. Dimitrios Koutmos (my brother) and I (Alex Koutmos) have been hard at work on writing a book on how you can use Elixir ...
New

We're in Beta

About us Mission Statement

Options

Thread Display Mode




Thread Preview

Skip Thread Previews