billylanchantin
Explorer - Series (1D) and dataframes (2D) for fast and elegant data exploration in Elixir
The Explorer library has been out for some time. But we just released the latest version (see below), and we thought we’d start posting updates to the forum.
Wait, Explorer is new to me. What are Series and DataFrames?
Explorer is a DataFrame library for Elixir.
DataFrame libraries are common in languages which have a focus on data manipulation, including:
If you’d like a more in-depth tutorial, there’s an excellent LiveBook called Ten Minutes to Explorer that you can play with:
But we’ll provide a quick overview here.
Briefly, you can think of a DataFrame like an in-memory table. Its purpose is to facilitate common data exploration and analysis tasks. As such, it’s a column-oriented table.
Column-oriented tables
If you’re unfamiliar with column-oriented tables, suppose you have a table of pet data like this:
| type | age | color |
|---|---|---|
| cat | 5 | black |
| dog | 2 | brown |
| dog | 3 | brindle |
A row-oriented organization of that data might look like this in Elixir:
rows = [
[type: "cat", age: 5, color: "black"],
[type: "dog", age: 2, color: "brown"],
[type: "dog", age: 3, color: "brindle"],
]
It matches the original table fairly one-to-one. But the column-oriented version might instead look like:
columns = [
type: ["cat", "dog", "dog"],
age: [5, 2, 3],
color: ["black", "brown", "brindle"]
]
It has same information, but “transposed”.
Column-orientation is beneficial if you’re asking questions that require a lot of number-crunching like “What’s the average age of all pets?”. In the row-oriented version, finding the average age would require first looking through the entire contents of the table to collect the relevant data. But in the column-oriented version, those values have already been co-located in memory.
Series and DataFrames: columns and tables
In dataframe parlance, a “series” is a single column and a “dataframe” is a collection of named series, aka a table.
Our example above would look like this:
type = Explorer.Series.from_list(["cat", "dog", "dog"])
age = Explorer.Series.from_list([5, 2, 3])
color = Explorer.Series.from_list(["black", "brown", "brindle"])
df = Explorer.DataFrame.new(type: type, age: age, color: color)
# #Explorer.DataFrame<
# Polars[3 x 3]
# type string ["cat", "dog", "dog"]
# age s64 [5, 2, 3]
# color string ["black", "brown", "brindle"]
# >
Some things to note:
- Each series has a corresponding data type or “dtype”, e.g.
typehas the dtypestring. - The word “Polars” appears. That indicates that this dataframe is using the backend powered by the fantastic Polars library (the default backend).
And if we really did want to know the average age of the pets, that would look like this:
Explorer.Series.mean(df["age"])
# 3.3333333333333335
Features and design
Preiminaries out of the way, here are Explorer’s high-level features:
-
Simply typed series:
:binary,:boolean,:category,:date,:datetime,:duration, floats of 32 and 64 bits ({:f, size}), integers of 8, 16, 32 and 64 bits ({:s, size},{:u, size}),:null,:string,:time,:list, and:struct. -
A powerful but constrained and opinionated API, so you spend less time looking for the right function and more time doing data manipulation.
-
Support for CSV, Parquet, NDJSON, and Arrow IPC formats
-
Integration with external databases via ADBC and direct connection to file storages such as S3
-
Pluggable backends, providing a uniform API whether you’re working in-memory or (forthcoming) on remote databases or even Spark dataframes.
-
The first (and default) backend is based on NIF bindings to the blazing-fast polars library.
The API is heavily influenced by Tidy Data and borrows much of its design from dplyr.
The philosophy is heavily influenced by this passage from dplyr’s documentation:
By constraining your options, it helps you think about your data manipulation challenges.
It provides simple “verbs”, functions that correspond to the most common data manipulation tasks, to help you translate your thoughts into code.
It uses efficient backends, so you spend less time waiting for the computer.
The aim here isn’t to have the fastest dataframe library around (though it certainly helps that we’re building on Polars, one of the fastest).
Instead, we’re aiming to bridge the best of many worlds:
- the elegance of dplyr
- the speed of polars
- the joy of Elixir
That means you can expect the guiding principles to be ‘Elixir-ish’. For example, you won’t see the underlying data mutated, even if that’s the most efficient implementation. Explorer functions will always return a new dataframe or series.
Links:
- HexDocs: Explorer — Explorer v0.11.1
- GitHub:
https://github.com/elixir-explorer/explorer
Acknowledgements
Explorer is an extensive library and there’s much more we could say. But for now, we’d just like to thank the dozens of contributors who’ve added wonderful improvements over the years. ![]()
Trending in Announcing
Other Trending Topics
Categories:
Sub Categories:
Forums
Popular Tags
- #ecto
- #liveview
- #troubleshooting
- #learning-elixir
- #deployment
- #library
- #erlang
- #testing
- #genserver
- #mix
- #absinthe
- #remote-other
- #otp
- #plug
- #how-to-question
- #macros
- #postgres
- #channels
- #elixirconf
- #exunit
- #discussion
- #code-sync
- #javascript
- #podcasts
- #onsite
- #dialyzer
- #docker
- #authentication
- #umbrella
- #full-time-contract
- #podcasts-by-brainlid
- #ecto-query
- #elixir-ls
- #phoenix_html
- #iex
- #blog-post
- #graphql
- #genstage
- #ai
- #elixirconf-us
- #websockets
- #supervisor
- #advent-of-code
- #distillery
- #processes
- #api
- #forms
- #metaprogramming
- #performance
- #security










First 10 of 16 Posts
billylanchantin
Explorer - version 0.8
Explorer has released version 0.8!
Added
Add
explode/2toExplorer.DataFrame. This function is useful to expand the contents of a{:list, inner_dtype}series into a “inner_dtype” series.Add the new series functions
all?/1andany?/1, to work with boolean series.Add support for the “struct” dtype. This new dtype represents the struct dtype from Polars/Arrow.
Add
map/2andmap_with/2to theExplorer.Seriesmodule.This change enables the usage of the
Explore.Queryfeatures in a series.Add
sort_by/2andsort_with/2to theExplorer.Seriesmodule.This change enables the usage of the lazy computations and the
Explorer.Querymodule.Add
unnest/2toExplorer.DataFrame. It works by taking the fields of a “struct” - the new dtype - and transform them into columns.Add pairwise correlation -
Explorer.DataFrame.correlation/2- to calculate the correlation between numeric columns inside a data frame.Add pairwise covariance -
Explorer.DataFrame.covariance/2- to calculate the covariance between numeric columns inside a data frame.Add support for more integer dtypes. This change introduces new signed and unsigned integer dtypes:
{:s, 8},{:s, 16},{:s, 32}{:u, 8},{:u, 16},{:u, 32},{:u, 64}.The existing
:integerdtype is now represented as{:s, 64}, and it’s still the default dtype for integers. But series and data frames can now work with the new dtypes. Short names for these new dtypes can be used in functions likeExplorer.Series.from_list/2. For example,{:u, 32}can be represented with the atom:u32.This may bring more interoperability with Nx, and with Arrow related things, like ADBC and Parquet.
Add
ewm_standard_deviation/2andewm_variance/2toExplorer.Series.They calculate the “exponentially weighted moving” variance and standard deviation.
Add support for
:skip_rows_after_headeroption for the CSV reader functions.Support
{:list, numeric_dtype}forExplorer.Series.frequencies/1.Support pins in
cond, inside the context ofExplorer.Query.Introduce the
:nulldtype. This is a special dtype from Polars and Apache Arrow to represent “all null” series.Add
Explorer.DataFrame.transpose/2to transpose a data frame.Changed
Rename the functions related to sorting/arranging of the
Explorer.DataFrame.Now
arrange_withis namedsort_with, andarrangeissort_by.The
sort_by/3is a macro and it is going to work using theExplorer.Querymodule. On the other side, thesort_with/2uses a callback function.Remove unnecessary casts to
{:s, 64}now that we support more integer dtypes.It affects some functions, like the following in the
Explorer.Seriesmodule:argsortcountrankday_of_week,day_of_year,week_of_year,month,year,hour,minute,secondabscliplengthsslicen_distinctfrequenciesAnd also some functions from the
Explorer.DataFramemodule:mutate- mostly because of series changessummarise- mostly because of series changessliceFixed
Fix inspection of series and data frames between nodes.
Fix cast of
:stringseries to{:datetime, any()}Fix mismatched types in
Explorer.Series.pow/2, making it more consistent.Normalize sorting options.
Fix functions with dtype mismatching the result from Polars.
This fix is affecting the following functions:
quantile/2in the context of a lazy seriesmode/1inside a summarisationstrftime/2in the context of a lazy seriesmutate_with/2when creating a column from aNaiveDateTimeorExplorer.Duration.Contributors
Thank you to everyone who opend up a PR:
And thank you to the first-time contributors!:
Changelogs
Full Changelog: Comparing v0.7.2...v0.8.0 · elixir-explorer/explorer · GitHub
Official Changelog: Changelog — Explorer v0.11.1
billylanchantin
[Blog] Explorer 0.8: The
dtypereleasehttps://cigrainger.com/explorer-0-8-0/
billylanchantin
[Release] Explorer - v0.9
Added
Add initial support for SQL queries.
The
Explorer.DataFrame.sql/3is a function that accepts a dataframe and a SQL query. The SQL is not validated by Explorer, so the queries will be backend dependent. Right now we have only Polars as the backend.Add support for remote series and dataframes.
Automatically transfer data between nodes for remote series and dataframes and perform distributed garbage collection.
The functions in
Explorer.DataFrameandExplorer.Serieswill automatically move operations on remote dataframes to the nodes they belong to.The
Explorer.Remotemodule provides additional conveniences for manual placement.Add FLAME integration, so we automatically track remote series and dataframes returned from
FLAMEcalls when the:track_resourcesoption is enabled.See FLAME for more.
Add
Explorer.DataFrame.transform/3that applies an Elixir function to each row. This function is similar toExplorer.Series.transform/2, and as such, it’s considered an expensive operation. So it’s recommended only if there is no similar dataframe or series operation available.Improve performance of
Explorer.Series.from_list/2for most of the cases where the:dtypeoption is given. This is specially true for when the dtype is:binary.Changed
Stop inference of dtypes if the
:dtypeoption is given by the user.The main goal of this change is to improve performance. We are now delegating the job of decoding the terms as the given
:dtypeto the backend.Explorer.Series.pow/2no longer casts to float when the exponent is a signed integer. We are following the way Polars works now, which is to try to execute the operation or raise an exception in case the exponent is negative.Explorer.Series.pivot_wider/4no longer includes thenames_fromcolumn name in the new columns whenvalues_fromis a list of columns. This is more consistent with its behaviour whenvalues_fromis a single column.Explorer.Series.substring/3no longer cycles to the end of the string if the negative offset surpasses the beginning of that string. In that case, an empty string is returned.The
Explorer.Series.ewm_*functions no longer replacenilvalues with the value at the previous index. They now propogatenilvalues through to the result series.Saving a dataframe as a Parquet file to S3 services no longer works when streaming is enabled. This is temporary due to a bug in Polars. An exception should be raised instead.
Contributors
And thank you to the first-time contributors!:
Changelogs
billylanchantin
[Release] Explorer - v0.10
Added
Add support for the decimals data type.
Decimals dtypes are represented by the
{:decimal, precision, scale}tuple,where precision can be a positive integer from 0 to 38, and is the maximum number
of digits that can be represented by the decimal. The scale is the number of
digits after the decimal point.
With this addition, we also added the
:decimalpackage as a new dependency.The
Explorer.Series.from_list/2function accepts decimal numbers from thatpackage as values *
%Decimal{}.This version has a small number of operations, but is a good foundation.
Allow the usage of queries and lazy series outside callbacks and macros.
This is an improvement to functions that were originally designed to accept callbacks.
With this change you can now reuse lazy series across different “queries”.
See the
Explorer.Querydocs for details.The affected functions are:
Explorer.DataFrame.filter_with/2Explorer.DataFrame.mutate_with/2Explorer.DataFrame.sort_with/2Explorer.DataFrame.summarise_with/2Allow accessing the dataframe inside query.
Add “lazy read” support for Parquet and NDJSON from HTTP(s).
Expose more options for
Explorer.Series.cut/3andExplorer.Series.qcut/3.These options were available in Polars, but not in our APIs.
Fixed
Fix creation of series where a
nilvalue inside a list * for a{:list, any()}dtype -could result in an incompatible dtype. This fix will prevent panics for list of lists with
nilentries.Fix
Explorer.DataFrame.dump_ndjson/2when date time is in use.Fix
Explorer.Series.product/1for lazy series.Accept
%FSS.HTTP.Entry{}structs in functions likeExplorer.DataFrame.from_parquet/2.Fix encode of binaries to terms from series of the
{:struct, any()}dtype.In case the inner fields of the struct had any binary (
:binarydtype), it wascausing a panic.
Changed
Explorer.Series.cut/3andExplorer.Series.qcut/3to not have “break points” column in the resultant dataframe.
So the
:include_breaksis nowfalseby default.Contributors
New Contributors
Changelogs
tmbb
I have a question regarding the development of new “primitive” functions.
I want to write functions that sample vectors (series?) of values from probability distributions. I’m most interested in the Beta and Dirichlet distributions.
I’ve thought of a format that handles both univariate (e.g. Beta) and multivariate (e.g. Dirichlet) distributions.
For univariate distributions it would be something like this:
For multivariate distributions it would be something like this:
Using a dataframe instead of a series allows me to have more or less uniform encoding for both univarite and multivariate distributions that plays well with the Explorer functions.
In order to implement this efficiently, I am building the dataframe fully in rust using a rust function and binding it to the Native module.
My questions are:
Is there any way of implementing this outside of Explorer? I believe there isn’t, because the Rust types are defined inside Explorer’s source tree
Is there any way the base types could be split into a rust crate so that they could be used in different elixir packages? I understand it’s probably more work than it’s worth
Would Explorer accept the inclusion of these functions as part of a larger group of statistical functions (random variable sampling, CDF and PDF evaluation, etc.)
If so, should those functions live in the
Explorer.DataFramemodule (because they generate dataframes) or in a newExplorer.Distributionmodule or something (likeExplorer.MathorExplorer.Statistics)Some final thoughts: I wonder if dataframes are the right API for this, but the truth is that Elixir doesn’t have bindings to a good array library, and even if it did, representing multidimensional arrays as a columns in a dataframe indexed by multiple columns doesn’t seem that terrible, and it might lead to an API which is actually better than some messy things I’ve had to do with NumPy arrays
josevalim
I wonder if you could implement these functions in Nx? This means you get to implement them in Elixir and they should be highly efficient as well.
Although we discussed ideas for exposing an Elixir API that would allow you to pass user defined functions, implemented in Rust, to Polars, so you can extend them.
tmbb
I’ve just checked and Nx does provide a function to generate uniform random number between 0 and 1, so yeah, I could probably sample from whatever distributions I wanted. However I have never actually been able to make Nx work with the EXLA compiler (or with any other compiler backend, IIRC). I didn’t actually try that hard, though.
I can try to make it work again, I guess…
How would that work? Would you use the C-ABI to interface between dynamically linked rust libraries from different Elixir packages?
billylanchantin
Polars has a plugin feature. IIRC the plan was to try and piggyback off that. But we’d not worked out the details.
I’ll also echo that this sounds more like a case for
Nx.Tensors. I would start by trying to implementdraw_from_beta,draw_from_dirchlet, etc. asdefns as that seems like the more natural setting to me.tmbb
Hm… Maybe, I don’t know… Drawing random numbers and putting them in a dataframes for further analysis (let’s say those numbers are posterior samples of a parameter in the context of a Bayesian analysis) seems like quite a natural thing to do.
But I’ll give Nx a try, anyway
tmbb
I still can’t get the EXLA compiler to run on my system (64bit WSL running on Windows), so I went ahead and implemented the random distributions in a fork. I’ll polish it a bit and make a case for inclusion in Explorer.
I think Explorer should also support KDE for random samples, with decent bandwidth estimations. I have successfully implemented KDEs using the Silverman’s rule for bandwidth estimation in Elixir (on top of Explorer, of course), and it’s quite performant, but it’s important to support better methods like the improved Sheather-Jones method, which definitely needs to be implemented in “raw” rust with access to sophisticated rust packages.
Having good support for KDEs (or for binning in histograms) is important for plotting random distributions.