162 karma · joined May 7, 2017
I agree with some comments I've seen recently that too much content about Elixir is telling and not showing how great it is. So here's a case study in how consolidating around Elixir was a win.
[1] https://www.youtube.com/watch?v=HP86Svk4hzI
[2] https://dockyard.com/blog/2024/02/06/5-benefts-amplified-saw...
We were able to completely eliminate a few python services and consolidate to all Elixir for ETL and ML.
I'm really pleased with how it turned out. Inspired by the Nx architecture, which uses pluggable backends, I built a thin user-facing API that calls into (theoretically, as polars is the only extant backend) pluggable dataframe backends. What really excited me about this approach is that it permits a similar approach to what we see in dplyr [5], where you can manipulate in-memory data frames using the same API as remote databases or spark dataframes.
Next up is to move to lazy-by-default. To be as unsurprising as possible, Explorer dataframes are (for all intents and purposes) immutable. This has required a fair amount of copying when using the eager API from Polars, and because Rustler NIF resources use atomic reference counting and the GC only sweeps intermittently, there can be some pretty bad memory performance. Fortunately Polars also has a lazy API. The plan is to use that with 'peeking' for display.
After that, I'd like to move into additional backends. I'm particularly keen on Ecto (database) and Apache Arrow/Ballista for distributed and OLAP work. There is also work underway for a pure Elixir backend so the library can ship without a Rust dependency. Speaking of which, there's work on prebuilt binaries underway as well.
I'd love feedback on the API! I aimed for a dplyr-ish API as I think it melds better with a functional language than pandas. Generally I find dplyr more intuitive than pandas. The philosophy here is to get from brain to data as simply and intuitively as possible.
Finally, contributions and any other feedback are super, super welcome. It's early days and I'm also a startup founder so I haven't been able to dedicate as much time as I'd like, but I try to get some work done and add features at least once a week.
Thanks for looking!
--
[1] https://github.com/elixir-nx/nx
[2] https://github.com/elixir-nx/axon
[3] https://github.com/pola-rs/polars
As a point of shameless self-promotion, I'm the founder of a company seeking to improve the patent system, starting with prior art search [1][2]. We're not one of those companies that will do it for you for less. Instead, we empower you to do it yourself, then have your attorney go deeper if needed. The idea being that you don't need to be a patent expert or lawyer to do it; anyone with knowledge of their technical area can search effectively. Plus we've got some collaboration features so when you're ready to hand off to your $500/hr attorney, you can. And not just hand off -- collaborate.
If a patent troll is shaking you down, reach out to us and we'll do you a generous deal, no strings attached. We very much do not like patent trolls.
[2] https://www.amplified.ai/en/blog/22062020/introducing-amplif...
A couple of tech points: we're using Elixir and the app is entirely Phoenix LiveView. Building with LiveView has been such a pleasure. I'll write another post about taking the bet on LiveView, but it's been a great one for our small team.
On the machine learning side, everything is bespoke. My PhD work was on innovation economics and integrating machine learning into econometric models. We jumped into transformers early and we've been working on advances in handling long sequences that we'll be submitting to conferences over the next year.
We really want to make the patent system work the way it's meant to work. That means reducing the number of bad patents that get granted and making it easier for anyone to grapple with patent prior art. We're also keen to help out anyone being harassed by patent trolls. If you are, drop us a line and we'd be happy to get you set up with access to Amplified for free.
I'm a cofounder of amplified ai, a distributed global startup where we want to change how the world innovates. We're starting with prior art search, and we're looking to expand our small team with a data engineer. We're looking for someone creative and talented, who can readily move across different data sources and choose the best tool for the job. We love functional programming, and use Clojure for a lot of data work. We also use Golang extensively for back end services. We're looking for someone who can finish setting up and continue to improve our distributed data stack. We have a lot of data and a lot more to come.
We value diversity, humility and intense curiosity. We're fast moving and willing to try promising new approaches.
Check us out on AngelList (https://angel.co/amplified-ai/jobs/251929-data-engineer) for more details.