On DataFrame datatype in Ruby (2016)
zverok.github.io
zverok.github.io
Ruby, Go, Scala, Java should all build these libs (Scala has Spark for cluster computing, but it needs a good Pandas-like single node lib) to attract the data engineers. I'd imagine it's relatively easy to add DataFrame support now cause PyArrow does the heavy lifting.
The DaRu (Ruby DataFrame) sub-project's last commit it around a year ago, but the daru-view sub-project (visualization and plotting for web apps) seems to be maintained and updated frequently:
https://github.com/SciRuby/daru
https://github.com/SciRuby/daru-view
While I like Ruby's syntax and community, it is really hard to compete with in high-level scientific programming market when you are competing against Python, R & Julia (and Mathematica and MatLab).
Just to provide a bit of the personal context: this article caused creator of DaRu (Sameer Deshmukh) to contact me and propose to work on DaRu together, and so I did (see @zverok here: https://github.com/SciRuby/daru/graphs/contributors). I also was, for some time, SciRuby/DaRu's mentor for Google Summer of Code (and, IIRC, it was my initial idea that daru-view grew from).
Also, since that article, an independent dataframe library https://github.com/ankane/rover was created by Andrew Kane, handling some of API and implementation in a cleaner way.
That being said, I am not sure that DaRu, or Rover (or "dataframe" idea in general) has enough visibility in the Ruby community. It is mostly thought as "some special scientific thing", while I believe in 2021 it should be seen as one of the necessary everyday high-level datatypes.
That's what I'd focus this article on if I'd written it today.
And I add that just the datatype is not enough, because is "this close" to be the foundation for relational programming (p.d: I'm building a relational language where you can say everything is alike data-frames/relations: https://tablam.org).
> It is mostly thought as "some special scientific thing"
This is part of the problem, truly. Being so focused on "science" when I think is better to frame it as data manipulation like you do in SQL tables/views, making it much more general than is used for...
> I believe in 2021 it should be seen as one of the necessary everyday high-level datatypes.
I agree. It is one of those data structures that developers tend to use by reimplementing badly using arrays/lists and dictionarys, usually even unaware that they are doing so.
E.g. Apache Arrow.
E.g. Julia's constellation of different dataframe (Tables.jl) libraries that are mutually compatible.
E.g. efforts at standardizing dataframes in python.
E.g. tidyverse etc in R
(In _some_ defense of the post, at the time of the writing it was targeted _only_ to the Ruby community. But still, probably, if I'd included "how others do it" section it would bear more weight.)