1,918 karma · joined May 13, 2015
close tab
Models are excellent at code but they are not people managers. My guess this will ultimately lead a bunch of companies to revert to average returns or fold.
Cal Newport was right.
I guess that's fine, but after awhile I get a spidey-sense reading something that feels like a Claude session.
FWIW I think if you are just doing pure analytics and nothing else, Parquet will probably continue to do the job for you just fine, and you don't need to touch your workloads at all.
These new formats I think will find a niche where people aren't just running Spark jobs, but doing lots of systems building over large tables. If you're building a PB-scale data warehouse, you care a lot about the file format b/c it is a big factor in your performance curve, and you're willing to ship new experimental codecs in response to new datatypes you want to support that the system wasn't originally designed for, or you want to use a newly invented compressor.
This is really more of an expectation that has been put on file formats by the query engines. Spark/Datafusion/DuckDB wouldn't really know what to do with a multi-table file.
> Parquet is unfortunately very good just by virtue of being first, and so widely supported
IMO that is not how technology works. It is great that Parquet is so good at a lot of things, but that does not mean just because it came first that it deserves to be the only analytic file format forever.
> Its main result seems to be improved random access which, although certainly welcome, is not the point of columnar storage, as columnar storage was invented to exchange random access for something else: fast analytics
Fast analytics, as well as newer ML-shaped workloads, are inherently mix of batch scans and random access.
Some of the authors of F3 previously authored another paper that goes into the details of the shortcomings of Parquet
https://www.vldb.org/pvldb/vol17/p148-zeng.pdf
All of the newer formats that popped up recently (Vortex, Lance, F3 now) have been working on solving the problems outlined in that paper.
Lance has some interesting ideas, Vortex focuses on extensibility and performance by replacing all of Parquet's black-box encoders with fully transparent encodings. This solves the tradeoff between bulk and element decoding, allowing you to have efficient full scans and really fast random access.
E.g. Langchain recently rebuilt a system that used to be all Parquet files to use Vortex and saw a massive speedup, which they talk about more here: https://www.langchain.com/blog/introducing-smithdb
Disclaimer: I work on Vortex, so a lot of these questions about "what is the point of building a new format" are things that I have grappled with myself.
I recognize that this is a luxury product but I kind of laughed out loud at this testimonial. The amount of privilege you need to have to grow up and live in *Arizona* without ever learning how to drive is insane.
TMK there is no current government order to eliminate large swaths of Lebanon from maps. So the fact that Apple is doing this (seemingly on its own, despite all other mapping services reflecting the original place names) is the thing I'm explicitly calling out as being weird.
If you look at the Apple Maps satellite layer, you see thousands of structures spread across the area.
It is a reasonable assumption that these population centers were labeled and Apple (or one of its data partners) has withdrawn the labels.
Bing: https://www.bing.com/maps?cp=33.185932%7E35.321974&lvl=11.9&...
Google: https://www.google.com/maps/@33.1649913,35.2506666,11.55z
OSM: https://www.openstreetbrowser.org/#map=11/33.1554/35.2890
I'm not sure why they would do this for US users unless the US government requested it.
Very cool demo though!
Just vice signaling all the way down.
I read this blog post and to help wrap my head around it I put together a simple TCP-based KV store with group commit, helped make it click for me.
Or it will get figured out in the niche fields where people are willing to figure out really hard stuff to squeeze out max performance (PE, hedge funds, intelligence)
Either way agree, it's hard to get mass adoption without the software ecosystem feeding back in
It's only after hours of scouring my EOBs and being on the phone with my insurance that I then come back to the practice's office with evidence in hand, and they dismiss the charges.
I'm pretty sure this is just a racket because they expect most people not to put up a fight and just pay, or get sent to collections hell.
The amount of work you need to do as a patient in our health system is so dumb.
> brew install micromamba
> mamba install qgis
It's really crazy the number of open geospatial data feeds that exist out there from NASA, NOAA, and ESA. If you're interested in checking any of this stuff out, I highly encourage following Mark Litwinchik's blog, this guy is a legend and he does most of his work with open tools like QGIS and DuckDB
The motivation is to move the IP and trademark into a separate organization so it's no longer owned by Spiral. This means we can't re-license it later, we'd have to fork it, because the Vortex trademark and all that is controlled by LF.
Basically in the US you need a legally recognized entity to hold intellectual property. "Donating" the project involves setting up a "Series LLC" that is nested underneath the top-level Linux Foundation corporation, and donating the IP into it.
Checkout https://docs.linuxfoundation.org/lfx/project-control-center/... and ctrl-f "LF Projects, LLC"