The databricks people suggest that a better "cache format" is needed, but I don't see how that's different from ETLing the data into a regular warehouse.
655 karma · joined August 31, 2016
The databricks people suggest that a better "cache format" is needed, but I don't see how that's different from ETLing the data into a regular warehouse.
But one guy came back from retirement after 9/11 and moved into a trailer in the Lockheed parking lot.
But I constantly underestimate the tendency to take a minor capability and exaggerate it into a replacement for 50 years of heavily studied technology.
Even if people somehow get this to work like a real database, they will quickly become frustrated with the same features that the data lake avoids: access controls, locks, consistency, centralization, a schema, you name it.
Somewhere in my github there's an indefinite-width multiply that only adds bits while there's a risk of carry into the 1's digit; the check for that is quite cheap.
I came up with a suffix-sorting index for this domain that's interestingly simple. Most algos for this use a generalized suffix tree that's built by concatenating all the strings into one giant string and feeding it into a conventional suffix sort, but that has some big constants on the indexing throughput, due I think to the overhead of handling one giant string instead of a bunch of small independent records.
In the latter case, by making the structure slightly simpler and search slightly harder, I can get indexing throughput in the GBps, at least for the sorting part.
The output of that in its simplest form is a 4n or 8n-sized set of int's, but it can be fed further into a compressed rank/select data structure for various space/indexing time/retrieval time tradeoffs, and I don't think those are slow (eg Roaring Bitmaps)
I'll post this on show HN if anybody's interested; I'm still writing up the details, as I've barely gotten the POC code working.
The UK is looking increasingly like a third world country.
The funny part of the Kakao CEO asking to be called Brian is that there was a K-drama (Search WWW) with a fictional tech company, but they also made the CEO name Brian. I suspect if this idea had gone on every CEO would be called Brian.
My knowledge of microcontrollers is dated, but the frameworks for Arduino etc. seemed limited on their ability to do event or message-based programming; most example apps at least were a polling loop. The classic architecture of setting up interrupt/event handlers and going to sleep was not there.
https://db.cs.cmu.edu/papers/2024/whatgoesaround-sigmodrec20...
Rate limiters apparently use count-min over a fixed time interval, which is bursty, but I took the same hash structure and came up with "timestamp-min" that allows even pacing (also without "fill up"). The one-sided error is also useful for checking if a cache entry is too old.
It doesn't fix everything (eg DDOS), but it can prevent any single client from over-requesting, or stealing another client's requests (collisions can be made as hard as necessary).
https://dspace.mit.edu/bitstream/handle/1721.1/34929/MIT-CSA...
Silicon Valley doesn't have a good record in the DB/DWH space; producing a fully-featured DBMS doesn't seem to fit the VC model.
It appears that a lot of attention is now directed at the folks doing 100 MB queries, and the high end has moved past everybody's radar. My idea of an exciting product is Ocient, who have skipped over Cloud and gone for hyperscale on-prem hardware. Yellowbrick is also a contender here.
I have a lot of experience with Vertica, and they seem to have gotten stuck in this niche as well, with sales tilted towards big accounts, but less traction in smaller shops, and a difficult road to get a SaaS or similar easy-start offering.
There's a crossover point where self-managed is cheaper than cloud, but nobody seems to have any idea where it is. Snowflake will gladly tell you that your sub-$1M Vertica cluster should be replaced by $10M of sluggish SaaS, and that you are saving money by doing so. These decisions seem more in the realm of psychology or political science.
DHH's cloud exit was a refreshing take on the expense issue, even if it wasn't strictly in the database space -- the cost per VCPU and so forth that he documented is a good start for estimating savings, and he debunked a lot of the "hidden costs" that cloud maximalists claim.
In the business/financial space the biggest news to me was the correction in Snowflake's stock price, which seemed to indicate that investors were finally noticing metrics like price-performance, but they added a little more AI and went back into irrationality.
I'm heavily in favor of DuckDB, Hudi, Iceberg, S3 tables, and the like. Mixing high-end and low-end tools seems like the best strategy (although settling on one high-end DWH has also worked IME), and the low end is getting better and cheaper, squeezing out the mid-range SaaS vendors.
In research I found Goetz Graefe's work in offset-value coding exciting -- he's wired it into query operators in a way that saves a lot of CPU on sorting and joins/aggregation. This is a technique that I've applied favorably in string sorting, and it was discovered in the DB community decades ago but largely forgotten. (This work precedes 2024, but I'm a slow study.)
The linked paper is clearer:
"However, the inner loops in the join are typically accelerated with existing primary key indexes or temporary indexes built on the fly."
"Note that SQLite probes the part table index for every tuple in the lineorder table."
The Bloom filter does not reduce the cardinality of the join, it simply replaces each B-tree probe and filter with a Bloom probe on a pre-filtered key set.
This technique is well-known; the paper cites several examples, and Bloom filter pushdown is common to many commercial systems.
We had one outage where key rotation had been enabled on reboot, so data partitions were lost after what should have been a routine crash. Overall, for data warehousing, our failure rate on on-prem (DC-hosted) hardware was lower IME.
Bitstring rank/select is a well-known problem, and the BMI and non-BMI (Hacker's Delight) versions are available as a reference.
Thanks for the reference; this question has been on my mind for some time.
SELECT
pickup.zone AS pickup_zone,
dropoff.zone AS dropoff_zone,
cnt AS num_trips
FROM
(select pickup_location_id, dropoff_location_id, count(*) as cnt from taxi_data_2019 group by 1,2) data
INNER JOIN
(SELECT * FROM zone_lookups WHERE Borough = 'Manhattan') pickup
ON pickup.LocationID = data.pickup_location_id
INNER JOIN
(SELECT * FROM zone_lookups WHERE Borough = 'Manhattan') dropoff
ON dropoff.LocationID = data.dropoff_location_id
ORDER BY num_trips desc
LIMIT 5;
┌───────────────────────┬───────────────────────┬───────────┐
│ pickup_zone │ dropoff_zone │ num_trips │
│ varchar │ varchar │ int64 │
├───────────────────────┼───────────────────────┼───────────┤
│ Upper East Side South │ Upper East Side North │ 536621 │
│ Upper East Side North │ Upper East Side South │ 455954 │
│ Upper East Side North │ Upper East Side North │ 451805 │
│ Upper East Side South │ Upper East Side South │ 435054 │
│ Upper West Side South │ Upper West Side North │ 236737 │
└───────────────────────┴───────────────────────┴───────────┘
Run Time (s): real 0.304 user 1.791931 sys 0.132745
(unedited query is similar to theirs, about .9s)And, doing the same to the optimized query drops it to about .265s.