1,510 karma · joined August 24, 2015
[ my public key: https://keybase.io/refset; my proof: https://keybase.io/refset/sigs/pyTK0thu8g5O4Zs7EA3HrQYKkYZjeEz_TUr8nZV2hlY ] c0ba44896414421c9009bc8a77a86ceb
meet.hn/city/51.0612766,-1.3131692/Winchester
Socials: - github.com/refset
---
This is the truth. My favourite example of this is MemoryDB's "multi-AZ (multi-datacenter) durability" for Valkey - there's a good write-up here https://brooker.co.za/blog/2024/04/25/memorydb.html
LMDB springs to mind.
> LMDB was designed to resist data loss in the face of system and application crashes. Its copy-on-write approach never overwrites currently-in-use data. Avoiding overwrites means the structure on disk/storage is always valid, so application or system crashes can never leave the database in a corrupted state.
https://en.wikipedia.org/wiki/Lightning_Memory-Mapped_Databa...
To be clear my comment is specifically only relating to signature schemes, not encryption.
> The state is enormous
The scheme I linked to points towards efficient "pebbling" and "hash chain traversal" algorithms which minimize the local state required in quite a fascinating way (e.g. see https://www.win.tue.nl/~berry/pebbling/).
> tracking state across components and through distribution channels
Assuming you have reliable ordering in those channels I don't see how the stateful nature of such schemes makes it hugely more complex than the essential hard problem of key distribution.
[0] https://web.archive.org/web/20110401080052/https://www.cdc.i...
[1] https://news.ycombinator.com/item?id=33925383 I wrote about this "Dahmen-Krauß Hash-Chain Signature Scheme" (DKSS) algorithm previously in a comment a couple of years ago
[0] "FDB - a reactive database environment for your files" https://www.youtube.com/watch?v=EvAFEC6n7NI
Not entirely fair - see https://github.com/rockset/rocksdb-cloud (a fork of RocksDB with a separation of storage and compute, using S3 and Lambda-based compaction)
What counts as large or small definitely varies a lot depending on the context of the conversation/analysis.
MotherDuck's "Big Data is Dead" post [0] sticks in mind:
> The general feedback we got talking to folks in the industry was that 100 GB was the right order of magnitude for a data warehouse. This is where we focused a lot of our efforts in benchmarking.
Another point of reference is [1]
> [...] Umbra achieves unprecedentedly low query latencies. On small data sets, it is even faster than interpreter engines like DuckDB
> TPC-H Small Dataset = 866k tuples, sf 0.1
[0] https://motherduck.com/blog/big-data-is-dead/
[1] https://db.in.tum.de/~kersten/Tidy%20Tuples%20and%20Flying%2...
Incremental View Maintenance engines might be the solution we've been waiting for here.
Another example of this I saw recently using SQLite (compiled to Wasm) in the browser: https://docs.sqlitecloud.io/docs/sqlite
And if you ever want something similar for more general backend APIs (without relying on Wasm or the browser to run the software), https://codapi.org/ looks very slick e.g. as demonstrated in https://antonz.org/sql-upsert/ - discussed on HN previously [0]
Inspired by https://www.db-fiddle.com/ my colleagues ended up building a fairly bespoke setup for XTDB's docs (XTDB doesn't yet compile to Wasm) shortly before I came across Codapi, although our requirements were even more particular, e.g. see https://docs.xtdb.com/tutorials/financial-usecase/time-in-fi... - the backend here is https://github.com/xtdb/xt-fiddle which runs purely on top of Lambda Snapstart, and embedded within docs based on Astro's Starlight [1] and Web Components
> [Chandler] is inspired by a PIM from the 1980s called Lotus Agenda, notable because of its "free-form" approach to information management. Lead developer of Agenda, Mitch Kapor, was also involved in the vision and management of Chandler.
[0] https://news.ycombinator.com/item?id=39070631 / "For a moment there, Lotus Notes appeared to do everything a company needed"
It's X.T. (as in 'Cross-Time' / https://xtdb.com), but thank you! :)
> 1: https://aidanhogan.com/docs/ring-graph-wco.pdf
Oh nice, I recall skimming this team's precursor paper "Worst-Case Optimal Graph Joins in Almost No Space" (2021) - seems like they've done a lot more work since though, so definitely looking forward to reading it:
> The conference version presented the ring in terms of the Burrows–Wheeler transform. We present a new formulation of the ring in terms of stable sorting on column databases, which we hope will be more accessible to a broader audience not familiar with text indexing
> Letos join
God-Emperor Join has a nice ring to it.
[0] "Simple Adaptive Query Processing vs. Learned Query Optimizers: Observations and Analysis" - https://www.vldb.org/pvldb/vol16/p2962-zhang.pdf
> If you can’t trust, you have to pay attention in order to verify – and verifying is expensive
> We’d like to believe that goodness brings us together, but that’s not what the data reveal. According to group studies, we don’t come together because we trust: we come together because we align our intent to mistrust.
> We stick together because our interests align and we suffer the mistrust of others until we can no longer justify it.
Interesting perspectives.
Makes me wonder... what is the state of the art in software systems that could help us to "align our intent" in this "impending cyborg age"?
My assumption is that the author is painting out a longer-term vision both for Confluent and its audience.
It probably makes more sense in the context of other posts, e.g.
> This trend towards object storage is not just happening at Confluent but across the data ecosystem. Many different types of cloud data systems are integrating object storage into their architecture, and I have covered a few of these such as Neon and ClickHouse Cloud. There is also a flurry of start-ups doing logs (such as the Kafka API) over object storage directly (six and counting). There is money being plowed into data systems that use object storage as the primary, and sometimes, only storage layer.
https://jack-vanlightly.com/blog/2024/5/2/hybrid-transaction...
In other words: managed Kafka is becoming a commodity and Confluent should "start with why" to figure out where to go. I figured the advice and perspective is of general enough interest.
> [H2] provides a way to enforce usage of parameters when passing user input to the database. This is done by disabling embedded literals in SQL statements. To do this, execute the statement:
> SET ALLOW_LITERALS NONE;
> Literals can only be enabled or disabled by an administrator