HNHacker News
TopNewBestAskShowJobs

refset

1,510 karma · joined August 24, 2015

https://github.com/refset

[ my public key: https://keybase.io/refset; my proof: https://keybase.io/refset/sigs/pyTK0thu8g5O4Zs7EA3HrQYKkYZjeEz_TUr8nZV2hlY ] c0ba44896414421c9009bc8a77a86ceb

meet.hn/city/51.0612766,-1.3131692/Winchester

Socials: - github.com/refset

---

submissionscomments
refset··on Ask HN: What's Prolog like in 2024?
Compiling Datalog to SQL with Logica is possibly the easiest path if you need a production ready open source Datalog setup (i.e. choose your favourite managed Postgres provider): https://logica.dev/
refset··on Paper Trails
I haven't followed Arweave closely at all recently but the creator seemed genuine and capable enough to warrant attention. As to tech itself, assuming commodity storage costs continue to drop ~indefinitely and somewhat predictably it feels intuitive to me how such a scheme can be viable. It's ultimately a bet on ever-reducing storage costs.
refset··on Paper Trails
Arweave has a compelling design for long-term public archival along these lines, with an explicit concept of 'endowment' to predictably cover the costs of future storage.
refset··on A write-ahead log is not a universal part of durability
> durability is a spectrum

This is the truth. My favourite example of this is MemoryDB's "multi-AZ (multi-datacenter) durability" for Valkey - there's a good write-up here https://brooker.co.za/blog/2024/04/25/memorydb.html

refset··on A write-ahead log is not a universal part of durability
Specifically, that modern theory is "protocol-aware recovery" - e.g. https://www.usenix.org/conference/fast18/presentation/alagap...
refset··on A write-ahead log is not a universal part of durability
> I'm not sure if there are any databases that do your 'with a lot of care' option

LMDB springs to mind.

> LMDB was designed to resist data loss in the face of system and application crashes. Its copy-on-write approach never overwrites currently-in-use data. Avoiding overwrites means the structure on disk/storage is always valid, so application or system crashes can never leave the database in a corrupted state.

https://en.wikipedia.org/wiki/Lightning_Memory-Mapped_Databa...

refset··on Quantum is unimportant to post-quantum
> use symmetric crypto

To be clear my comment is specifically only relating to signature schemes, not encryption.

> The state is enormous

The scheme I linked to points towards efficient "pebbling" and "hash chain traversal" algorithms which minimize the local state required in quite a fascinating way (e.g. see https://www.win.tue.nl/~berry/pebbling/).

> tracking state across components and through distribution channels

Assuming you have reliable ordering in those channels I don't see how the stateful nature of such schemes makes it hugely more complex than the essential hard problem of key distribution.

refset··on Quantum is unimportant to post-quantum
There has been research on the intersection of IoT and PQ signatures specifically at least, e.g. see "Short hash-based signatures for wireless sensor networks" [0] [1]. Unlike SPHINCS+ which is mentioned in the article, if you're happy to keep some state around to remember the last used signature (i.e. you're not concerned about accidental re-use) then the scheme can potentially be _much_ simpler.

[0] https://web.archive.org/web/20110401080052/https://www.cdc.i...

[1] https://news.ycombinator.com/item?id=33925383 I wrote about this "Dahmen-Krauß Hash-Chain Signature Scheme" (DKSS) algorithm previously in a comment a couple of years ago

refset··on Python toolkit for quantitative finance
Beacon is more PaaS than SaaS from what I've seen, but it's all very neatly integrated and they even wrote their own compute scheduling engine. The data model is interesting: https://www.beacon.io/wp-content/uploads/2021/05/5.-WhitePap...
refset··on HyperCard Simulator
The joy of Flash was the ease of scripting alongside vector-based illustration with keyframe animations and audio (plus intuitive asset management, onion skinning etc. etc.). Flash was downright fun indeed!
refset··on Local First, Forever
Along similar lines of "just use your preferred cloud-based file-syncing solution", see: https://github.com/filipesilva/fdb - the author spoke about it recently [0]. The neat thing about this general approach is that is pushes all multi-user permissions problems to the file-syncing service, using the regular directory-level ACLs and UX.

[0] "FDB - a reactive database environment for your files" https://www.youtube.com/watch?v=EvAFEC6n7NI

refset··on OpenAI Acquires Rockset
> The technology is not open-source

Not entirely fair - see https://github.com/rockset/rocksdb-cloud (a fork of RocksDB with a separation of storage and compute, using S3 and Lambda-based compaction)

refset··on Tetris Font (2020)
The Advanced Data Structures course lectures are great, including his very own research on "Retroactive Data Structures" https://courses.csail.mit.edu/6.851/spring21/lectures/L01.ht...
refset··on When are SSDs slow?
> small dataset (100GB)

What counts as large or small definitely varies a lot depending on the context of the conversation/analysis.

MotherDuck's "Big Data is Dead" post [0] sticks in mind:

> The general feedback we got talking to folks in the industry was that 100 GB was the right order of magnitude for a data warehouse. This is where we focused a lot of our efforts in benchmarking.

Another point of reference is [1]

> [...] Umbra achieves unprecedentedly low query latencies. On small data sets, it is even faster than interpreter engines like DuckDB

> TPC-H Small Dataset = 866k tuples, sf 0.1

[0] https://motherduck.com/blog/big-data-is-dead/

[1] https://db.in.tum.de/~kersten/Tidy%20Tuples%20and%20Flying%2...

refset··on When are SSDs slow?
Slightly off-topic but CedarDB is extremely exciting. It's the commercialization of the widely cited Umbra research DBMS [0] that has been in the works for several years, which benchmarks faster than DuckDB for OLAP [1] whilst simultaneously being really strong for transactional workloads. Also discussed recently here [2].

[0] https://umbra-db.com/

[1] https://cedardb.com/blog/ode_to_postgres/

[2] https://news.ycombinator.com/item?id=40241150

refset··on Half a century of SQL
> Many constraints are extremely expensive to enforce at scale to the point of being prohibitive

Incremental View Maintenance engines might be the solution we've been waiting for here.

refset··on Half a century of SQL
Integrating live code editors within docs and tutorials is great.

Another example of this I saw recently using SQLite (compiled to Wasm) in the browser: https://docs.sqlitecloud.io/docs/sqlite

And if you ever want something similar for more general backend APIs (without relying on Wasm or the browser to run the software), https://codapi.org/ looks very slick e.g. as demonstrated in https://antonz.org/sql-upsert/ - discussed on HN previously [0]

Inspired by https://www.db-fiddle.com/ my colleagues ended up building a fairly bespoke setup for XTDB's docs (XTDB doesn't yet compile to Wasm) shortly before I came across Codapi, although our requirements were even more particular, e.g. see https://docs.xtdb.com/tutorials/financial-usecase/time-in-fi... - the backend here is https://github.com/xtdb/xt-fiddle which runs purely on top of Lambda Snapstart, and embedded within docs based on Astro's Starlight [1] and Web Components

[0] https://news.ycombinator.com/item?id=38663717

[1] https://starlight.astro.build/

refset··on DuckDB Doesn't Need Data to Be a Database
Steampipe demonstrates a rather impressive range of scenarios for using FDWs + SQL in place of more traditional ETL and API integrations: https://steampipe.io/
refset··on GTFL – A Graphical Terminal for Common Lisp
FlowStorm really looks like it delivers on "execution is data", per https://www.scattered-thoughts.net/writing/the-shape-of-data
refset··on Agenda: a personal information manager (1990) [pdf]
While trying to recall the relationship between Lotus Notes and Lotus Agenda ("which came first?" etc.) I just dug up this[0] comment from a different Notes-related HN discussion a few months ago that also mentions 'Chandler'[1] - an OSS implementation written in Python:

> [Chandler] is inspired by a PIM from the 1980s called Lotus Agenda, notable because of its "free-form" approach to information management. Lead developer of Agenda, Mitch Kapor, was also involved in the vision and management of Chandler.

[0] https://news.ycombinator.com/item?id=39070631 / "For a moment there, Lotus Notes appeared to do everything a company needed"

[1] https://en.wikipedia.org/wiki/Chandler_(software)

refset··on Try Clojure
Give https://github.com/basilisp-lang/basilisp a try
refset··on What "Follow Your Dreams" Misses [video]
Whatever the motivations (and FWIW it does seem like a channel very worthy of support), I can confirm it is at least the official 'Store' linked to from the official website: https://www.3blue1brown.com/
refset··on Computer scientists invent an efficient new way to count
> Nice work with TXDB btw

It's X.T. (as in 'Cross-Time' / https://xtdb.com), but thank you! :)

> 1: https://aidanhogan.com/docs/ring-graph-wco.pdf

Oh nice, I recall skimming this team's precursor paper "Worst-Case Optimal Graph Joins in Almost No Space" (2021) - seems like they've done a lot more work since though, so definitely looking forward to reading it:

> The conference version presented the ring in terms of the Burrows–Wheeler transform. We present a new formulation of the ring in terms of stable sorting on column databases, which we hope will be more accessible to a broader audience not familiar with text indexing

refset··on Computer scientists invent an efficient new way to count
Having read something vaguely related recently [0] I believe "Lookahead Information Passing" is the common term for this general idea. That paper discusses the use of bloom filters (not HLL) in the context of typical binary join trees.

> Letos join

God-Emperor Join has a nice ring to it.

[0] "Simple Adaptive Query Processing vs. Learned Query Optimizers: Observations and Analysis" - https://www.vldb.org/pvldb/vol16/p2962-zhang.pdf

refset··on Jepsen: Datomic Pro 1.0.7075
Thank you for the explanations. Do you happen to know why transactions ("transaction requests") are represented as lists and not sets?
refset··on Jepsen: Datomic Pro 1.0.7075
I don't know whether it was intentional or not, but IIRC DataScript opted for sequential intra-transaction semantics instead.
refset··on How shall we live? (2023)
> Trust is the decision to forgo attentiveness

> If you can’t trust, you have to pay attention in order to verify – and verifying is expensive

> We’d like to believe that goodness brings us together, but that’s not what the data reveal. According to group studies, we don’t come together because we trust: we come together because we align our intent to mistrust.

> We stick together because our interests align and we suffer the mistrust of others until we can no longer justify it.

Interesting perspectives.

Makes me wonder... what is the state of the art in software systems that could help us to "align our intent" in this "impending cyborg age"?

refset··on Homoiconic Python
3) Basilisp (https://github.com/basilisp-lang/basilisp, "A Clojure-compatible(-ish) Lisp dialect targeting Python 3.8+")
refset··on The Sisyphean struggle and the new era of data infrastructure
> What is the relation between "start with why" and "databases are becoming a commodity"?

My assumption is that the author is painting out a longer-term vision both for Confluent and its audience.

It probably makes more sense in the context of other posts, e.g.

> This trend towards object storage is not just happening at Confluent but across the data ecosystem. Many different types of cloud data systems are integrating object storage into their architecture, and I have covered a few of these such as Neon and ClickHouse Cloud. There is also a flurry of start-ups doing logs (such as the Kafka API) over object storage directly (six and counting). There is money being plowed into data systems that use object storage as the primary, and sometimes, only storage layer.

https://jack-vanlightly.com/blog/2024/5/2/hybrid-transaction...

In other words: managed Kafka is becoming a commodity and Confluent should "start with why" to figure out where to go. I figured the advice and perspective is of general enough interest.

refset··on North Yorkshire Council to phase out apostrophe use on street signs
H2 offers quite a comprehensive solution for dealing with this:

> [H2] provides a way to enforce usage of parameters when passing user input to the database. This is done by disabling embedded literals in SQL statements. To do this, execute the statement:

> SET ALLOW_LITERALS NONE;

> Literals can only be enabled or disabled by an administrator

https://www.h2database.com/html/advanced.html

← PreviousPage 5 of 16Next →