HNHacker News
TopNewBestAskShowJobs

refset

1,510 karma · joined August 24, 2015

https://github.com/refset

[ my public key: https://keybase.io/refset; my proof: https://keybase.io/refset/sigs/pyTK0thu8g5O4Zs7EA3HrQYKkYZjeEz_TUr8nZV2hlY ] c0ba44896414421c9009bc8a77a86ceb

meet.hn/city/51.0612766,-1.3131692/Winchester

Socials: - github.com/refset

---

submissionscomments
refset··on How Programming Languages Got Their Names
According to Wikipedia:

> The original name SEQUEL, which is widely regarded as a pun on QUEL, the query language of Ingres, was later changed to SQL (dropping the vowels) because "SEQUEL" was a trademark of the UK-based Hawker Siddeley Dynamics Engineering Limited company. The label SQL later became the acronym for Structured Query Language.

refset··on The disaggregated write-ahead log (2023)
Thanks, yes protocol-aware recovery was the context. Pretty sure I first heard it described in Joran's QCon London 2023 talk here: https://youtu.be/_jfOk4L7CiY?t=1460

> If you want your distributed database to maximise availability, how your local storage engine recovers from storage faults in the write-ahead log needs to be properly integrated with the global consensus protocol.

refset··on Implementing system-versioned tables in Postgres
I would guess so, yes.
refset··on The disaggregated write-ahead log (2023)
Consensus protocols, durability and transactional semantics are (should be) closely coupled. I recall TigerBeetle discussing somewhere how they could achieve better throughput and durability guarantees by combining replication/recovery with the consensus protocol, instead of layering it above. I.e. disaggregating the log can be expensive. There's a reference in [0] that might elaborate.

> TigerBeetle is “fault-aware” and recovers from local storage failures in the context of the global consensus protocol, providing more safety than replicated state machines such as ZooKeeper and LogCabin

[0] https://github.com/tigerbeetle/tigerbeetle/blob/main/docs/DE...

refset··on Implementing system-versioned tables in Postgres
I have never researched the timeline properly to understand at which point the concept/code got ditched, but it was part of Stonebraker's 1985 vision: https://dsf.berkeley.edu/papers/ERL-M85-95.pdf

> POSTQUEL allows users to save and query historical data and versions. By default, data in a relation is never deleted or updated. Conventional retrievals always access the current tuples in the relation. Historical data can be accessed by indicating the desired time when defining a tuple variable.

I looked this up during another thread which also has some other sources if you want to dig further: https://news.ycombinator.com/item?id=37956687

refset··on Implementing system-versioned tables in Postgres
The issue is the complexity of figuring out appropriate JSON-compatible serializations and getting the implementation correct for every single column in use. A simple example would be round-tripping the Postgres money type using salary::numeric and later (user_data->>'salary')::money

A much more complex example would be representing a numrange where you have to preserve two bounds and their respective inclusivity/exclusivity, in JSON.

refset··on Implementing system-versioned tables in Postgres
It's definitely a reasonable trade-off given the circumstances. Do you know whether Supabase teams who use this extension simply avoid using any non-JSON Postgres types? Or do they lean on workarounds (e.g. JSON-LD encoding)?
refset··on Implementing system-versioned tables in Postgres
It's especially tragic knowing that Postgres did originally have system-time-like versioning built-in. Instead we get to enjoy being upsold on proprietary ETL-to-Redshift, AlloyDB, etc.
refset··on Implementing system-versioned tables in Postgres
Limited to JSONB types though, by the looks of it.
refset··on PostgreSQL is enough
> which is probably why it's promoted so much

I think what you're actually observing is simply that Postgres is by far the most vendor-neutral DBMS (/API) available, and therefore the volume of conversation & marketing around it stacks up very disproportionately.

In contrast, asides from MySQL all other DBMS options require getting invested in ~one company and relying entirely on the whims & fortunes of their commercial support organisation.

A relevant article and comment thread: https://news.ycombinator.com/item?id=31425872

refset··on PostgreSQL is enough
Better still, take a look at Nile's "tenant virtualization" concept: https://www.thenile.dev/
refset··on What I talk about when I talk about query optimizer (part 1): IR design
> the Query Graph Model (QGM) representation is quite abstract and hardcodes many properties, making it exceptionally difficult to understand. Its claimed extensibility is also questionable.

I don't know much about the context, but it was interesting to note that Materialize scrapped their QGM code last year: https://github.com/MaterializeInc/materialize/pull/17139

Also, a couple of interesting projects in the IR space:

- https://substrait.io/ is a cross-language serialization for Relational Algebra

- https://www.lingo-db.com/ is an MLIR-based (LLVM) query engine described extensively in this paper https://db.in.tum.de/~jungmair/papers/p2485-jungmair.pdf?lan...

refset··on Shouldn't FROM come before SELECT in SQL? (2011)
On page 4 of "A Critique of Modern SQL And A Proposal Towards A Simple and Expressive Query Language" [0] (recently published at CIDR 2024) there's a great diagram covering this and many other "semantical ordering" confusions with SQL. It's an interesting paper in general, by a couple of the biggest names in modern database research (...though admittedly perhaps not language design).

Something like 'SaneQL' (which the paper introduces) deserves to succeed outside of the lab. Source is here [1].

[0] https://www.cidrdb.org/cidr2024/papers/p48-neumann.pdf

[1] https://github.com/neumannt/saneql/

refset··on Embracing Common Lisp in the modern world
> Windows hegemony

:chefs-kiss:

refset··on Embracing Common Lisp in the modern world
I replied with this already to another sibling comment raising the same issue:

> I think he simply meant that Java was marketed and therefore prospered as a counter to Microsoft's dominance more generally.

i.e. ".NET" being used casually (and incorrectly) as a catch-all for anything Microsoft was doing in this space roughly around that period ...possibly suggests there's been an anti-M$ bias!

refset··on Embracing Common Lisp in the modern world
Assuming current multi-tenant infrastructure trends continue, over-provisioning of RAM and slow start up times are likely the much bigger efficiency issues.
refset··on Embracing Common Lisp in the modern world
I think he simply meant that Java was marketed and therefore prospered as a counter to Microsoft's dominance more generally. You're right that .Net (and the CLR) came later in the story ...as a counter to Java.
refset··on Caches: LRU vs. Random (2014)
Thanks for the pointer :) it seems there's a follow-up to that paper in turn, with an explicit focus on that eviction strategy: "Write-Aware Timestamp Tracking"

https://www.vldb.org/pvldb/vol16/p3323-vohringer.pdf

refset··on Caches: LRU vs. Random (2014)
In a similar vein, there's a neat "Efficient Page Replacement" strategy described in a LeanStore paper [0] that combines random selection with FIFO:

> Instead of tracking frequently accessed pages in order to avoid evicting them, our replacement strategy identifies infrequently-accessed pages. We argue that with the large buffer pool sizes that are common today, this is much more efficient as it avoids any additional work when accessing a hot page

> by speculatively unswizzling random pages, we identify infrequently-accessed pages without having to track each access. In addition, a FIFO queue serves as a probational cooling stage during which pages have a chance to be swizzled. Together, these techniques implement an effective replacement strategy at low cost

[0] https://db.in.tum.de/~leis/papers/leanstore.pdf

refset··on Ask HN: Who's writing recursive queries / Datalog?
> Do people really write recursive queries in the wild (non-recursive Datalog doesn’t count)?

I believe a lot of RDF systems demand fast, recursive inferencing. This might be helpful context: https://docs.oxfordsemantic.tech/reasoning.html#materializat...

> What’s the most interesting use of WITH RECURSIVE in SQL you’ve seen?

https://www.sqlite.org/lang_with.html#outlandish_recursive_q...

refset··on Learn Datalog Today
I asked Rich about his thoughts on query optimizers last year (not in the context of Datomic specifically) and his only reservation was around the practical implications for the operational experience. Specifically, that database systems should always provide the means for developers to control exactly how/when/whether existing (cached) execution plans get re-optimized, otherwise query optimizers can actually be a source of greater problems than they solve, particularly for applications with extremely rapid changes in data.
refset··on Learn Datalog Today
Datalog can be very effective for expressing certain kinds of problems and for generating efficient solutions to those problems. Particularly anything that is even mildly recursive, and therefore especially "knowledge graphs" that rely heavily on rules to infer, model and retrieve information. However if your problem domain amounts to CRUD storage without a need for complex recursion then mature SQL systems usually have all the advantages (asides from the syntax!). For a more formal answer:

> The intersection of databases, logic, and artificial intelligence gave raise to deductive databases. Deductive database systems are database management systems built around a logical model of data, and their query languages allow expressing logical queries. A deductive database system includes procedures for defining deductive rules which can infer information (in the so-called intensional database) in addition to the facts loaded in the (so-called extensional) database. The logic model for deductive databases is closely related to the relational model and, in particular, with the domain relational calculus. Datalog is the most known deductive query language (which syntactically is a Prolog subset) where constructed terms are not allowed as other non-declarative constructs such as the cut.

> Also following the relational model, relational database systems are well-known and widespread nowadays. Their formal query languages include relational algebra and relational calculi but, in practical systems, the de-facto and ANSI/ISO standard SQL is the language of choice of every relational database vendor. Whilst SQL and relational formal languages implement a limited form of logic, deductive database languages implement advanced forms of logic.

https://www.fdi.ucm.es/profesor/fernan/des/html/manual/manua...

refset··on Learn Datalog Today
> Datomic has a notion of rules which are mostly syntax sugar and do not support this sort of recursive reasoning.

> Why is that a big deal? When rules are run automatically, you can build live, reactive systems, not just a database that sits around waiting for you to query it.

There was at least one serious attempt to bring these worlds together: https://github.com/sixthnormal/clj-3df

refset··on Learn Datalog Today
For comparison, I previously translated that cart parts scheduling example on the Flix homepage to Datomic-style Datalog syntax: https://gist.github.com/refset/21b3fc1dec9a6928943073809e133...
refset··on Learn Datalog Today
RDFox offers a rather impressive sounding Datalog inferencing engine: https://www.oxfordsemantic.tech/rdfox

> We present a novel approach to parallel materialisation (i.e., fixpoint computation) of datalog programs in centralised, main-memory, multi-core RDF systems. Our approach comprises an algorithm that evenly distributes the workload to cores, and an RDF indexing data structure that supports efficient, ‘mostly’ lock-free parallel updates.

> Materialisation is PTIME-complete in data complexity and is thus believed to be inherently sequential. Nevertheless, many practical parallelisation techniques have been developed [...]

There have been several papers and patents describing their approach, e.g. http://www.cs.ox.ac.uk/dan.olteanu/papers/mnpho-aaai14.pdf

refset··on The Library of Consciousness – Alan Watts
That's awesome! Ram Dass has been a real hero for me. I also love how excerpts from Watts and Ram Dass have been picked up by so many (very talented) musicians over recent years, e.g. https://www.youtube.com/watch?v=nwT7AobR2No
refset··on Constraint-Driven Innovation [pdf]
The narrative here is that the evolution of database systems has always been in response to a rapidly changing landscape of "constraints", e.g.

- The challenge of open R&D collaboration -> Postgres & MySQL

- The ineffectiveness of one-size-fits-all DB engines -> 19 database services at AWS

- Shared disk architectures being difficult to scale -> cloud object storage

- High storage costs limiting the applicability of data analysis -> cloud data warehousing

...and so on. Followed by a summary of the biggest constraints that databases face today:

> What limits the application of infinite cores?

> 1. Data: inability to get data to processor fast enough

> 2. Power: cost rising and will dominate

Conclusion (spoiler!):

> ML central to DB going forward + opportunities with H/W specialization = big database innovations still coming

refset··on My New Computer
> Why do you not trust Ventoy?

> Ventoy is a great tool from what I’ve seen online and the use case it fills does save time and resources. However, I have some reservations about it. If I had to compromise a bunch of critical systems over a long time period, then publishing a great tool and having it tamper with your OS installation media silently would be a really good pick. At this time, I don’t trust the developers of the tool enough and I don’t have the time or skills to perform repeated audits of the software every time they release a new version.

...from the addendum to https://ounapuu.ee/posts/2023/02/15/shrinkflation/ (which was itself discussed previously [0]) - I don't have strong opinions of my own but reading this review was enough to put me off installing Ventoy the other day.

[0] https://news.ycombinator.com/item?id=34800830

refset··on Thinking in an array language
Ah I was thinking about running-on, but compiling-for is probably relevant also :)

> I don't know if FPGAs would have an advantage over GPUs, since the bottleneck is often memory bandwidth and from a quick search it seems FPGAs are worse there

Thanks, that makes sense, though I believe (and hope) the situation can reverse.

refset··on Thinking in an array language
Thank you for the link! That's a very interesting overview. Do you think the commodification of FPGAs could add another angle of interest/motivation here? Or is that hardware model likely to be incompatible with how these compilation techniques work?

I'm personally rather curious about the intersection of compilers and database query engines, where ~cheap dynamic/incremental compilation across extremely diverse runtime workloads is often a basic requirement.

← PreviousPage 7 of 16Next →