HNHacker News
TopNewBestAskShowJobs

refset

1,510 karma · joined August 24, 2015

https://github.com/refset

[ my public key: https://keybase.io/refset; my proof: https://keybase.io/refset/sigs/pyTK0thu8g5O4Zs7EA3HrQYKkYZjeEz_TUr8nZV2hlY ] c0ba44896414421c9009bc8a77a86ceb

meet.hn/city/51.0612766,-1.3131692/Winchester

Socials: - github.com/refset

---

submissionscomments
refset··on Netflix Revamps Tudum's CQRS Architecture with Raw Hollow In-Memory Object Store
> Hollow employs compression techniques

Given how prominently 'compression' gets mentioned I was excited to learn more, but it looks like it simply amounts to using GZIPInputStream/GZIPOutputStream within the blob APIs, or am I missing something...?

refset··on Dyna – Logic Programming for Machine Learning
The dissertation covers extensive details, but on your last point at least it describes:

> 2.2.1 Evaluation by Default. One of the major syntactic differences between Dyna and other logic programming languages is that Dyna evaluates an expression in place by default. The reason for this change is that most terms have a meaningful value, much like how a function returns a value in a functional programming language. Conversely, in logic programming languages such as Prolog or Datalog, terms only “return” the value of true.

refset··on Dyna – Logic Programming for Machine Learning
> Dyna3 — A new implementation of the Dyna programming language written in Clojure

There are some epic looking Clojure namespaces here, e.g. this JIT compiler https://github.com/argolab/dyna3/blob/master/src/clojure/dyn...

refset··on A Quantitative Model of Trust as a Predictor of Social Group Sizes
> In the end, one way of stating the conclusion of this model is that a social group is a form of tool for economic benefit, and the other tools we create build on the cognitive underpinnings of social interactions. Indeed, in the modern world, we often spend more time working with and getting to know our tools and processes than with friends and family members.

> One important implication of our findings from the Wikipedia editors study is that work groups (sets of individuals collaborating on a task) are strictly limited in size to around four individuals. Larger groups fail to coordinate effectively, are more prone to disagreements and conflicts and consequently shed members rather than recruit new ones. This finding has profound implications for how we organize work groups in order to maximize production of technology.

It would be interesting to validate this with GitHub contribution data.

I came across this paper while skimming through Mark's recent blog post on knowledge graphs and LLMs: https://mark-burgess-oslo-mb.medium.com/the-role-of-intent-a...

refset··on Structuring large Clojure codebases with Biff
Datomic offers the ability to declare a "composite index" which can help to accelerate some kinds of access patterns but can't solve 6NF join overheads entirely. If you want guaranteed read performance then denormalized views are the way to go, and perhaps even an IVM engine like Materialize - or this looked promising at one time: https://github.com/sixthnormal/clj-3df
refset··on Future Data Systems Seminar Series – Fall 2025 (draft schedule)
Ah, apologies if I jumped the gun - but seeing the page was live and with all the great talks lined up already ...I couldn't resist sharing!
refset··on Ukrainian hackers destroyed the IT infrastructure of Russian drone manufacturer
The UK at least https://en.m.wikipedia.org/wiki/Poisoning_of_Sergei_and_Yuli...
refset··on ClojureScript from First Principles [video]
The thing that tipped me over the edge into learning ClojureScript was the news almost exactly a decade ago about achieving self-hosted compilation with eval [0] (kudos David!). In the end that specific capability was not quite practical enough for what I needed, but it proved a level of sophistication and maturity in the stack that has only increased since.

Today we also have sci/scittle/cherry for anyone who's seeking that runtime Clojure->JS eval vision. And now with Jank (LLVM Clojure) on the horizon this year it's never been a better time to try Clojure, regardless of which hosted runtime you're enthusiastic to use - Basilisp on Python, ClojureCLR, ClojureDart etc.

[0] https://swannodette.github.io/2015/07/29/clojurescript-17/

refset··on Local-first software (2019)
Lotus Notes always deserves a mention in these threads too, as 1989's answer to local-first development. CouchDB was heavily inspired by Notes.
refset··on XTDB 2.0 is Now Generally Available
These things should become plausible once the engine supports recursive CTEs and incremental view maintenance. In the meantime you could exploit Arrow interop with other engines (while keeping XTDB as the source of truth) - people have been building in this direction, e.g. using Polars in "Chrontext: Portable SPARQL queries over contextualised time series data in industrial settings" [0]

[0] https://www.sciencedirect.com/science/article/pii/S095741742...

refset··on HTAP is Dead
I agree the opportunity is still there, although the long game keeps getting longer.

Prof. Viktor Leis suggested [0] that SQL itself - being so complex to implement and so ineffectively standardized - may be the biggest inhibitor to faster experimentation in the field of database startups. It's a shame there's no clear path to solving that problem directly.

[0] https://www.juxt.pro/blog/sane-query-languages-podcast/

refset··on HTAP is Dead
The HTAP vision was essentially built on the traditional notion that a database is a single 'place' where both transactions happen and complex queries run.

Rich Hickey argued [0] that place-orientation is bad and that a database should actually just be an immutable value which can be passed around freely. That's fairly in line with the conclusions of the post, although I think much more simplification of the disaggregated stack is possible.

[0] https://www.infoq.com/presentations/Deconstructing-Database/

refset··on Databricks acquires Neon
It's not clear to me that the _entire_ Neon stack is OSS and available to self-host (though they do share a lot of OSS code, which is great), and in any case, it's not currently supported/documented beyond some "local development" instructions, e.g. "We do not officially support use of autoscaling externally" [0]

> Can’t you use a cloud provider and have them host this for you?

If it really is all OSS, then I guess the moat is the impressive execution of this team.

[0] https://github.com/neondatabase/autoscaling

refset··on Databricks acquires Neon
> you use S3 as bottomless storage for Postgres [...] Why are people paying?

It's vastly more complicated to do this efficiently than you might imagine. Postgres' internal architecture is built around a very different set of assumptions (pages, WAL, local disk etc.) than what the S3 API offers.

refset··on Yagri: You are gonna read it
> Maybe there are better patterns out there to make this cleaner

SQL:2011 temporal tables are worth a look.

refset··on Append-Only Programming
David Harel, the creator of statecharts, also developed the Behavioral Programming [0] 'b-thread' model motivated by a similar vision for append-only programming - it has been discussed on HN previously e.g. https://news.ycombinator.com/item?id=42060215

[0] https://cacm.acm.org/research/behavioral-programming/

refset··on How about trailing commas in SQL?
In XTDB you can now use (trailing-friendly) commas in this scenario too, instead of ANDs: https://github.com/xtdb/xtdb/pull/3985
refset··on Immutability Changes Everything (2016) [pdf]
> It's relatively easy to implement your own temporal tables on most existing databases

It gets tricky when you need to change the schema without breaking historical data or queries. SQL databases could do a lot more to make immutability easier and widespread.

refset··on Every System is a Log: Avoiding coordination in distributed applications
> a log entry must be appended that says the log is closed

Related concepts: 'epoch' (distributed consensus), 'watermark' (out-of-order stream processing)

refset··on Five years of React Native at Shopify
Explicit query plan pinning helps a lot, alongside strong profiling and monitoring tools.
refset··on Five years of React Native at Shopify
> just unrolling several unnecessary nested subqueries, and adding a more selective predicate

And state of the art query optimizers can even do all this automatically!

refset··on PostgreSQL is the Database Management System of the Year 2024
I wonder where things will stand in 10 years from now. Will many orgs still be consuming vanilla Postgres, or will most workloads have shifted to ~proprietary implementations behind cloud services like "Aurora PostgreSQL Limitless Database" and "Google AlloyDB for PostgreSQL" due to unrivalled price-performance? In other words, can progress in OSS Postgres keep up with cloud economics or will things devolve into an even messier ecosystem centred purely around the wire protocol and SQL dialect?
refset··on Some programming language ideas
For anyone happy enough to consider dealing with the JVM instead of C, and Clojure instead of SQL, I think this CINQ project can deliver on much of what you're looking for here: https://github.com/wotbrew/cinq

> I just write a SQL query instead with joining with seeing the internal data representation of the software as an information system instead of bespoke code

This sounds very similar to how CINQ's macro-based implementation performs relational optimizations on top of regular looking Clojure code (whilst sticking to using a single language for everything).

refset··on Some programming language ideas
If you would like some exposure therapy: https://databasearchitects.blogspot.com/2024/12/advent-of-co... [0]

[0] Recent discussion https://news.ycombinator.com/item?id=42577736

refset··on Databases in 2024: A Year in Review
More specifically, DBOS Inc. raised a $8.5 million seed round [0] and is backed by Michael Stonebraker (the creator of Postgres). I initially assumed Andy was alluding to this when he wrote "the most famous database octogenarian splashing cash" :)

[0] https://techcrunch.com/2024/03/12/new-startup-from-postgres-...

refset··on Reads Causing Writes in Postgres
Interesting! MVCC mechanics aside, it's also worth remembering that work_mem is only 4MB by default [0], so large intermediate results will likely spill to disk (e.g. external sorts for ORDER BY operations).

[0] https://www.postgresql.org/docs/current/runtime-config-resou...

refset··on How to use Postgres for everything
An EAV table is usually a symptom of a wider set of issues with schema management, and in contrast, most people I've spoken to who have implemented their own EAV tables on top of a regular SQL database have ended up regretting it because the approach is too hard to scale and maintain. In your experience, was the EAV model limited to a subset of the overall schema?

I agree a special database shouldn't be necessary at all, and instead, convenient syntax for immutable DML and temporal support should be built into Postgres already. But short of a miracle it will probably take a new ('special') database in order for Postgres to evolve in response. Therefore, in the meantime, we believe there's a gap in the market for organisations who value the 'safety' (foolproof complexity reduction) that native bitemporality in a database can offer above the raw query performance offered by existing update-in-place databases: https://xtdb.com/blog/but-bitemporality-always-introduces-co...

refset··on PostgreSQL High Availability Solutions – Part 1: Jepsen Test and Patroni
Writing Clojure without spaghetti isn't too hard, and definitely more practical than waiting for a Jepsen alternative to come along.

The Jepsen author gave a great talk on all the performance engineering work that has gone into it, Jepsen is near enough an entire DBMS in its own right https://www.youtube.com/watch?v=EUdhyAdYfpA

refset··on How to use Postgres for everything
I meant 'scale' mostly in the sense of 'complexity' (sorry!). If you only have a small number of tables you need/want this versioning for then the DIY approach is workable, but if you want to apply this across an entire schema then things can get complicated fast.
refset··on How to use Postgres for everything
How do you find it when you scale it up to every table, every query?
← PreviousPage 2 of 16Next →