HNHacker News
TopNewBestAskShowJobs

refset

1,510 karma · joined August 24, 2015

https://github.com/refset

[ my public key: https://keybase.io/refset; my proof: https://keybase.io/refset/sigs/pyTK0thu8g5O4Zs7EA3HrQYKkYZjeEz_TUr8nZV2hlY ] c0ba44896414421c9009bc8a77a86ceb

meet.hn/city/51.0612766,-1.3131692/Winchester

Socials: - github.com/refset

---

submissionscomments
refset··on Vanna.ai: Chat with your SQL database
It's not a public facing product, but there was a talk from a team at Alibaba a couple of months ago during CMU's "ML⇄DB Seminar Series" [0] on how they augmented their NL2SQL transformer model with "Semantics Correction [...] a post-processing routine, which checks the initially generated SQL queries by applying rules to identify and correct semantic errors" [1]. It will be interesting to see whether VC-backed teams can keep up with the state of the art coming out of BigCorps.

[0] "Alibaba: Domain Knowledge Augmented AI for Databases (Jian Tan)" - https://www.youtube.com/watch?v=dsgHthzROj4&list=PLSE8ODhjZX...

[1] "CatSQL: Towards Real World Natural Language to SQL Applications" - https://www.vldb.org/pvldb/vol16/p1534-fu.pdf

refset··on Thinking in an array language
I'm far from an expert on such things, but Co-dfns appears to be genuinely cutting edge research that could have utility for many language ecosystems (or do you disagree that this direction of GPU-based compilation holds promise?). If that work has to look like a "mess" to non-APLers in order to get done then so be it.
refset··on Thinking in an array language (2022)
Nothing has ever convinced me about the potential for array languages in practice quite like watching Aaron Hsu describe how he develops his parallel APL compiler [0] using two Notepad.exe windows side-by-side: https://www.youtube.com/watch?v=gcUWTa16Jc0&t=860s

He has also written many comments (arcfide on HN) about this stuff before, e.g. discussing "semantic density" [1]

> The compiler is designed so that I can see as much as possible with as little indirection as possible, so that when I see a piece of code I not only know how it works in complete detail, but how it connects to the world around it, and every single dependency related to it in basically one single half screen full of code (usually much less than that) without any jumps, paging, scrolling or any movement. [...] The idea of semantic density is critical to this point. The semantic density of the APL code I'm using to solve the problem is at a certain rate. I maintain a consistent density rate by choosing my variable names in such a way that they visually align with the expressivity per character of the built in primitive symbols.

And

> [...] idiomatic programming methods that are so concise, they can begin to be read as we read and chunk English phrases. By doing so, it becomes actually easier to just write out most algorithms, because the normal name for such an algorithm is basically as long as the algorithm itself written out. This means that I start to learn to chunk idioms as phrases and can read code directly, without the cost of name lookup indirection. I can get away with this because I've made reusability and abstraction less important (vastly so) because I can literally see every use case of every idiom on the screen at the same time. It literally would take more time to write the reusable abstraction than it would to just replace the idiomatic code in every place.

[0] https://github.com/Co-dfns/Co-dfns

[1] https://news.ycombinator.com/item?id=13571159

refset··on Show HN: Dbeel – A distributed thread-per-core db
This post on Glommio is insightful - "Glommio is a cooperative thread-per-core crate for Rust & Linux based on io_uring": https://www.datadoghq.com/blog/engineering/introducing-glomm...

The key insight being that in this model "locks are never necessary", which is huge source of efficiency when working with sharded data across multiple cores.

> Thread-per-core architectures are friendly to modern hardware, as their local nature helps the application to take advantage of the fact that processors ship with more and more cores while storage gets faster, with modern NVMe devices having response times in the ballpark of an operating system context switch.

refset··on Databases and why their complexity is now unnecessary
"YugabyteDB Anywhere" is open core though right? Half the value & complexity is in the orchestration stack. It's definitely a step forward, but such licensing is likely still too restrictive to ever supplant Postgres itself (it terms of breadth of adoption).
refset··on Databases and why their complexity is now unnecessary
If Postgres was already horizontally scalable and supported incrementally maintained recursive CTEs (like Materialize can do) then I could see how Rama would be mostly uninteresting to a seasoned SQL developer, but as it is I think Rama is offering a pretty novel & valuable set of 3GL capabilities for developers who need to build scalable reactive applications as quickly as possible. Capabilities which other SQL databases will struggle to match without also dropping down to 3GL APIs.

> A purely relational database that jettisons SQL doesn't have to have the limitations the author is poking at.

Agreed. Relational databases can take us a lot further yet.

refset··on Databases and why their complexity is now unnecessary
I can imagine Codd saying the exact inverse: any sufficiently complex data model quickly becomes intractable for developers to assemble ideal indexes and algorithms together each time in response to new queries, which kills productivity and reduces iteration speed. Particularly as the scale and relative distributions of the data changes. The whole idea of declarative 4GLs / SQL is that a query engine with cost-based optimization can eliminate an entire class of such work for developers.

Undoubtedly the reality of widely available SQL systems today has not lived up to that original relational promise in the context of modern expectations for large-scale reactive applications - maybe (hopefully) that can change - but in the meantime it's good to see Rama here with a fresh take on what can be achieved with a modern 3GL approach.

refset··on Databases and why their complexity is now unnecessary
> If all business happens in one transactional system, your semantics are dramatically simplified.

100% agreed. One of the biggest issues SQL databases have faced is that the scope & scale of "one transactional system" has evolved a lot more quickly than any off-the-shelf database architecture has been able to keep up with, resulting in an explosion of workaround technologies and approaches that really shouldn't need to exist.

We're now firmly in the cloud database era and can look at things like Spanner to judge how far away Postgres remains from satisfying modern availability & scaling expectations. But it will be great to see OSS close the gap (hopefully soon!).

refset··on Databases and why their complexity is now unnecessary
Maybe https://plato.stanford.edu/entries/relations/
refset··on Trade-offs between Different CRDTs
Operational Transform (an approach that requires a central server to coordinate edits)
refset··on Show HN: I made an app that consolidated 18 apps (doc, sheet, form, site, chat…)
It runs on L̶o̶t̶u̶s̶S̶c̶r̶i̶p̶t̶ JavaScript. Speaking of, Mitch Kapor's 1984 memo is turning 40 this year:

> With the formal commencement of the "Notes" project upon us, it seemed appropriate to set down a few brief notions about the project, its scope, and its strategic importance to Lotus. This material should be regarded as more than highly confidential.

https://web.archive.org/web/20180225100127/http://www.kapor....

Also discussed previously: https://news.ycombinator.com/item?id=13168969

refset··on Databases in 2023: A Year in Review
Definitely sarcasm, it's a running gag - see the 2021 and 2022 posts also :)
refset··on Databases in 2023: A Year in Review
> Anything loosely connected to AI + LLMs received the bulk of the attention (rightly so, as it is a new chapter in computing)

Beyond LLMs and vector search, the scope for applying AI / machine learning _within_ databases is enormous: join planning, learned indexes, compression, workload prediction, configuration tuning etc.

Given the current pace of advances perhaps a dedicated "AI in Databases in 2024" review will be on the cards this time next year.

refset··on Thoughts on PostgreSQL in 2024
> building and managing an active-active system is extremely complicated

Could Postgres yet evolve to become a Spanner-like multi-writer system?

refset··on Adventures with compression
Compression is, ultimately, AI.

https://news.ycombinator.com/item?id=38399753

https://news.ycombinator.com/item?id=31923231

refset··on Fake Trees: Using Indents for Simpler UIs
Take a look at the kidvis function in hn.js
refset··on The value of canonicity (2020)
I'm curious about the usage of CockroachDB at Nubank, given it's not mentioned as a "paved road" (or otherwise) in the post alongside Kafka/Datomic/Spark, but is described here: https://www.cockroachlabs.com/customers/nubank/
refset··on Moving from relational data to events
Absolutely, there's no beating the RUM Conjecture.
refset··on Moving from relational data to events
"normalize until it hurts, denormalize until it works" is evergreen advice for scaling both reads and writes. Synchronously enforcing referential integrity and other forms of normalized constraints is what gets expensive.

Pat Helland has some really good writing on this stuff, e.g. https://pathelland.substack.com/p/i-am-so-glad-im-uncoordina...

refset··on Moving from relational data to events
> Datomic [...] has temporal support

Note it only supports "transaction time" (or "system time" per SQL:2011) but not "valid time" (~"event time") which is needed for a bitemporal data model. [1]

[1] https://vvvvalvalval.github.io/posts/2017-07-08-Datomic-this...

refset··on S3 Express Is All You Need
I am eager to hear how it will affect their latency numbers:

> Engineering is about trade-offs, and we’ve made a significant one with WarpStream: latency. The current implementation has a P99 of ~400ms for Produce requests because we never acknowledge data until it has been durably persisted in S3 and committed to our cloud control plane. In addition, our current P99 latency of data end-to-end from producer-to-consumer is around 1s

via https://www.warpstream.com/blog/kafka-is-dead-long-live-kafk...

refset··on The Three Projections of Doctor Futamura (2009)
For practical implications see Truffle & Graal:

> Truffle works by taking your interpreter and generating a compiler by partial evaluation. So I believe they have written a JVM bytecode interpreter and Truffle makes the JIT.

> I'm super impressed by how the Truffle/GraalVM team has been able to turn this theoretical concept into a system that yields production grade compilers

via https://news.ycombinator.com/item?id=25842901

refset··on RSS can be used to distribute all sorts of information
https://yakread.com
refset··on Make real, the story so far
I was fortunate enough to watch Steve present some of these demos live last night at the Future of Coding meetup in London where the whole room was buzzing - the potential here feels immense. Especially when adding a more rigorous FSM/statechart engine into the mix, which @davidkpiano has already demonstrated: https://twitter.com/DavidKPiano/status/1725522630457840037 (as opposed to relying on GPT-generated JavaScript...based on a png of a FSM diagram!)
refset··on Pinball implemented using Squint, a ClojureScript dialect
The Wordle clone is less exciting but you can scroll down to see the source code and JS output more easily: https://squint-cljs.github.io/squint/?src=https://gist.githu...
refset··on From Datalog to SVG
This is really neat. Reminiscent of how drawio.com can embed the source inside exported PNGs (see also https://github.com/Fusion/pngsource)

I would love to see this approach extended to add some animation too.

refset··on Rules of schema growth (2017)
Views seem like the ideal solution in theory, but SQL implementations of views are often problematic in practice due to planning complexity & overhead. In theory they should also be a good mechanism for handling writes (see "updateable views" / "writable views") but the list of caveats is long and many developers are understandably nervous about pushing lots of logic into TRIGGERs.
refset··on Show HN: Light implementation of Event Sourcing using PostgreSQL as event store
"retroactive events" is probably the thing to look for, e.g. https://www.infoq.com/news/2018/02/retroactive-future-event-...
refset··on Why you should probably be using SQLite
Exactly, the n+1 problem is really about query optimization, latency just makes it more pronounced.
refset··on Temporal Databases (1986) [pdf]
I really like Kent Beck's recent re-framing of this overly technical topic in terms of "Eventual Business Consistency": https://tidyfirst.substack.com/p/eventual-business-consisten...

> In a nutshell, we want what’s recorded in the system to match the real world. We know this is impossible (delays, mistakes, changes) but are getting as close as we can. The promise is that if what’s in the system matches the real world as closely as possible, costs go down, customer satisfaction goes up, & we are able to scale further faster.

← PreviousPage 8 of 16Next →