HNHacker News
TopNewBestAskShowJobs

brof

2 karma · joined June 29, 2018

submissionscomments
brof··on What if SELECT, FROM, WHERE were functions?
Got it! One more question if I may: in https://arxiv.org/pdf/2607.26356, all the JOB queries are around a second or less, if I read correctly.

My understanding is the query:

movie

    .with(keyword.eq("my-kw"))

    .select(title)
is inlined to something by the compiler approximately like

for movie in 0..movie_count {

    let lo = keyword_offsets[movie];

    let hi = keyword_offsets[movie + 1];

    for pos in lo..hi {

        let kw = keyword_ids[pos];

        if keyword_text[kw] == "my-kw" {

            emit(movie, movie_title[movie]);

            break;

        }

    }
}

So there's no index lookup on keyword, right? E.g., to use a hash index on keyword to find the resulting movie rows. If I wanted to do so, would I re-normalize the data in some way? This is what surprised me: that even without the secondary indexes, it is still the same (or more) performant than DuckDB.

Perhaps the index-lookup version would use something like

let movies_by_keyword: HashIdx<_, _> =

      keyword.text().inv().collect();
but it does not seem like it does (even though I assume DuckDB may).
brof··on What if SELECT, FROM, WHERE were functions?
I've been looking into Prela the last few days and really like it :)

Conceptually, the two key operators `.select()` and `.and()` connect somewhat to the lineage of arrow notation [1], fork algebras [2], etc.

Left-to-right composition:

(>>>) :: (b -> c) -> (c -> d) -> (b -> d)

Fanout:

(&&&) :: (b -> c) -> (b -> d) -> (b -> (c, d))

The database-specific detail is enumerating the universe of entities to maintain shared correlations and using projections as ways to access information, so, in the above notation:

movie :: Movie -> Movie

title :: Movie -> String

movie >>> title :: Movie -> String

The specialization to then have Rust types be able to generate inlined query execution code at compile-time is quite neat. But, I wonder where the limitations of this approach are seen? In contrast to something like Soufflé which has a much more complicated optimizer but generates C++ code ultimately, it seems difficult to always rely on the Rust compiler. Likewise, some other projects like Crepe rely on Rust macros more heavily. And, the embeddability is a plus, but it does not look too easy currently to use from other languages as you would need to compile the query independently and provide something like an FFI wrapper.

Relatedly, some other posts about relational language design from the last couple days: https://news.ycombinator.com/item?id=49342530, https://news.ycombinator.com/item?id=49363617

And this tutorial looks interesting: https://northeastern-datalab.github.io/relational-language-t...

--

[1] https://ghc.gitlab.haskell.org/ghc/doc/users_guide/exts/arro... [2] https://www.cosc.brocku.ca/Faculty/Winter/JoRMiCS/Vol1/PDF/v...

brof··on Sandboxing
> On macOS, Zed uses Apple's Seatbelt sandbox through sandbox-exec.

From `man sandbox-exec`, "The sandbox-exec command is DEPRECATED." However, a lot of AI harnesses seem to use it anyway - I wonder what this says about long-term Apple support.

brof··on Container Is Not a Sandbox
Good summary article but the headline and certain conclusions seem overstated. This article appears to be about AI/serverless compute providers running a multi-tenant environment where untrusted code from multiple customers can be colocated onto a single machine. I don't think anyone would seriously suggest containers are enough for that use case. OTOH, VMs have escapes too, and if you are a compute provider, you are probably relying on additional failsafes like VM-in-container with locked down capabilities, SELinux, and more.
brof··on Optimizing Datalog for the GPU
Yannis is one of the foremost experts in the field, and is responsible for Doop, which as far as I know, is one of the main developments that led to a resurgence of interest in Datalog for program verification :) Him calling himself a "heavy Datalog user" is quite modest.
brof··on Taking Notes with Joplin
Maybe you'll like wiki.vim (not to be confused with vimwiki). I, like you, tried a lot of solutions, and it's the only one that's really stuck over the past couple of years. I think it's because it doesn't try to do too many things and is just a thin wrapper over a collection of text files (you can specify to use Markdown). Obsidian is fancier but I find that I've been able to replicate most of what I need and don't miss the extra bells and whistles.

Oh, and I've found using a tool like gitwatch really helpful for keeping the repo synced without having to remember to commit.