HNHacker News
TopNewBestAskShowJobs

kgeist

5,099 karma · joined March 7, 2021

submissionscomments
kgeist··on Don't classify, hallucinate
It's basically a variation of HyDE (Hypothetical Document Embeddings), and the rationale is that the embedding of the query is not necessarily close to the embedding of the answer. If you generate a hallucinated answer, it can line up with the actual document better (in the embedding space, via BM25, or hybrid).

But honestly, it only works for common knowledge that's already in the LLM. If the target document contains very niche or private information, then the hallucinated answer's embedding can be even farther away than the query's.

kgeist··on Don't classify, hallucinate
A common case I have is when you don't have classifications to begin with. For example, you need to find what users complain about most. I take embeddings of all records, then cluster the embeddings into semantic groups, then ask an LLM to take a random sample from each clustered group and create a classification for that group.

This method is sensitive to the thresholds (what is the maximum distance between embeddings for them to be still considered part of the same semantic group), so I run it all in an agentic loop where an agent tries different thresholds and clustering algorithms until it's satisfied with the result, plus it may deduplicate some groups.

I run it all on self-hosted hardware, so it costs nothing to leave it running for, like, a night, and as a bonus, none of the corporate data leaves the office. I think a rigid set of manually created classifications may not capture all the possible classifications that can exist. Needs a review by a human, though.

kgeist··on Claude users are mad that Anthropic's new watermarks will catch them using it
I wonder how it impacts code generation. It shouldn't impact prose in general because of synonyms and whatnot, but code requires exact reproduction. That is, what happens if you ask an LLM to recite a large, human-written excerpt as is, without modifications? Wouldn't the modified token sampler try to change some tokens here and there (for the watermark to work)?

For example, what if I say, "Repeat this text verbatim: %long_human_written_text%"? Would the output be recognized as AI-generated or human-generated?

kgeist··on Nvidia Nemotron 3.5 Lightning and NeMo Switchyard
More compute/larger datasets during training != larger models. The Bitter Lesson was that just scaling things up beats custom hand-crafted optimizations. Up until 2024, we thought that meant scaling up the parameter count, but then that started to plateau. After o1 was released, we thought it was about scaling up test-time compute. Now, seeing how 30b models can easily outperform 200b models from a few years ago, it seems like what we need to scale is RL, at least for agentic capabilities. It looks like 30b is already enough for good agentic capabilities. Larger models aren't considerably "smarter" (especially since they're mostly MoE anyway, with something like the same 30-50b active parameter range), they just know more (better world knowledge), which lets them make more informed decisions. Maybe we just need to scale up the retrieval layer.
kgeist··on Go is an ideal language for AI-assisted software engineering
I agree with the article, but there's one thing Go has that doesn't help LLMs: structural typing. An LLM has to grep a little more to understand which interfaces a struct implements.
kgeist··on Stealing Reasoning Traces from Proprietary LLM APIs
In the BlackHat presentation on the HuggingFace incident, OpenAI showed some excerpts from the reasoning traces, and they had that grug speak too (skipped articles, etc.). So the OP's method must have indeed found the actual reasoning traces.
kgeist··on Show HN: DeepSeek-V4 Latent Reasoning – moving "thinking" into latent space
LLMs already think in latent space. The generated reasoning tokens are only the surface of what's happening internally. An LLM may write one thing in the traces but decide differently in the latent space. The whole token-based "reasoning" thing was just a clever hack to extend the existing architecture without completely redoing it. In one of Anthropic's recent papers, they added an additional subnetwork trained to map internal states to readable text, so that's probably the vector of further development.
kgeist··on Show HN: DeepSeek-V4 Latent Reasoning – moving "thinking" into latent space
The article is AI-written as well, with all the "honest problems" and "real weak spots". The code must be AI-generated too, so without a human properly verifying it, I can't take the project seriously. It may just be AI hallucinations congratulating themselves on imaginary achievements.
kgeist··on Shopify replaced Redis with MySQL for inventory reservations–and it scaled
The social network VK internally uses highly specialized database engines per business domain. They don't use stock DBs. They have a DB engine for posts, a DB engine for likes, etc. They have a team of DB engineers. Their DB load was around 250 mln RPS 3 years ago. Stock DBs were harder to scale for them. I guess if you have immense highload, having a team of DB engineers can be cheaper because you can save a lot on servers. I reviewed their code. A DB engine's source code is pretty compact and simple (relatively speaking) because they deal with very specific domain entities, so they don't have to account for all the possible user query combinations that a general-purpose DB would have to support. It was mostly shards+binlog+snapshots+views in RAM. Considering that Telegram was founded by former VK engineers, I suspect they have something similar.
kgeist··on DeepSeek V4 Flash 0731
I'm not sure it's a fair test either to compare the "low" setting of one model with the "low" setting of another. They're completely different settings that just happen to have the same name.
kgeist··on US strikes $1.2B deal to pay German firm to halt offshore wind projects
It's an old, small model, so that's expected. It's more of a prototype.
kgeist··on Improving GPT‑5.6 Sol in ChatGPT, expanding GPT‑5.6 Luna access for free users
>avoid extra detail when it does not help

I wonder if they actually do it to optimize inference. I maintain a corporate AI server and one of the tricks to reduce the load was to modify the system prompt to be as terse as possible so the average response completes faster and requests queue up less often.

kgeist··on Is AI reasoning right for the wrong reasons?
Transformers lack recursion and are limited by the network's fixed depth, so "reasoning", IMHO, is basically a way to emulate deeper recursion. As we go through the layers, concepts are pattern-matched and refined, but at some point we have to stop and cannot refine them any further (no more layers). Usually, this refinement continues during the generation of the next token (the previous intermediate results needed to continue the refinement are still in the KV cache).

But some problems require a substantial number of pattern-matching and refinement steps. The problem is, we also have interference from the fact that the model is trained to model language using mostly non-reasoning data of varying semantic lengths. Because of that, it may stop generating text before the abstract refinements are fully completed, simply because the pretraining data tells it to.

So we have to additionally train models to produce "reasoning traces" so that the emulated recursion continues for longer than what is typically found in pretraining data, allowing the model to build richer and more complex abstractions and surface more concepts. The ability to split problems into steps and logically connect concepts is already present in non-reasoning models (the original CoT trick), because some of it exists in the pretraining data, but not enough to support much longer recursion (hence the premature stops).

As for whether it is "true reasoning" or not, I think that is just arguing semantics for the sake of it. LLMs can demonstrably solve various complex problems. Yes, they often make stupid mistakes, but don't we have the saying, errare humanum est? Don't humans make mistakes too? Don't we also have around 200 cognitive biases showing that we "simply pattern-match" too? I think we still cannot get rid of the Great Chain of Being idea.

kgeist··on What Even Are Microservices?
In my experience, almost all the problems that microservices advertise solving can also be solved with a modular monolith plus some tooling to enforce certain rules (say, one module shouldn't be able to peek into another module's internals, bypassing an agreed-upon "clean" public interface; that alone solves most spaghetti-code problems)

There are two things monoliths can't easily offer:

* Using different frameworks, languages, etc. But in my experience, it's pretty rare for a team to use many programming languages at once. Usually, it's just a few highly performance-sensitive services that need to be written in another language (say, a proxy in Rust while the rest is in Python). For that, I prefer an architecture with one main monolith plus a few high-performance satellite services. No problem there.

* More optimized scaling in certain scenarios. Say I have a module that processes files and can use all available CPU. I might want to put it in a separate container on another node so that the processing doesn't destabilize the core web server. Technically, monoliths support this too, just run the monolith in a different mode (say, behind an `--image-process` flag of sorts), and you can schedule it on another node in the same way. The only downside is that it may use more RAM than necessary for the extra binaries or scripts that won't be used

What else am I missing?

kgeist··on Agent swarms and the new model economics
The knowledge is lossy, and code generation itself is non-deterministic (temperature), so the operator-tree executor must be interference from other DB implementations, because it's uncommon to have a bytecode interpreter
kgeist··on Nativ: Run frontier open models locally on your Mac
We've shipped some code generated by Qwen3.6 27B to production (under OpenCode). It lacks the breadth of knowledge of models like Opus, but if a change is fully inferable from the prompt and the surrounding code, it works very well. It won't be able to write something from scratch that requires niche knowledge (say, a performant inference engine tailored to Blackwell GPUs), but if it's just a PR adding a new use case to an existing project (which is usually just "load from the DB, do some invariant checks, modify the entities, store them back"), it works as well as Sonnet (provided you have the correct configuration, like recommended temperature and top-p settings, the model isn't over-quantized, you have at least 150k tokens of context available, etc.).
kgeist··on Agent swarms and the new model economics
Even if no Rust code for it was seen during training, an LLM can trivially transpile SQLite's C codebase to Rust on the fly. For example, I just asked ChatGPT to write John Carmack's famous Fast Inverse Square Root algorithm in Erlang, without searching online or thinking, and it transpiled it immediately (while also extracting the knowledge in the same step). SQLite's semantics/code are stored in the middle layers of an LLM, and the last layers are able to convert it into any representation, as conditioned by the prompt. Cursor's experiment is deeply flawed because they merely extracted the model's compressed, lossy knowledge of SQLite's codebase and then just ran a bunch of tests/fixing rounds to make up for the lossiness. The claim that the agents built it from scratch is false.
kgeist··on Lobste.rs is now running on SQLite
I reproduced it in Sqlite with short-lived reads/writes though. Other DBMSes seem to not have this issue (IIRC MySQL will block a write if WAL falls behind)
kgeist··on Lobste.rs is now running on SQLite
On the page you linked:

>However, if a database has many concurrent overlapping readers and there is always at least one active reader, then no checkpoints will be able to complete and hence the WAL file will grow without bound.

>This scenario can be avoided by ensuring that there are "reader gaps": times when no processes are reading from the database and that checkpoints are attempted during those times.

Dunno, maybe Rails has a built-in workaround for this.

My workaround was to run a separate thread that monitored the WAL size on disk every second. If it went above the target size of 8 MB, my framework would enter "slow down" mode, where all reads and writes were artificially delayed by calling "sleep()", starting at 16 ms and gradually increasing the sleep time based on a few heuristics.

This allowed the application to have short gaps with no reads or writes, so the checkpointer could actually proceed.

kgeist··on Lobste.rs is now running on SQLite
They use WAL in SQLite. If I continuously perform reads/writes so that they overlap with no gaps, I can make their VM go down because SQLite will not have time to initiate a checkpoint to trim the WAL file. SQLite waits for a time window without any active reads/writes before starting a WAL checkpoint. If there isn't one, the WAL will grow indefinitely, eating up all the disk space on the VM.

It's in SQLite's documentation, and almost no one switching to Sqlite seems to be aware of it because no one discusses it in blog posts like these. I guess most projects switching to SQLite have very low traffic and no malicious users (yet)

kgeist··on Show HN: Getting GLM 5.2 running on my slow computer
Ollama uses 4 bit quants and a very short context window by default. It can easily break on anything more complex than a simple chat.
kgeist··on Hy3
Qwen3.6 below Q8 often can't exit a reasoning loop (until it hits max output token count), forgets to insert a tool call, often mistakenly inserts them inside the thinking block... It's still usable though.
kgeist··on Show HN: Getting GLM 5.2 running on my slow computer
How was qwen3.6 launched?

The thing is, everyone has their own variant of "qwen3.6 27b" depending on the launch parameters, ranging from "SOTA in its class" to "completely broken"

kgeist··on Why developers are ditching GitHub for Codeberg and self-hosting alternatives
VPN, accessible only from inside the corporate network
kgeist··on Why developers are ditching GitHub for Codeberg and self-hosting alternatives
Yes, Rocket.Chat
kgeist··on Why developers are ditching GitHub for Codeberg and self-hosting alternatives
We've been self-hosting GitLab for about a year now, and I don't remember it ever going down or being unavailable. We self-host almost everything else too (except for online meetings), and it's all been pretty stable as well. Some of the tools we self-host do go down occasionally, but it's usually just a matter of restarting the VM or adding more storage.
kgeist··on Rewriting Bun in Rust
It converges to "almost deterministic" on highly predictable outputs (i.e. code) with the right sampling params (say, you only sample the most probable token without randomness/high temperature) and with self-correction loops
kgeist··on Rewriting Bun in Rust
>they turned it into something unreadable

Did you compare the code before/after? It's a mechanical line-by-line port, and most of the code is identical to the old version, just with Rust syntax. They have an example in the blog post.

kgeist··on SWE-1.7 Reach Near GPT 5.5 and Opus Intelligence
On artificialanalysis.ai, Kimi 2.7 Code is way worse than GLM 5.2 at everything (general intelligence, coding, agentic tasks).

But here, both Kimi 2.7 and its derivative SWE-1.7 are ahead of GLM 5.2. This tells me the benchmarks they use are cherry-picked.

kgeist··on 98% Isn't Much
>Pragmatically, often users without new browsers and OSses are not the best clients

Hmm, it could be fat enterprise clients with locked-down software versions (legacy, security etc.) That's where most of the money is, isn't it?

← PreviousPage 2 of 34Next →