HNHacker News
TopNewBestAskShowJobs

zwaps

5,489 karma · joined June 26, 2017

submissionscomments
zwaps··on Clef: our open-source decision models
No mention of calibration. Is it just another llm finetune?
zwaps··on Burnham must be clear: no slavery reparations will be paid
There is museum in London called the British Museum which has zero artifacts on display from Britain and is entirely filled with stolen goods from other countries No one there finds this strange, bad, or anything that should be adressed

Brits are entirely deranged when it comes to their former empire.

zwaps··on If AI coding is lowering your code quality, you're not managing quality right
You are holding it wrong!
zwaps··on I built non-autoregressive decision models with RL a year ago
I am sympathetic but this buries the lede, hard.

You are competitive with Jev only if you fine tune on the train dataset and calibrate per question.

As much as I dislike literally everything about typesafes behavior, they have an API model that works on any problem without fine tuning, and that is the key.

To be honest everyone who can finetune can likely finetune a BERT for a specific task and get similar results to yours. And that has been true for years

The key to Jevs success is that it works without fine tuning

zwaps··on Alan's Random Insult Generator (1999)
Hey, don’t copy my daily affirmations
zwaps··on Hyper-Markdown – Markdown for knowledge graphs
Thanks Claude for writing that
zwaps··on Show HN: Shared memory graph for Claude and ChatGPT, over MCP
I literally struggle to read it
zwaps··on Handbook.md shows that long policy documents do not reliably govern agents
If the authors are reading: Why did you let an AI model author parts like "Design Principles"? It's not good writing and it is obvious.
zwaps··on Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard
But then Sol doesn’t really follow the plan or rather invents things it should add or cut on top of the plan. How does one deal with that?
zwaps··on Show HN: Axtary – Content Authorization for AI Agents
Is anyone else having immense problems with this website due to Claude writing it?

Beyond the AI writing style and the pure amount of sentences, nothing seems to transmit the right information at the right time. You need to read the entire page twice to get what is going on.

For example, the first text you read is: —- Axtary checks the exact diff, message, query, or tool payload before a connector executes. Routine actions follow policy; higher-risk actions require approval of that exact payload. —-

Without having read the rest, nothing in this paragraph seems to tell me anything.

Only after I read the rest of the website do I somehow get what is even meant.

zwaps··on Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard
Sol is a complete mess for me.

It only works on end to end tasks in fresh codebases.

Otherwise it cannot follow instructions, changes and deletes unrelated features or does sloppy work to mark a task completed while leaving a compromised codebase.

I could not get Sol to finish a feature in a complex code base without several loops of fixing and reverting

zwaps··on AIs don't do what you want. This is bad
Gpt 5.6 in codex is unusable for me.

It is no longer able to execute well defined changes. It does stuff that wasn’t asked for, deletes pieces of functionality unrelated to the task and introduces regressions everywhere.

I blame it on benchmaxxing. I fear coding models no longer work in a large complex codebase

zwaps··on Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber
Sure, but there's no sota alternative from Google. That's it, and its beaten by GLM 5.2 on every measure except somewhat speed.

I find that quite staggering. GLM is open weights

zwaps··on Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber
according to AA it's not better than GLM 5.2 and that's surprising to me
zwaps··on Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber
Here's the issue:

GLM 5.2 is better, also cheaper, and almost as fast.

So essentially, a big L for Google. Combine this with them not being able to produce a frontier model this generation... hmm implications

zwaps··on Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber
More likely they don't manage to advance benchmarks on the SOTA level anymore. In other words: They can't beat 5.6 nor Fable
zwaps··on Xiaomi-Robotics-1
Coining the term: Slopfold

When you Robot folds your clothes but it's kinda wrinkly and sloppy but you accept it anyway because it's easier than precisely folded and pressed clothes.

zwaps··on Harness Engineering
You call your own blog post seminal and quote yourself in the repo.

Are you … alright?

zwaps··on A Beautiful Theory Falls to Ugly Data
Obviously the theory is decades old. I think nowadays a Game Theorist would not go and claim that the fixed-point convergence is an actual market process with consumers and monopolists trying to outsmart each other in some kind of bizarro bazar game.

An equilibrium for a given game is - depending on the equilibrium concept (bummer, even more conditions) - is a stable outcome of some sort with usually no claims as to how it would actually be reached.

By that, you can already see that this is not really an actual theory of an empiric situation, but rather a mathematical model of a certain solution structure.

If you were to write this paper today as an economist and your goal was to claim that is actually, really holds in reality, then you'd not only have to produce the theory but you'd also have to build some sort of empirical model that you can estimate with somewhat plausible identification conditions and structure, or be able to show it in a (pseudo-)experiment setup that is believable enough. Suffice to say that there are very few such claims made on reality in modern microeconomics (that is to say, Game Theory by and large)

As it turns out, these sort of mathematical models have quite a bit of value in a normative setup, say if you go and design a market or an auction. Less so as a theory to explain all of reality.

I think in Coase's time, it was easier to write a 6 page paper from your bathtub and claim something about the world. Wasn't there an xkcd comic like this?

zwaps··on A Beautiful Theory Falls to Ugly Data
Not really. Game Theory (in this iteration at least) is about identifying equilibria, not about the process of reaching them. This is one of several "deviations" of Game Theory from "reality". The fact that equilibria are fixed-points and can be explained in some sort of bargaining process doesn't really mean that's how we should actually imagine them. If we do so, we both overstate the theory (claiming some sort of actual behavioral process) and understate it being a (possibly quite general) fixed point to many possible market and non-market processes.

If you want to make Game Theory collide with reality, the actual convergence to an equilibrium is only one of many venues where there is a large divide. Other assumptions of these models - from rational behavior to uniform prior assumptions - are equally problematic.

Game Theory models are nevertheless very helpful because they require you to actually lay out your assumptions or - when you observe something else - reason about "what else is going on" in any of these areas. As it turns out (as another person has said), it is also immensely helpful when designing mechanism (i.e. games) like a Steam store or an ad auction, which is why tech companies hire quite a few Game Theorists.

zwaps··on A Beautiful Theory Falls to Ugly Data
Somewhat that is it.

The issue if, of course, that marginal revolution overstates the contribution of a single empirical study here.

Of course everyone is aware that the original model doesn't hold in reality. The contribution of showing this in the ebook market is... not zero, but certainly not the implied "We killed the theory!!!!"

Instead, there are decades of papers poking holes in the Coase model and producing ideas as to why the conjecture doesn't hold. In my mind, these are the more interesting contributions. The authors mention two, but I think far more tangible are time preferences, time horizon limits, pertubations, non-uniform prior assumptions and bounded rationality.

zwaps··on A Beautiful Theory Falls to Ugly Data
Some of the papers are linked. These papers, decades old, helpfully even show cases where the result does not hold.
zwaps··on A Beautiful Theory Falls to Ugly Data
Yes, the logic applies to all sorts of bargaining situations - that's what later papers mentioned in the article is about
zwaps··on Dispersion loss counteracts embedding condensation in small language models
Is this not just embedding anisotropicicity?

Big topic early 2020s

zwaps··on Good results fine tuning a local LLM like Qwen 3:0.6B to categorize questions
Has anyone compared recently doing something like ModernBERT plus classifier vs. full or lora FT of a small LM like qwen?
zwaps··on Identity verification on Claude
Oh no that’s terrible. They even say the data is used by the external company to train and use however.

Shit now i have to cancel my account

zwaps··on Running local models is good now
Does not apply to oss models
zwaps··on Ask HN: Did we witness the "Trinity moment" for AI?
It will not. We are unable.
zwaps··on Home alone: Remote work, isolation, and mental health
In UK and Spain, computer engineers earn the median income?
zwaps··on Ferrari Luce
“Sound waves are captured from electro-mechanical vibration in the axles“

Finally! Electronic sound is fine if it is the actual sound of the car instead of some fake recording of a v8

Page 1 of 34Next →