6,841 karma · joined August 20, 2020
contact: hn@kortlepel.com
Do you have a source on this? I live in Germany and most people I know either invest into stocks (safe, boring ones), invest into a house/apartment/land or similar, or are too broke to even save anything at all (most people).
This is not acceptable. Healthcare should be a basic right, and a lot of European countries treat it as such (e.g. in Germany, everyone is insured, no matter their employment status, and its mandatory and equal for everyone regardless of their job).
People also can't be fired for being sick.
I run it like `sbh --net pi` or `sbh --net --docker pi` depending if I want the agent to have docker access. The result is that ~/src/my-project gets mounted as `/w/home/lion/src/my-project` and any LLM I've tried understands that this is a sandbox implicitly.
I strongly suggest everyone who uses a harness, of any kind, to copy it and edit it to better suit one's setup.
Testing on a single device, or two devices, is not enough. I'm on a laptop with 96 GB ram and a Ryzen AI 9 HX PRO 370, and it's running at 3 FPS.
The crucial part is browser performance and network latency. You need to test across multiple browsers, but also evaluate frame times, CPU load, GPU load, etc. so you can make a reasonable estimate whether it's fast or not.
Then, when you realize that its not good enough, e.g. on firefox, you use firefox's profiler to see which parts are slow. LLMs would guess, probably, but you have to measure.
Welcome to software engineering.
> On FrontierCode, which evaluates whether coding agents produce changes ready to merge into real codebases, GPT‑6 Sol improves substantially over GPT‑5.6 Sol, and is able to match Claude Fable 5.1 xhigh at much lower cost.
I continue to appreciate OpenAI's attempt at some honesty here, showing that they are capable enough and have skilled engineers to a point where they can recognize that slop is hated for good reason, and that there is a real issue. Compare this to anthropic, where e.g. in the Opus 5.5 announcement[1] one of the first points on the page is
> One tester completed a 680,000-line code migration in less than a day—work that would have taken an engineering team weeks. It’s good at finding and fixing inefficiencies in software: when we asked it to cut load times across every page of a web app, Opus 5.5 succeeded 39 of 40 times, while Opus 5 made smaller improvements that also altered the app’s behavior. A different tester had several Claude models build a game from a single prompt; Opus 5.5 scored higher than any other model on the strength of its graphics and polish.
This is the kind of shit that is the very reason why I stick to OpenAI and deepseek. OpenAI is simply more honest and reasonable about their models' capabilities, while delivering models that still have solid value.
Notice how the OpenAI announcement doesn't make use of anecdotes.
When humans do this, they inadvertently learn something, too, but when an LLM reproduces or derives and implementation of a chess engine, it in no way implies that the LLM can follow the rules in its own "train of thought" and consistently apply the rules in its "head".
Let's say you want to evaluate my algebra skills. You make me solve some algebra challenges. If I then whip out a computer and write a calculator, or take some sticks and stones and take a couple hours to build an abacus, and then solve the algebraic challenges, this would not constitute a good solution, and would defeat the entire point of the test. If, instead, I do the algebra in my head or on paper, it might seem like there's no difference, but you can derive all sorts of information from that.
For example, you could time it, check for recurring errors I make, for interesting mistakes like mistaking 7 and 1 for one another due to bad hand-writing, etc.
If that was the goal, then me writing a calculator or crafting an abacus defeats the point of the test. Yes, me writing a calculator shows that I'm intelligent, and I understand the algebraic rules, but if the test is about applying the rules, I have not passed.
In the very same way, an LLM writing a chess engine to solve a chess benchmark that is all about LLM's reasoning capability is complete bogus and defeats the entire point.
Writing a well understood engine for a super popular problem does not count as reasoning about the problem.
If you let the LLM write a chess program, which it can ONLY do because there are already so many chess programs out there, then the benchmark becomes about recall of popular program source code, not chess.
Ads don't just steal your time either, they invade your head. You are influenced by them and you have no choice as long as you watch them.
And they look nice, feel nice. Some of them even have a glass/crystal back, so you can watch the mechanism, relentlessly ticking on.
Get an analog automatic watch. It doesn't need to be second-accurate. Everyone else is keeping track of nanoseconds, you'll be okay not knowing.