329 karma · joined June 28, 2025
All of which is picked up from my apartment buildings backyard.
one would hope that they disable websearch and internet access (maybe all tools?) when doing generalization testing?
Just as aside, lithium-battery powered smoke detectors are a godsend. The battery lifespan lasts the recommended life span of the detector, i.e. around 10 years. No battery replacements needed, although you still ought to test them once in a while.
> On Apple's unified memory that transfer compiles to almost nothing. It is still written down, because data locality should be provable by reading the source rather than by profiling the binary.
How does profiling the binary play into this? Elsewhere the hypothesis is "it crashes", whereas profiling is about performance.
> Most languages treat the accelerator as infrastructure: you write math, and a large opaque runtime decides how to ship it. Vx treats it as semantics.
What does that mean? Vx doesn't treat the accelerator as infrastructure? Vx doesn't have a large opaque runtime? "You write math, Vx treats it as semantics". I don't understand that.
so the contrast here is the budget available, not the quality of the talent, if we accept the premise that DeepMind and the universities have approximately similar level of talent.
I don't quite understand your point here. OpenAI has a consistent history of releasing cheap/small models - first nano/mini, then luna/terra. Of course, those are now more capable than half a year ago, but I don't see a behavior change from OpenAI here.
I'd expect that their business strategy is to compete in more markets, and if successful, they can capture more value. This is the "easiest" for them as they already have GPUs, a training environment etc. For that platform it's not the worst if there's an internal customer team that can help shape the future and provide immediate feedback, and if it results in a good model, even better.
Other things I'd not be surprised they offer in the future in the same vein: A multi-model harness, coding agent (cloud and local), and maybe at a later point in time even a CPU-only cloud compute product.
To be fair, though, I mainly ran this experiment to distinguish models, not to see if they were better than humans or not, so for that purpose it didn't really matter if the basic conditions were the same or not.
Comparing strategies, though, Astra does something pretty different from the highest scoring submission. It doesn't predict the RNG, instead it reacts to the visuals on-screen, planning ahead by estimating velocity of all objects on screen. Unlike the winning solution, it actively flies the ship, whereas the winning solution basically only rotates and teleports. So Astra behaves more like a regular "perfect" player would, rather than one that breaks the PRNG.
The highest scoring submission that won the contest had a high score of around 137k. Last week, I had GPT-6 Astra, Sol and Luna implement and hill-climb on this task, as I wanted to see how big the difference in smartness is. Luna implemented something, but never exceeded ca. 20k points, with a large variance. Sol got something in the area of the humans implementation.
Astra, which finished fastest, had a highscore of around 1.7Mm when the game seemed to fairly reliably crash. On the way, it disassembled parts of the ROM to extract information about the game.
I didn't do a ton work to document and measure the specifics, but it was very impressive.
Sounds like this is a more serious and phone-centered take on what Steve demoed. We sure live in interesting times.
Surely that's something that can be trained for.
Glad that that's no longer the case!
Also, if you have persistent CI workers with a persistent bazel instance, you save on some network roundtrips, but that's obviously harder to set up and make bulletproof.
Tsgo, oxlint, caching dependencies etc. what linear outlined in their blog post would be more impactful for the average TS project I've worked on.
For a less complex project (1 programming language, still shipping to all 3 major OSes), with my knowledge and agents I got the bazel conversion done in 2 weeks.
The setup cost for bazel just went down by a lot, and I don't think the industry as a whole is aware of that yet.