for those who need a primer on this: https://endash.us/toolkit/items/mini-shai-halud (yes an ai-assisted build obvi, but not just an ai blog)
12 karma · joined August 19, 2026
for those who need a primer on this: https://endash.us/toolkit/items/mini-shai-halud (yes an ai-assisted build obvi, but not just an ai blog)
"We (the researchers and the agents) did not see any of their work through any means until they released it publicly — in particular, no specific user data was accessed in order to solve this problem. While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models . However, our proofs differ significantly and even the precise results proved are different in the Euler case (forced vs unforced)." https://openai.com/index/navier-stokes-solution/
NOTE: there are a couple duped threads around this. i replied on a different one first before seeing this one
What's most striking to me, and what may or may not be true, is the "we cannot rule out" bit.
"We (the researchers and the agents) did not see any of their work through any means until they released it publicly — in particular, no specific user data was accessed in order to solve this problem. While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models . However, our proofs differ significantly and even the precise results proved are different in the Euler case (forced vs unforced)." https://openai.com/index/navier-stokes-solution/
Silly, stupid, whatever. But it's open-source and able to be cloned/forked to answer your "real" questions with ease.
Oh, and you can submit your own questions to be added to the list.
Built with assistance from n-dx
For those who aren't familiar with teamPCP's big attack recently, we made a Scroll to highlight how it was so effective: https://endash.us/toolkit/items/mini-shai-halud
Classic hindsight bias.. And it's funny that Sam is saying this.. just a few months ago i saw his interview where he claimed to be very good at seeing the future and predicting what will happen
First thing we do is to map out an existing codebase (or a piece of a codebase if it's a monorepo). We have the models synthesize what's in there and document the architecture AND the intent of the app (retroactively building a PRD based on codebase). We use n-dx for this and it has worked very well. There's still bugs and plenty of improvements to be made with the tool (n-dx), but it's open-source so it's easy to make improvements whenever it falls short during a set of tasks
We break down our tasks pretty granularly before they get picked up by a model, and for that workflow we've found that sonnet 4 and opus 4 are still quite effective, and debatably more effective than the 5s
for reference, we use n-dx (https://n-dx.dev) for our workflow
It taught me a ton, and I wanted to take a stab at making a free online version of it
This is meant to mirror the classic MIT game as closely as possible, but able to be played remotely and by anyone who has the link
Anyone can play a full round right now, with no signup required. If you're teaching a class or running a workshop, a free En Dash account lets you host sessions (and you get the full debrief/readout persisted to your account)
Built by me, with some help from other folks at En Dash Consulting.
Play a round and let us know what score you managed :D Would love to hear any thoughts!
this also reminds me of what got me hooked on CS in the first place: a simple java steganography app in cmsc150
- satisfying the needs of parallel agentic development is wholly aligned with the optimal DevSecOps CI/CS/CD WhateverTerm models out there. And that's rad, bc a lot of orgs have a reference frame to map to.
- microservices, monoliths, monorepos, mammoths, whatever.. The code and services can be structured however, so long as the release capabilities are modular and governable/manageable/auditable/flexible/transparent/etc. A killer workflow allows for tight independent releases, but not chaotic, with proper add'l structure/scaffolding to satisfy that list above. Microservices and smaller repos can help with the context window bit initially, but you can rig up and kind of local llm-focused setup to allow for selective context and holistic context (across N repos or N projs within monorepo)
- strategy: use the robots to fix the problems in your PDLC/CICD/ABC so that the robots can help you out more, and keep iterating on that
i also did something similar with a learning game for kids, where moving the car requires typing in the correct keys on the keyboard.
we put together a Scroll on how this manifested with node, if anyone finds it helpful to understand that attack vectors & mitigation steps: https://endash.us/toolkit/items/mini-shai-halud
i don't think there'll be a drop in demand, but with supply chains.. beware the bullwhip effect https://beergame.endash.us/bullwhip-effect
and glad that this led me to OpenKnowledge.. hadn't seen that before