328 karma · joined October 9, 2018
I don't want to drag the discourse away from this achievement, but I hate how this is announced with blatant corporate advertising (our internal model, here are the benchmarks, gpt astra TM yours now for the low low price of £200pcm). I just didn't think Navier-Stokes falling would be sponsored by McDonald's.
Still. I am crying right now. Navier-Stokes is solved.
What I have found is that getting a model to rewrite a badly written passage is hard, because it seems to key off what it reads. It might swap some vocabulary around ok, but it doesn't fix structures very well. So getting it close to the preferred style in the first place is better.
To take this further, if you must fix existing bad prose, write a clean-prose skill which extracts the bare structure of the prose with none of the style, hands it to an author subagent who isn't poisoned with the original bad prose, then hands the output to a reviewer subagent.
Opus 5 writing is horrendous, so I have been experimenting with improving the output!
Both the article and the parent comment treat "happens to have Kik account" as an independent discovery that affects our Bayesian inference.
No. The innocent was identified exactly _because_ they have a Kik account, so the conditional probability they have a Kik account is 1.
For example, a child learns that "foreigner" means someone from outside their country. Then, when they're 11, they go on their first holiday abroad and realise "Wait! _I_ am a foreigner here!"
So, maybe one way to frame what you're saying, is that LLM output tends towards being easily assimilable.
I have a degree in mathematics followed by twenty years in software development (pillory me, if you like). My conclusions after 6 months of using LLMs every day are remarkably similar to the author's. I increasingly think in shapes, architectures, data structures and ideas; less and less in lines of code.
And I fully understand his point that, once the architecture of an idea is settled, reading LLM code does not feel worse than reading human generated code. Especially if you have a strong style and conventions guide.
The idea is the hard part, and it's the right place to focus your effort.
- write gherkin features for new features; update them for enhancements; don't touch them for refactors. Label your PRs with these nouns.
- use pre-push hooks for type checks, linting, unit tests, and other quick, scriptable validations.
- make a viteperess subsite in your repo, have the agents maintain it - document important principles, architecture, etc.
- make a cli command which lists all pages along with the yaml frontmatter description so agents can choose what to read without blowing up the context window.
- use ddd and monorepo - write your logic in headless layers, and compose layers into apps. agents navigate layers very successfully.
- use zod (or your language equivalent) and contract-first API development; this is my favourite bit tbh, I use orpc
- make a single skill called "code" which describes the lifecycle: open a worktree, setup .env to guarantee no conflict with other agents (choose unused ports etc - docker is good here), write or update feature file (this is where you negotiate the spec), implement, validate (e.g. using playwright mcp), pre-push checks, push and wait for review, tear down and fast forward main
- testcontainers is great for ensuring multiple agents can run tests that don't conflict
Seriously I only have one skill that's it. Everything else is in the docs. I'm feeling very productive like this, in a "making good software" sense not a LoC sense.
For each new feature, I open a worktree, spar with Claude to work up a gherkin spec with @todo on each story. Each agent pushes commits to a WIP PR in GitHub where I review and leave comments or questions. Once the spec is done we mainly interact on the PR. @todo becomes @wip and @done as the agent progresses. I really like gherkin for agentic engineering, it's very clarifying.
I have about 2-4 agents running at a time. Large test suite, linters and formatters enforced on push.
I'm confident this is human.
Far from being motivated by some applications, the most useful discoveries in mathematics are usually discovered "for their own sake" and their application is only discovered later. Sometimes centuries later!
The fact that the derivative of this accumulator function is equal to the original function, this is the fundamental theorem of calculus, and I violently agree with you that this part is shockingly, unexpectedly beautiful
A "draft" is a row in a database with live preview. Users can click a button to make a checkpoint (git commit, by GitHub API, but they don't know that). When they click "publish", the PR for their draft is merged.
Writers in my team can use a nice Tiptap editor with custom components. I get the change management of git.
The API for reading content and editing drafts is also exposed over MCP meaning AI can collaborate in the authoring process from anywhere that can connect to MCP.
Loving it so far.
I just don't think "instant feedback" is as important as we think in mathematics education, and might even rob us of moments to practice mathematical behaviours like justifying, communicating and accommodating. Slow feedback does have benefits.
I am a tech enthusiast to put it mildly. I also taught maths in schools from roughly 2010 to 2020 so saw the iPad/app revolution in my classrooms. Anecdotally, I think it made my lessons and my students worse. Books, paper and each other are the best tools (in my very personal opinion).