1) verify identical query results
2) run repeatedly to get average, worst, best, etc duration of runs
Sped up so many legacy things that none of us were ever going to bother with.
Source: C. Northcote Parkinson, "Parkinson's Law and Other Studies in Administration" (1957)
I’m assuming he’s using it to mean “cruft” or “low value features” but maybe there’s some meaning I’m missing.
I originally heard it in the context of picking paint color for the bike shed, which is probably more on point for software engineers, most of whom would not be involved in any discussion around building a nuclear reactor.
People will recognize a glimmer of something they can attach to, prompt and get a few pages of junk about this narrow topic, and then contribute it in a way that further reduces the signal-noise ratio of the discussion or project.
Like, dude, we have the same tools. We can get the same slop from the tap at any time.
Can only assume english might not be your first language.
Thankfully until this thread I’d never seen anyone use the term this way.
> Can only assume english might not be your first language.
This not-subtle insult does not add anything here.
An insult would suggest it's also related to autism.
Is the harness for the AI set up with any sort of instructions to generally prefer broadly-applicable first-principals query performance analysis over unintuitive results that may be local maxima due to certain things specific to that env?
Even with humans I've had to unwind "optimizations" before that worked great for low-volume envs by taking advantage of 'be inefficient in the small scale with memory to save CPU and wall clock time' or such in a way that caused pain deployed.
AI doesn't have enough senses yet. It's trapped in a box.
I can imagine the quality of bug fixing by unsupervised (we save money by layoffs) halucinating AI at current state of development.
"99 bugs in the code, 99 bugs in the code... Fix one bug, compile it again, 101 bugs in the code!"
(To the tune of 99 Bottles Of Beer)
But hey, one can't survive 1 week of step-by-step debugging without 99 bottles of beer.
If you want to learn darts or perfect parallel parking for example, most of the increase in accuracy comes from simply but very deliberately pointing out to your brain where you wanted to land versus where it did land.
We like to usually just do things and hope for the best. Defining success is not something we automatically do and naturally we don’t do it with AI either.
Also let me ask you why we need better and better and models if what we have already can produce good output with 'all the tooling to verify its hypotheses'
This is such a blanket dismissal that I can’t agree or disagree.
Maybe very few of YOUR problems are this way. At least mention some problem domains.
Recent experiences: compiler-related (helpful), UI-related (agree it isn’t testable but the design iteration is quick, easy, and correct), debugging technical configuration problems (useless; I basically have to solve each problem myself before the LLM recognizes it).
> Maybe very few of YOUR problems are this way. At least mention some problem domains.
i did in second part of my comment. why do you think billions are being poured into ai if ai can already do verifable tasks.
“Good” isn’t “perfect” and even if it was, the ability to produce perfect output with all the tooling to verify its hypotheses could still be improved, in time and token efficiency, by better models producing fewer spurious hypotheses, rejecting those it does generate faster, and taking fewer unnecessary steps in confirming its good hypotheses.
I’ve been working on UI component improvements and it was doing a lousy job until i specifically told it to test in a headless browser to validate it works. I think somewhere in an AGENTS.md i have an instruction to “don’t state your guesses as fact - validate findings and results”.
As usual, if you use anything but the best model available I’m going to state that the better ones do better. If you do use the best model available, then I’ll just mention that Fable still has limits and still needs some guidance.
One thing it does not do is deliberately build tests which test nothing at all, or which restate the code under test. I mention this because certain other models absolutely would.
They will mock things to no end. They will flat out REMOVE assertions (saying it's not needed). They can also write test to assert the wrong result.
You have to always review it, it's exhausting.
And I think I know why - neither do most people.
Claude will do some boneheaded things for sure, but it's pretty good about writing tests that are useful, and not removing or modifying tests just because they're in the way.
Claude is pretty bad about assuming that it couldn't have broken a test it didn't know about, as it has often told me "this is already broken on main" which is definitely NOT true.
Imagine you’re blind and deaf and have temporary retrograde amnesia. You “wake up” one moment with a memory of some words in your head like “what is the bug?” …but you don’t recall the context of that question, and nor can you look/listen around to observe the context.
So you don’t know whether you’re e.g. at the office, in front of your computer, in the middle of doing some pair-programming (where, yes, you’d in investigate the bug thoroughly with tools), vs. having a conversation with a colleague over lunch (where the expectation is for you to tap into your knowledge + intuitions to either guess or say you don’t know — not to pull out your laptop.
That’s what it’s like to be one of these LLMs being prompted by some agent harness. Unless the harness injects the proper context into its “recent memory”, it just doesn’t know.
Doing a pass where you just ask the AI to sanity-check the existing tests (against rules like “test against the spec, not the implementation” can also help.
I work on user facing applications, and since the models do not have good taste, testing the UX is essential.
If you spot a bug, usually the model will attempt to reproduce it in a new test case that does cover the actual issue.
2nd line, code review.
Do the first pass with an agent, ask it to bounce back vacuous or tautological tests. Ask it to verify that the tests verify what they claim to. Then read them yourself.
3rd line, mutation testing. If the tests don't actually catch broken code, kill the mutants.
I don't strictly mean junit unit tests.
Tests onli validate the presence of bugs, not their abscence (Djikstra).
I'll also add that tests look at outputs and don't care how those outputs are derived. E.g. code filtering the entire db in memory will be fine in tests.
Tests prove the things you thought of worked, and it often isn't hard to find the likely edge cases such that you have reasonable confidence everything works. You will be wrong from time to time, but not that often. You can prove code correct, but if the proof is wrong (common when a human is doing it), or the spec is wrong (most people have no clue how to write a comprehensive spec) it can still be wrong.
And then there's the endless code duplication, reinventing of existing code and libraries etc.
Tests =/= TDD
That'll get RL-d in, eventually.
Current AI isn't super effective at making the breakthroughs, but it sure is effective at democratizing the ones you can point it at.
That's democratization - shifting power and skills away from those with the wealth to purchase vulns and run offensive security teams. It's still not simple, but it's one of the reasons bigger vendors want to lock down the capabilities of models that present a threat to the wealthy and powerful.
As hardware specs get better, and model providers keep making different classes of mistakes that alienate users, more people will opt for or support local models. I tend to take Schneiers' "Attacks only improve" philosophy and apply it to damn near everything like this. Inference will get more expensive, but IMO cloud based inference is at or near the maximum that users can tolerate, across multiple spectra. There may be a marginal increase in cost, but as the different metrics between what is affordable for rapidly scalable, cloud based inference and the lag of local inference with open models and expensive hardware converge, prices will come down.
I am not the right person to say what the peak is going to look at, but when the companies with (practically speaking) unlimited compute resources and money are starting to take a good hard look, there is going to be a drive towards efficiency. I hear grumbling at work, and the appreciation from my leaders when I show up with a tool that shows how I reduced my token costs, or I keep asking the questions of what the token cost (real, and internal billing) are for tools my peers write are.
I also have absolutely eye-watering personal AI bills that are starting to make buying a better tier of local inference gear look more affordable, even with inflated hardware prices. The current generation of open models are very effective, and paying for multiple $200 dollar a month subs is not feasible longer term, and as one example, in one month, if I was paying enterprise token rates, I would have burned through $8000+ token budgets on one of those accounts, doing actual, practical useful work, it makes more sense to buy some much better hardware than to pay for one-time inference costs.