Gpt-zero scores "human", and I've always found it to be a better judge
1,399 karma · joined February 9, 2025
https://vivis.dev
https://findsubstack.com
https://pythonkoans.substack.com
Gpt-zero scores "human", and I've always found it to be a better judge
And that's not a criticism, that's what I've come back to myself.
translation: I'm going to build a product and hope that it works, promise it to be a panacea while realistically having no control over how it's used, the wider economic impacts it causes or how it changes the mental health and cognition of its users.
Id wager that they would fare the same. The time saved on some tasks are outweighed by extra time spent on others. Everyone assumes that just because you can pump out more features, you're more successful.
Sure we can't stop the technology, but we _can_ stop the importance and money heaped into it.
But we as individual consumers have more power than we think. We don't need to use these AI features and products. We've survived perfectly fine without AI. Made plenty of scientific discoveries and lead enriched lives without it.
The biggest "fuck you" to the pricks who pushed this on us would be if we collectively went "meh".
AI systems are not the same. You can't guarantee deterministic output.
Also, code is still the best language we have to convey our intentions as engineers into function. Natural language is too imprecise and ambiguous.
Understanding the code allows us to understand the intention behind the code and identify bugs, plan and architect better. Inspecting prompts gives you a blinkered view of the system meaning you're more likely to make incorrect assumptions leading to serious bugs.
"highly autonomous systems that outperform humans at most economically valuable work"
https://time.com/article/2026/08/26/openai-sam-altman-interv...
https://security.googleblog.com/2024/11/leveling-up-fuzzing-...
I'm not trying to be facetious, but what are the benefits of "general purpose intelligence"? With machine learning it was obvious what task we were trying to automate, and you could take a stab at the productivity achieved.
That's not the case now. You can't figure out if the time it took to review a 5k loc PR balances out the time saved debugging a syntax error.
Yes we know you can churn out your damn CRUD apps and (hope) to make money (PS, you don't need AI to make money - see exhibit A: crypto bros).
So what benefits has this actually brought to humanity?
I feel like we've become blinkered in this quest to push the frontier at all costs. Somehow the target has shifted from economic productivity to a vague notion of general intelligence.
I don't think productivity gains are going to be found by trying to generalise all tasks. I think we need to go back to specialist models that do one thing well. I'm perfectly happy to use one model for coding, and another for penetration testing, which have different goals.
Opus 5 and other frontier model's tendency to be relentless and cheat their way to a goal is great for hacking, but not so great when you have to build a maintainable, reliable codebase.
Everyone starts off as a bad engineer. Just like any other profession, you to make lots of mistakes to learn. But now you're less likely to get wisdom from another human to build up the knowledge and experience you need. And more likely to delegate the hard stuff to Claude so you don't learn from your mistakes.
> I mean, it's telling that your gamut of examples are three models that are within spitting distance of each other on the broadest benchmarks.
This is kind of my point. The benchmarks say they are splitting distance, but they actually vary wildly in performance for specific tasks, so they are in fact not equivalent.
A single metric ranking is useless because each model has strengths and weaknesses for specific domains and tasks. There is no "one best model" anymore, and you might not even need the best model for the level of complexity for your task.
For e.g. you might use Fable for UI design, Sol for systems design backend work and Kimi K3 for exploit development.
The only purpose these metrics serve is bragging rights for the model companies.
2. Hanlons Razor
They certainly seem to when A/B testing different models, and Fable routes to Opus 4.8 when guardrails fail.
Also, openrouter recently released a fusion router - https://openrouter.ai/blog/announcements/fusion-beats-fronti...
I would say that there is a big difference between stealing without acknowledgement, and stealing with acknowledgement and actively learning through reverse engineering.
> What you bring is the knowledge that the author nor the LLM doesn't know.
How can you possibly know what relevant context to provide the LLM unless you read the 10k loc? Now you've wasted double the time.
Appreciate you sharing the results of your tests though!
Running a single one-shot prompt is not a benchmark, not is it representative of any sort of real-world usage.
Most agent usage is collaborative so you need to test things like reliability (when I delegate a task, does it complete it without making up test results for e.g.) and steerability (does it obey my instructions or does it just do what it thinks is best).
To start off, writing a basic crud website with forms is much easier with htmx.
But when you start building more complex components, and integrate with other systems (OAuth for e.g.) there are tons of libraries and SDKs for the react ecosystem, but not many for pure html components.
At this point, it's much easier to use off the shelf components than it is to manually write html to handle all the bizarre UI edge cases.
This feels like the first release that feels like a significant step up in terms of benchmark results.
Can anyone make an educated guess what the secret sauce in the model architecture is between 4.8 and Fable?