Is there some sort of enumerated list somewhere that we can run as a test suite and then we can make a bigger deal about the percentage of that list that we're burning down as these models improve?
309 karma · joined July 6, 2017
Is there some sort of enumerated list somewhere that we can run as a test suite and then we can make a bigger deal about the percentage of that list that we're burning down as these models improve?
Its funny how fast we forget that when Google came out and arguably now many would laugh at the statement "all the good, companies like Google brought to the internet"
Text book reading in this course was 10-15% at baseline ... but this AI thing got 90% voluntary usage ungraded.
Even if its worse per-hour than a textbook, you're now teaching 6x as many students _something_ instead of teaching a small minority everything.
So really it just becomes an optimization problem at that point because most students are at least in the funnel/in the running to learn something.
The paper kind of proves this itself ... they tweaked the quize formats mid-semester and where able to iterate which you can't do on a textbook that nobody opens in the first place
I guess its also quite interesting that how they are framing these projects are opposite from how people currently perceive them and I guess that may be a conscious choice...
specifically, the GPT-5.3 post explicitly leans into "interactive collaborator" langauge and steering mid execution
OpenAI post: "Much like a colleague, you can steer and interact with GPT-5.3-Codex while it’s working, without losing context."
OpenAI post: "Instead of waiting for a final output, you can interact in real time—ask questions, discuss approaches, and steer toward the solution"
Claude post: "Claude Opus 4.6 is designed for longer-running, agentic work — planning complex tasks more carefully and executing them with less back-and-forth from the user."
With Codex (5.3), the framing is an interactive collaborator: you steer it mid-execution, stay in the loop, course-correct as it works.
With Opus 4.6, the emphasis is the opposite: a more autonomous, agentic, thoughtful system that plans deeply, runs longer, and asks less of the human.
that feels like a reflection of a real split in how people think llm-based coding should work...
some want tight human-in-the-loop control and others want to delegate whole chunks of work and review the result
Interested to see if we eventually see models optimize for those two philosophies and 3rd, 4th, 5th philosophies that will emerge in the coming years.
Maybe it will be less about benchmarks and more about different ideas of what working-with-ai means
For developers, academics, editors, etc... in any review driven system the scarcity is around good human judgement not text volume. Ai doesn't remove that constraint and arguably puts more of a spotlight on the ability to separate the shit from the quality.
Unless review itself becomes cheaper or better, this just shifts work further downstream and disguising the change as "efficiency"
In that situation saying "i resolve problems non-violently every day" stops being relevenat. The mechanisms that allow you to do so (enforcement, law, etc) have been removed as they were for those fighting for civil rights.
You may still personally choose non-violence in this case, but I'd bet you would understand/sympathize/maybe-even-join those who decided to break into their apartments by force and grab the things that are rightfully theirs.
nobody is secretly violent ... just normal peaceful channels stoped working.
Recognizing that distinction isn't justifying violence its just explaining why nonviolence provides leverage in the first place
History obviously shows that that "moral audience" was certainly the minority then.
MLK was already forcing that confrontation and by most accounts was succeeding slowly-but-surely. But it wasn't until his assassination that people were forced to confront the contrast he had been trying to illuminate all along.
Even his disciplined non-violence he was met with brutal force (as were the peaceful protesters) and this forced some sort of moral reckoning for those who had deferred or were complicit
https://www.youtube.com/watch?v=YKnJL2jfA5A&feature=youtu.be
His strategy worked because it existed alongside MANY other voices, IMO the most underrated of which is Malcolm X, that rejected this "gradualism" outright and refused endless delay.
They weren't organizing violence but they were instead making it credible that there is a world where those "peaceful" people do not accept complicity or "no" for an answer.
This shifted the baseline of what a "compromise" could look like (as we today see baselines shift very frequently often in a less just direction)
Seen that way, nonviolence wasn't just a moral stance, it was one side of a coin and once piece of a broader ecosystem of pressure from different directions. King's approach was powerful because there were alternatives he was NOT choosing.
You cannot have nonviolence unless violence is a credible threat from a game-theory perspective. And that contrast made his path viable without endorsing the alternatives as a model
They weren't primarily organizing armed revolt.. it was more about the idea that they were articulating moral clarity. They were, in the most credible way, refusing to accept endless delay.
This allowed them to shift the baseline of what was politically tolerable.
In that sense, the movements worked collectively because of a kind of good-cop/bad-cop dynamic. MLK JR offered a path to reform that felt (to some) constructive and legitimate _because_ there was a visible alternative that many people udnerstood as worse.
I think violence is already far to prominent today, but I think successful movements do need both moral persuasion (if morality is still a thing that persuades) and _also_ a credible way of making inaction feel unsafe.
Presumably the harness cant be doing THAT much differently right? Or rather what tasks are responsibilities of the harness could differentiate one harness from another harness
From my personal experience I find cursor to be much more robust because rather than "either / or" its both and can switch depending on the time or the task or whatever the newest model is.
It feels like the same way people often try to avoid "vendor lock in" in software world that Cursor allows freedom for that, but maybe I'm on my own here as I don't see it naturally come up in posts like these as much.
Anyway I think this would be an amazing thing to let other people contribute to as this is an entire industry of hypercasual games which could easily be ported to this minus the annoying ads
We already delegate accountability to non-humans all the time: - CI systems block merges - monitoring systems page people - test suites gate different things
In practice accountability is enforced by systems, not humans.. humans are defintiely "blamed" after the fact, but the day-to-day control loop is automated.
As agents get better at running code, inspecting ui state, correlating logs, screenshots, etc they're starting to operationally be "accountable" and preventing bad changes from shipping and producing evidence when something goes wrong .
At some point humans role shifts from "i personally verify this works" to "i trust this verification system and am accountable for configuring it correctly".
Thats still responsibility, but kind of different from whats described here. Taken to a logical extreme, the arguement here would suggest that CI shouldn't replace manual release checklists
Teams generally don't keep merging code that "doesn't work" for long... prod will brake, users will push back fast. So unless the "wrongness" of the AI-generated code is buried so deeply that it only shows up way later, higher merged LOC probably does mean more real output.
Its just not directly correlated there is some bloat associated too.
So that caveat applies to human-written code too, which we tend to forget. There's bloat and noise in the metric, but its not meaningless
AI removes boredome AND removes the natural pauses where understanding used to form..
energy goes up, but so does the kind of "compression" of cognitive things.
I think its less a quesiton of "faster" or "slower" but rather who controls the tempo
People want things to be simpler, easier, frictionless.
Resistance to these things has a cost and generally the ROI is not worth it for most people as whole
Would be nice to add something like that... i think review mode is a huge reason why chess.com puzzles / chess.com is so popular because you leave a little smarter than you came
If an LLM were acting as a kind of historian revisiting today’s debates with future context, I’d bet it would see the same pattern again and again: the sober, incremental claims quietly hold up, while the hyperconfident ones collapse.
Something like "Lithium-ion battery pack prices fall to $108/kWh" is classic cost-curve progress. Boring, steady, and historically extremely reliable over long horizons. Probably one of the most likely headlines today to age correctly, even if it gets little attention.
On the flip side, stuff like "New benchmark shows top LLMs struggle in real mental health care" feels like high-risk framing. Benchmarks rotate constantly, and “struggle” headlines almost always age badly as models jump whole generations.
I bet theres many "boring but right" takes we overlook today and I wondr if there's a practical way to surface them before hindsight does
Results: Claude: ~10s, perfect working demo ChatGPT: ~20s, solid solution Grok 4: ~1000s, failed completely, gave me a truncated base64 blob
This wasn't some obscure edge case... it was basic data visualization that any decent model should handle. Yet somehow Grok 4 is "competing with humans" and has "99% tool accuracy"...
I don't buy it..
links: Claude: https://claude.ai/share/7a413a6a-5c01-44a1-aaed-8b237e5e9e94 Chatgpt: https://chatgpt.com/canvas/shared/687a9f9d4304819187ac7d98d3... Grok 4: https://grok.com/share/c2hhcmQtMw%3D%3D_20b61291-e1bb-45e5-a...
These benchmarks are either just wrong or measuring something completely divorced from practical utility imo...
What insiders are controlling chess?
We've added a lot of features like tags/labels, integration with tracing: https://github.com/pyroscope-io/otel-profiling-ruby, integration with CI/CD (in rails), etc.
Feedback very welcome!
For example, I created an issue requesting a flamegraph visualization in grafana[1] and now it makes sense that they didn't initially respond because they were building it internally in secret and didn't want to spoil the big reveal (when they did respond they did mention that it was a secret).
They're also less incentivized now to tend to issues and PRs that help others outside of their ecosystem (i.e. competing logs, metrics, tracing, profiling, etc products).
Which TLDR is they are reporting agents for each respective language that send profiling data to our server.
I'd picture the bounty as us just giving an input and an output in the form of a set of unit tests and just saying: "if you create a java agent that passes these tests we will pay you $X"
Seems "mechanical" enough right?
My suspicion is that people contributing to open source don't really just contribute randomly. It is most likely that they use a library or repo and then they end up on the "issues" page because they had a problem themselves.
It's at this point where I would imagine one is most likely to convert into actually picking up an issue and when notice of a "bounty" might push someone over the fence (?)