"AI models will surely get better at writing "enjoyable proofs,"" why? Why is that surely true? they've increased in all other capacities at shocking rates while still writing awful, slippery, turgid prose. Very silly to assume that this will just go away.
i think you still don't understand what it means. there's nothing personal about it. the point you pick along the frontier is what's personal to you and your needs. the frontier itself is not.
https://frontierharness.org/ allegedly this is on the Pareto frontier, though I don't know how good of a benchmark this really is. it seems to focus on one-shot type tasks, whereas the real utility of one harness over another seems to make itself known in long running tasks.
also this is a very limited static snapshot with one model as backend, I wish there were more consistently refreshed and diversified harness benchmarks.
if you're paying API prices, you probably already know this. everyone else is using a subscription which is massively subsidized rel API prices. or am I missing a third case?
I wish we could use just the frontend of omp harness-agnostically, like the provider switching, stats dashboard, subscription pooling etc. are excellent features, but it doesn't seem inconceivable that we can have all that with a swappable "actual" harness. I've tried to make omp more pi-like by lazy loading most of the tools rather than dumping them into system prompt. haven't benchmarked this yet.
"When we learn how this curve was obtained that might change." IF they tell us. AFAIK Anthropic never released the reasoning chain for the Jacobian conjecture counterexample, and they might not release anything for this either.
yes of course LLMs are already wreaking havoc by scaling up the bot farms of yore but that's not really how I interpreted hyperpersuasion, I interpreted it as a specifically targeted one-on-one conversation with belief in a specific notion as the desired outcome. I believe that LLMs can logos and pathos-max but (depending on the interlocutor) they cannot overcome by sheer force of logorrhea their lack of reputation, credibility, presence, etc.
"He looked at me incredulously. Surely, a smart person like me should know that AI, or better said, AGI will be hyperpersuasive soon – already on a bunch of benchmarks it exceeds professional debaters at persuasion." I'm always tickled by this tic of this kind of person, that human traits are something that we can maximize to infinity. That if we trained a robot to tell jokes, it would eventually tell jokes so funny we'd die laughing etc. You can't convince someone who doesn't want to be convinced.
Any project I see on HN now, I have no idea if the developer will still be interested in it in a month. I'll stick with ghostty because of its reputation.
the people who crow in the comments of each of these posts about AI advances making human beings useless seem to bizarrely identify themselves with the AI, but none of them seem to have had any hand in building this technology. at best, they're power users. pure ressentiment.
not necessarily disputing any of his conclusions, but I find his style of argument frustratingly slippery, always generalizing broadly from weak evidence
Just ran it against Whisper-Large-V2 on a math lecture (my primary use case for ASR is subtitling math lectures), and it was substantially faster and only slightly worse. Very usable for live transcription though I'll probably stick with whisper for the time being since I don't really need the subtitles to be generated in real time.