Why wouldn’t the AI be able to answer questions about the proofs or organize a talk? Those are pretty much the things it excels at, exploring and summarizing previous knowledge. The problem is the volume of questions and data a human needs to step through to arrive at the same understanding.
herdr might be the culprit there, it can significantly slow down tui apps with its terminal capture. This happens on mac/linux, WSL will surely make it worse.
Off-topic, but what a horrible mobile experience on this site. The screen is >50% covered in ads, and both close buttons are a trap that let clicks through and open a new tab.
Those arrows also initially align perfectly on top of the featured photo, making it look like a gallery, but actually navigates to another article (more ad views, yay).
I wish these would stop using JevBench. It focuses way too much on text classification tasks, and some of the models perform very poorly on tasks that need actual intelligence.
These are pretty much what a human would draw. Sun rises from the east. A cloud makes the background “sky”. Three lines is the minimum to interpret as movement. Two feathers is standard on every cartoon and illustration.
You could say so from a consumer perspective, but it is not what the MTBF you quoted measures. Read errors are not considered a hardware failure. Not an opinion.
That's a lot of work to avoid accepting your previous comment was incorrect.
This practice of having the one provider should be eliminated. Companies self-inflict lock-in to large platform providers, prevent their own teams from using better technology options and stifle innovation. It's crazy that even with a pile of SOC/ISO/PCI/HIPAA/NIS certificates, procurement is still a months-long process, it should be much easier to do business.
Plan mode has become pointless since Opus 5 came out, they know when to switch between planning and execution now. But that iteration/discussion is still necessary unless you're building completely blind - the model cannot read your mind.
I've had it running 8h+ of non-stop optimizations, chasing a performance target, rewriting systems or building a series of prototypes for research. All it needs is a clear goal.
So.. they continue having access to private identifiers, while you willingly give it up to "protect privacy"? Piping all of that data into a massive central database instead of your nginx logs? How is this supposed to be better?
Because with a tiny model you're skipping all the intelligence and world knowledge that makes it useful without fine tuning. `typed-decisions` is almost entirely text classification tasks.
In my experiments decider-4B performs better than Kev with significantly lower latency. It's remarkably good for it's size, shame it wasn't included in the benchmarks. Laya on the other hand shouldn't even be featured - despite being 'the original' decision model, it can only do simple text classification and is nowhere near usable performance for anything else.
I enjoyed Pi for a few weeks, but ultimately moved to other harnesses - the plugin ecosystem became a sea of slop, large vibecoded projects that don't work at all, and yet have thousands of stars. At some point I gave up trying to get subagents working.
This is the case since Opus 5, the latest models (from all providers) favor using shell tools instead of the View/Edit tools available in the harness, and “accept edits” doesn’t let those calls through. Auto mode is the best option.
You might be overusing subagents. Especially with a chatty model like DS, you’ll be wasting millions of tokens on re-discovering the project and facts instead of actual reasoning.