This benchmark expects models to not only do the happy path, but also edge cases, even though the ticket they give to the model doesn’t specify if. This to align more to real world tickets. Usually with benchmarks it’s the other way around: only implement what’s asked. So I welcome this type of benchmark because this is how I use the models.
Also, handing over the work to AI robs you of learning and understanding. Let’s say you write a ticket for a feature and that description is good enough that an AI agent can implement it from start to finish. The AI will then discover things while implementing the feature, and will use those learning to make the feature work. That learning will then be discarded once the AI is done. No one will be able to partake or share the knowledge with others. Sure, some of it can be saved in form of comments, but not everything.
I don’t see why they just don’t allow a smaller model to answer the question while letting the bigger one vet it. The vetting can be asynchronous and can be delivered after a few seconds (if it’s an easy query). If it’s a hard query, the UI can show the answer is currently being vetted or something.
What surprised me is that caffeine gets converted to other compounds that _also_ binds to the same receptors thus causing you to feel less tired. And that compound has its own half life. Now I’m just guessing, but it could be that one is more affected by one of the byproducts so to speak than the caffeine itself.
Why conscious thought as described in the article was thought to be signs of a delirious person was because that’s where you could find it in real life, is my guess. If someone blabbing about without a listener, they’re insane. So to describe someone’s inner thought would render them insane to the reader. But with psychology came greater understanding and value of the inner thoughts and now a character no longer was deemed insane by having innate thoughts.
Finally an alternative to the big dogs that a company can use. People have been asking for a way to run the Chinese models from a trusted provider. Here GitHub delivered!
The performance, if we trust the benchmarks, put it at Sonnet 4.6.
The Vibe CLI is really bad on Windows, sure they don’t officially support it, so can’t blame them, but a FYI for anyone wanting to try it. It can’t get find and replace right.
Oh man, this was the release I wanted to link, as it has a new feature (tiled follow link) that I actually started using right away. A new browser feature I find useful didn’t happen that often for me, so I got excited.
I have this nagging feeling I’m more and more skimming text, not just what the LLMs output, but all type of texts. I’m afraid people will get too lazy to read, when the LLM is almost always right. Maybe it’s a silly thought. I hope!
One of the more annoying software that does this is the copilot Office 365 on the web. Every time (!) I open it, it shows a popup on how to add files to the context. That itself would be annoying, but it also steals focus! So you would be typing something and suddenly you’re not typing anymore for M$ decided it’s time for a popup.
I finally learned to just wait for the pop up and then dismiss it with esc. Ugh!
I built this recently. I used nvidia parakeet as STT, open wake word as the wake word detection, mistral ministral 14b as LLM and pocket tts for tts. Fits snugly in my 16 gb VRAM. Pocket is small and fast and has good enough voice cloning. I first used the chatterbox turbo model, which perform better and even supported some simple paralinguistic word like (chuckle) that made it more fun, but it was just a bit too big for my rig.
Gave it four of my vibe questions around general knowledge and it didn’t do great. Maybe expected with a model as small as this one. Once support in llama.cpp is out I will take it for a spin.
I’ve tried the voice clinking and it works great. I added a 9s clip and it captured the speaker pretty well.
But don’t do the fake mistake I did and use a hf token that doesn’t have access to read from repos! The error message said that I had to request access to the repo, but I’ve had already done that, so I couldn’t figure out what was wrong. Turns out my HF token only had access to inference.
I’ve recently bought the LG with 4th generation OLED, and for me that works for long coding sessions (I use it for work). They shifted or did something with the pixel arrangement for this generation just for text legibility.
Interesting experiment. I would hazard to guess that Google is on top when it comes to these sorts of things (spatial ability), then OpenAI and last Anthropic. I would like to see the same experiment using Google’s Live view or whatever it’s called in their Gemini App.
“ Starting at approximately 16:00 UTC, we began experiencing DNS issues resulting in availability degradation of some services. Customers may experience issues accessing the Azure Portal. We have taken action that is expected to address the portal access issues here shortly. We are actively investigating the underlying issue and additional mitigation actions. More information will be provided within 60 minutes or sooner.
This message was last updated at 16:35 UTC on 29 October 2025”
Nice idea, I’ve been toying around the idea of consuming news only once per day. But for me I think I want an actual newspaper with in depth articles rather than short news posts from online news.