Software developer with years of professional experience designing and building software for startups and businesses. I love talking Technology & Business.
I really like what Framework is trying to do but this story is very hard to ignore. If a failed BIOS update can brick the machine, and the official solution is a CA$500+ motherboard (that's about $350+ USD I belieive), there's still a pretty big gap between being modular and being truly repairable
Not to mention there are no "alternatives" that you can look into to purchase these replacements. You are essentially just locked to purchasing from Framework only
The fact that a 17GB model can do all of this locally is honestly kind of crazy. A year ago this would have felt like something you’d need a very expensive hosted model for, and now it can run on a reasonably specced PC
5.6 Sol looks nice, but the Gemini 3.5 Flash comparison is interesting. It’s cheaper and still came out ahead on detection and counting, which doesn't really give me much of a reason to use Sol since Flash is much cheaper and hence much easier to scale. Not to mention we now have 3.6 Flash too
I dont get why they'd make this available to free users when they're probably dealing with compute constraints considering their competition with the Chinese models. Giving a more capable model to a huge free user base looks like an expensive choice to me
This seems like a more natural use of agents than asking them to one shot and push to prod. Give them a measurable objective, let them experiment, and judge the outcome instead of the implementation. Similar to how you would deal with a junior or intern basically
I think the paper would have been stronger if it acknowledged how quickly the underlying evidence is becoming outdated. AI-assisted development in 2026 isn't just better models. The way many devs including myself work has changed and matured quite a bit as compared to last year
I'm liking the trend of companies are releasing smaller, focused models instead of trying to make one model do everything. A dedicated moderation model is much easier to reason about than hideden safety logic inside a general-purpose model which might not have had much training in that aspect at all
I think the article is right that developers become attached to tools because of trust. Ironically though, that's also why many people drifted away from Stack Overflow. The answers were trustworthy, but over time the experience of asking questions felt less and less welcoming than searching for existing ones
What's most impressive here is that that DeepSeek keeps showing how much of a performance improvement can come from post-training alone while the architecture stayed the same. It's a strong reminder that we're probably underestimating how much optimization is still left after pretraining
One thing I hope becomes more common after this incident is companies publishing what they would change even when the root cause wasn't their software. Anyways, I appreciate articles like this because it's easy for companies these days to say "we weren't the vulnerable component" and just move on
This is a good reminder that every software revolution eventually becomes a hardware and infrastructure problem. Someone still has to build the buildings, wire them, and keep them running
It feels like we're moving away from "bigger context is always better" toward "right-sized context". I love this not just because my wallet feels safer but because most of my coding sessions never come close to needing 1M tokens anyway
Thanks for sharing this. The personal observations were far more interesting than another discussion I read about hardware specs. I'd love to read more long-term experiences like this
I like the approach of building on existing standards instead of starting from scratch. That said, solving communication between users is a very different problem from replacing a system that has to work with literally everyone
The irony is that Google's success was built on crawling and indexing the open web. I understand wanting to protect your product, but once you remove affordable APIs and then object to third parties filling that gap, you're creating demand for the very behavior you're trying to discourage
I suspect the debate shouldn't be LLMs or no LLMs, but rather what level of human accountability is required.
We've accepted compilers, static analyzers, and code generators because the maintainer is still responsible for the final result. The interesting question is whether LLMs fundamentally change that responsibility, or just change the kinds of mistakes reviewers need to look for
I'd be interested to see whether this comes with a replacement for legitimate use cases. Removing a capability without offering an alternative tends to push developers toward even more fragile or in some cases "rule breaking" workarounds
I like that the conversation is moving beyond prompts, but I also hope the underlying principles stay portable. If "context engineering" becomes glued with one vendor's tooling, it'll be harder for developers to build workflows that transfer across models.
But well knowing these big AI companies that probably is their goal all along. To lock us to themselves
I wonder how much of this is handwriting versus simply removing distractions. A notebook doesn't have notifications, tabs, or an autocomplete constantly trying to finish your thoughts. I hope to to see studies comparing handwriting with typing on a distraction-free device
Cases like this are heartbreaking, but they're also a reminder that failures like this ones need to be published just as prominently as success. Gene editing is still a young field, and if negative outcomes remain hidden, other researchers can't properly assess risks or make improvements
I agree with this, but I'd add that good judgment comes from experience. Before you can evaluate whether AI-generated code makes sense, you need to have written enough code yourself to recognize the trade-offs and failure modes. AI is an incredible productivity multiplier, but it can also make it much easier to move quickly in the wrong direction without even knowing that you are moving in the wrong direction
I've a local dictation workflow for coding, and one thing I've learned is that transcription accuracy is only half or even less than half the problem now. The other half is latency. Once the delay gets low enough that you stop noticing it, voice input starts feeling much more natural. It'll be interesting to see where this lands compared to Whisper-based setups for continuous dictation