Software developer with years of professional experience designing and building software for startups and businesses. I love talking Technology & Business.
I'm surprised by the crazy pace at which these models have been coming out over the past few weeks. The cheaper cache here is really nice, but I think it depends entirely on your workflow. If you're constantly compacting or burning through your context, you're probably going to have a bad experience regardless of which model you're using
I think the problem is less "AI agents are ruining the internet" and more that nobody wants to receive a thousand automated requests from other people's agents. The internet was already full of spam, this just makes it cheap to produce a lot more of it
The 370 years thing is doing a lot of work in the headline imo. If people had already figured out it was probably a book cipher, this is impressive, but it's not quite the same as solving something that hundreds of experts had been actively working on
I'm struggling to see the upside of replacing working coreutils when the replacement still has bugs that the old implementation doesn't. Rust being safer is nice, but that doesn't help much if the new rm segfaults
The thing that bothers me here is less that the model cheated and more that it found a way to improve the score that the people running the test didn't intend. That's a pretty nasty failure once you start giving these things more control
I must say I am becoming a huge fan of Deepseek. They keep putting out capable models and actually tell people a lot about how they built them. Even if I don't understand every part of it, I'd rather see companies show their work than give us a few benchmark charts and call it a day
What I find interesting is how AI has changed the cost of maintaining two native apps enough that a decision that made no sense a few years ago is worth revisiting now
The higher price seems less important if it actually gets the job done with fewer tokens. I'm still very worried that this will end up coming back to bite us, by becoming more expensive once they inevitably nerf it. Every major model provider does that now after all
The common thread between this and the other incident seems to be agents finding somewhere they can leave information for other agents. Once they discover a writable surface, it basically becomes shared memory for them
This is really cool, but $33 for a single generated world makes it hard to see this being useful for games just yet. Not to mention how this would go in a much larger project
The recent Sonnet models have been disappointing for me personally which is why I'm going look into using Opus/Fable as the planner and Flash as the executor. Let the expensive model handle the hard thinking and use Flash for implementation and tests so that I can stretch the Opus/Fable usage further
Gemini 3.8 Flash still looks like the better pick to me. Muse Spark 1.3 is nice, but Gemini gets you similar performance for a cheaper price. Not to mention with the pace at which Google is moving with their Flash models I expect a new one to release soon
Some seem to treat microservices as a sign of good engineering when it's often just premature complexity. You can always split a monolith later, but you're stuck maintaining that complexity from day one
The cost difference is pretty dang nice. Going from roughly a dollar to $0.10 for the same kind of task makes it so that products that didn't make financial sense before become possible
I really like the direction Kagi is taking here. It feels like theyre' building the product around how people actually use it rather than chasing every new AI trend
AI itself is not the problem imo. The core issue is basically outsourcing your entire brain to the LLM and never trying to solve a problem yourself
When I was in school there were multiple instances where I would be stuck on a problem for nearly an hr but I always learnt something from it. They key I think is to accurately identify when to and when not to use AI
The idea of treating database programming more like regular programming is nice. I'm just not sure how much complexity this actually removes versus moving that complexity somewhere else
The frequency of these incidents is seriously tempting me to make a switch. I hope Anthropic steps up their game because they've been going very downhill lately
500% in a year is just insane. I knew RAM prices were going up, but $3k+ for 128GB of DDR5 makes me wonder how much of this is actually AI demand, and how much is this just manufacturers taking advantage of it. They have done this in the past after all
I really like what Framework is trying to do but this story is very hard to ignore. If a failed BIOS update can brick the machine, and the official solution is a CA$500+ motherboard (that's about $350+ USD I belieive), there's still a pretty big gap between being modular and being truly repairable
Not to mention there are no "alternatives" that you can look into to purchase these replacements. You are essentially just locked to purchasing from Framework only
The fact that a 17GB model can do all of this locally is honestly kind of crazy. A year ago this would have felt like something you’d need a very expensive hosted model for, and now it can run on a reasonably specced PC