Yup, and if someone is looking for a non-AI startup idea, take a use case with lots of data and a read-heavy workload, and do pretty much what Turbopuffer did for text and vector data. That seems like a very good way to go.
Pretty much this, I don't just see how ASI could happen with current LLMs without big theoretical breakthrough so that's why 'phasing the frontier' and 'distillation attacks'.
I think it's more about money and scaling, bus factor is pretty big if run very lean organization eg whatsapp 2014 and that point all those developers are kinda your cofounders and probably start asking much bigger piece of pie. With small teams you kinda trade scaling and availability to velocity. It's much easier to ship but running oncall 24/7 with small team is just nightmare.
I think we are starting be on that territory that regular software development is suffering, current models are great for benchmarks and one-shots but in daily development models are too eager and try to force patterns like excessive tests in every turn.
I think automation is coming but it will be way more gnarly than frontier labs want public to believe. Value is just too big, when you can automate most of eg customer support it will create huge savings and same time customer satisfaction will get better.
we aren't anywhere near lights off software factories, and writing software is easiest domain for modern llms, you can verify results rather easily, plenty of training data etc.
Yup, new SOTA models especially with high/xhigh/max reasoning too often overengineer solutions, good for benchmarks that usually measure task completion, bad for normal development where you don't want 'rewrite in rust and 1k LOC unit tests' style solutions when agent does mundane bug fixes.
Mistral just needs to good enough category think all those flash models or Qwen3.8 27b which they sadly aren't at the moment, that plus being European lab will mean that they will have very nice business. Even now these SOTA models feel too overkill for most tasks.
I have noticed that these cheaper and faster models are very great for Ops-work. Luna max is beast when you use some stronger model to write detailed instructions/run book what to do and when to stop.
But it was Google fault that they really cannot productize all that innovation that happened DeepMind. Google probably now fumbled with world models and this exodus will create next giant in that space.
It might be harness or prompting style. Personally I use opencode and my prompting style is very plain and terse . Where tasks are very small. Opus and Sonet too often are too verbose and go tangent. Where GPT5.5 is much stricter.
Generally why build your own CRM? ERP and other resource planning systems I get becouse you can tailor made those to your back office. But for CRM you need mostly reliability.
Difference is pretty big if it’s icy like breaking 100 meters vs 10 meters. Especially if there’s wildlife like reindeers/moose’s you are going to do emergency breathing semi regularly.
It’s more about operational resilience and serving customers than product development. If you run early WhatsApp like organisation just 1 person leaving can create awful problems. Same for serving customers especially big clients need all kinds of reports and resources that skeleton organisation can not provide.