5,740 karma · joined December 1, 2012
Just use Fable 5.1/Opus max for the hardest problems, GPT Sol high as your workhorse, and maybe terra for async batch stuff you don't really care about. Gemini 3.8 High also looks pretty good and is quite fast if you're already a GCP shop. You can basically benefit from open models without using them because they force the frontier models to be cheaper.
It's not just malicious software ads, they also sell ads to entities mitming every government service.
"Buy me the best Air Fryer" -> finds the best given my preferences and constraints and directly purchases it from the cheapest distributor
That said, I'd really like to see data to compare which approach works the best.
Maybe a long term play would be putting this out and creating a "graphics bench" to entice the labs to overfit on your DSL but that seems like a lot of work
You should license the backend to enterprises that want a cheaper/lighter weight option to kubernetes for the surge of agentically coded internal apps.
0) big companies already are very comfortable using contracts to trust other people with their data. Maybe if they're inflexible on the 30-day requirement for fable some orgs will opt out but by-and-large it's already happening and it's not a blocker.
1) Cost will be a blocker. The level of token spend is untenable and the pareto curve is flattening. Most orgs are going to default to either using a distilled model from China or a distilled model from the model companies (e.g, Sonnet 5). It'd behoove Claude/OpenAI to offer a model router before another vendor wins that area.
2) Karp is selling his book. No one knows or cares what an 'ontology' is. From what I can tell, company's product is a tool that helps governments bomb people.
This should probably be required - there is a different mindset and set of restrictions when you're expected to pick up a page. It also forces companies to use on-call judiciously - not every service needs a 5 min SLO.