936 karma · joined August 27, 2012
The latter going much further than most mainstream languages.
Google's AI offering is a complete nightmare to use. Three different APIs, at least two different subscriptions, documentation that uses them interchangeably.
For Gemini's API it's often much simpler to actually pay OpenRouter the 5% surchargeto BYOK than deal with it all.
I still can't use my Google AI Pro account with gemini-cli..
Edit: it seems this is a hosted version. Would be nice if they actually joined up some of their products.
For example, Google's inexplicable design decisions around libraries and APIs means it's often worth the 5% premium to just use OpenRouter to access their models. In other cases it's about which models particular agents default to.
Sonnet 4 is extremely good for tool-usage agentic setups though - something I have found other models struggle to do over a long-context.
https://www.newstatesman.com/politics/2025/07/the-british-we...
This implies it's not a hybrid model that can just skip reasoning steps if requested.
Anyone know what else they might be doing?
Reasoning means contexts will be longer (for thinking tokens) and there's an increase in cost to inference with a longer context but it's not going to be 6x.
Or is it just market pricing?
It's not clear but are there cases where this could be a significant price rise? If you exclusively had small objects (<512kb) being written and read then this could add up quickly.
Not to take away from their work but this shouldn't be buried at the bottom of the page - there's a gulf between completely new models and fine-tuning.
And the switch to zonal energy pricing will likely have a similar effect for other sources of generation.
Or an AI model pretending to be an expert in the field... (works well in a few niche domains I have used this in)
I'm pointing out that nearly every thread covering Deepseek R1 so far has been like this. Compare to the O1 system card thread: https://news.ycombinator.com/item?id=42330666
Very different standards.
OpenAI literally haven't said a thing about how O1 even works.
O1's reasoning traces aren't even shown, are you suggesting they've somehow exfiltrated them?
There are some significant innovations behind behind v2 and v3 like multi-headed latent attention, their many MoE improvements and multi-token prediction.
e.g Tests I want applied to anything retrieved from the database. What I'd like is to optimise the prompt around those (or maybe even the tests themselves) but I can't seem to express that in DSPy signatures.
It would be interesting to know the outright prices for those systems as well as their hourly rental rates at the moment.
(Page 16, 57A)
"A company must not be registered under this Act by a name that, in the opinion of the Secretary of State, consists of or includes computer code."