12,756 karma · joined September 1, 2012
Engineering at Known (https://known.com)
Co-founder and tinkerer at Supernotes (https://supernotes.app)
Gemini 2.5 Flash Lite is $500/Gt, Jev is $42/Gt. AKA an order of magnitude cheaper.
> BERT with more data
It is specifically not just that, in the same way that models which have been chat/task-optimized via RLHF (which made these models much more useful for a huge variety of tasks) are not just "the base transformer model with more data".
For example, if you feed in some context to Jev and Claude Haiku and say "make the appropriate tool call based on this context", Claude (or any other frontier LLM) will hallucinate tool calls some percentage of the time. Jev will not. While yes, the "will not" is constrained by Jev's (lack of) capabilities in some sense, this is actually a very real need for a wide variety of use-cases people are currently using off-the-shelf LLMs for at the moment.
Probably the better example is the whole probability thing, where even if you use something like constrained decoding to ensure an LLM only outputs a certain schema, and therefore can't hallucinate a class, if you ask for probabilities, the probabilities output by the model are just hallucinations. Jev meanwhile is outputting calibrated probabilities for different choices based on the actual landscape.
- short paragraph saying "we launched neki"
- short paragraph saying why neki was needed
- paragraphs about what neki is (with a header)
That seems like a very reasonable structure for a blog post like this one. Not being able to get to the third paragraph of a blog post seems like a "you" problem, not a blog problem.
This conversation is basically the "it's not happening" > "it might be happening" > "it's happening and it's good actually" meme, though admittedly not by one person.
Basically every movement/organization on earth has activists who support their activities, why would Hamas be any different? There are people everywhere who don't mind violence/killing and there are people everywhere who are anti-semitic. You think the intersection of those two cohorts is empty?
Pun intended?
A conversation with 20 turns, 50k tok growth per turn, 1m tok context at end would price out like this:
Fable 5 ($1/M cache reads) ; cache reads 9.5M tok × $1.00 = $9.50 ; cache writes 1M tok × $12.50 = $12.50 ; output 1M tok × $50 = $50.00 ; total = $72.00
Fable 5.1 ($0.25/M cache reads) ; cache reads 9.5M tok × $0.25 = $2.38 ; cache writes 1M tok × $12.50 = $12.50 ; output 1M tok × $50 = $50.00 ; total = $64.88
So yes, cheaper, but not massively.
- if you're gonna order the rest of the bar chart by rank, order your model accordingly.
- if you're gonna highlight a winner in a table of benchmarks, don't highlight your entire model row in the table.
Etc etc