679 karma · joined August 14, 2015
In coding? Software architecture? Math? General knowledge?
The diversity of model use-cases is so broad that comments about "model x is better" without any context are largely useless
While this model may share much with GPT-style models on the encoder side, it clearly has a different decoder architecture. So is a high-parameter count language model an LLM even when it doesn't have a GPT-style decoder? The definitions are in flux.
For all we know this might be a non-language-generative transformer e.g. a transformer where the decoder produces confidence scores rather than language. Please provide more likely architectures if you know them, I'm genuinely curious.
0.15/0.22 ≈ 0.68, meaning a roughly 32% reduction on inputs. The 50% reduction is only outputs and cached inputs.
Selling a car today gives Tesla the full profits of that hardware production today. Running that car as a robotaxi means that Tesla eats the costs of the hardware today in exchange for a larger total profit collected over several years.
So operating a fleet gives the company access to more long-term profits at the cost of decreasing the bank balance today (negative cash flow), where selling the cars lets the company fill the bank account right now (positive cash flow) at the cost of limiting long-term profitability.
The decision to prioritize immediate cash flow vs long term profits depends on the financial position and overall strategy of the company.
The following from Claude: """ A diff between two randomly chosen commits usually spans years of history, so it's not a few KB — the tree itself is ~1.5 GB of text, and a multi-year span rewrites a large slice of it. Call it 100–200 MB per pair on average: 8.5×10¹¹ pairs × ~2×10⁸ bytes ≈ 10²⁰ bytes, or ~150 exabytes """
Unlikely but quite interesting.
My original point is that by changing what metric and how much of the combo you look at the calculated loss can vary significantly: over an order of magnitude difference between the dog-only volume loss and the refilled drink calorie loss.
lies, damned lies, and statistics.
That's why I have percentages for both the dog alone and the combo.
The costco hotdog contains 580 calories, approximately 110 of which is the bun. If the bun shrinks by 30%, the calories of the entire dog shrinks by 5.7%%. But also the $1.50 is for a combo that includes a drink of up to 270 calories, so the caloric content of the combo shrinks by 3.9% -- substantially less than the 37% claimed volume loss (which does not account for the drink).
EDIT - The drink also allows for a refill which would add another 270 calories bringing the total to 1120, and the caloric shrinkage down to 2.9%
When looking at protein the meal loses 1 gram out of 23, or 4% loss, which is less than inflation. Good for all the gym rats.
Edit 2 - fixed percentage calculations
Gotta love single points of failure...
For most problems solvable in code, there are literally infinite valid code representations of possible solutions (at least in higher-abstraction languages, the space is much more bounded at the assembly level). If you consider the subset of those representations that result in sufficiently performant execution, does it really matter which one is running? This is where most programmers will start mentioning clean abstractions and readability and maintainability, but those are concerns of making the code understandable for people reading and modifying that code. The people are still the driving force.
I'm not arguing for slop here - I definitely care about code quality - but it's important to stay grounded in the reason that it matters: people.
Anthropic trains models on AWS's (and GCPs, and Microslop's) infrastructure, then skims margin off of selling inference also on the infrastructure owned by the other companies.
These are extremely different businesses.
If the AI is communicating to me and can't select the appropriate jargon level, it's failing at communicating effectively.