8,814 karma · joined May 4, 2016
They think that sampling is an inherent part of Transformers.
Even on this site, it is regurgitated with confidence.
It is not good for them in the sense that they will short circuit the RL path that actually improves general capability.
If someone wanted to use an 8x B300 as their daily driver...go ahead.
It would still be the best way to serve a given model.
What? LLMs are best served from a massive PD disaggregated cluster of B300s connected via NVLink.
If you're running LLMs on a Mac Mini, it's because you want to run local, not because it's the best setup.
What planet do you live on?
It's also very not true that farmers are poor in the US (median household income about 30% greater than general population). This mindset is like a century old.
It's just a low pass filter for the job market. Students aren't actually interested in learning anything, they have been condition for a long time to jump through the hoops.
It would be much better if we actually returned to exam and performance based evaluation of students, and stopped giving everyone As just for trying hard. But parents look at college fees as buying their kid a job opportunity, and so grade inflation has basically ruined the SNR of academic performance for all but the lowest performers.
Convictions at trial are much lower.
There are ~4x as many cases dismissed by judges before getting to trial. And the vast majority (90%) of defendants enter into plea bargains.
Prosecutors only bring charges when they feel they have a strong case. There are many many cases which are never pursued because of this, and people also get upset about that.
Many countries do not have a plea bargain system the way the US does. And if you look at conviction rates at trial, they are smack in line with much of e.g. Western Europe.
When the M1 was released, I couldn't believe how fast it was.
You may have better and more reproducible results using browser controls that aren't image based, or writing tools that completely sidestep browser use.
The future is not a single chat bot session of bs=1. The future is many agents performing many tasks in parallel for a single user. Large GPU clusters will always have the edge in efficiency.
Frontier labs have already been doing this for a while, verified in smuggled traces from OAT/Ant.
Simply a way to reduce tokens.
27 dense is far more capable than 35A3.
Less than half its price.
More than 50% discount.
Greedy decoding a single batch in most libraries will give you mostly deterministic outputs. Higher batch sizes can increase variance.
But all of this is down to CUDA and/or kernel implementation issues.
When OAI released gpt-oss it was released as an mxfp4 checkpoint.
OAI, Ant, et al are also obviously employing QAT.
It's almost certainly full parameter post training of the original model weights.