There are plenty of other ways to do zero shot classification that would result in more "token usage" (really just having to reprocess everything for each class), but the pricing and the way they describe it narrows it down somewhat.
1,844 karma · joined January 12, 2022
There are plenty of other ways to do zero shot classification that would result in more "token usage" (really just having to reprocess everything for each class), but the pricing and the way they describe it narrows it down somewhat.
But if they had to show how well their product worked they might give away the whole game... because they'd have to compare their "noul" class against an NLI benchmark for instance, and possibly show they're losing to cross encoders and give away the fact that they are just rebranding NLI. Or rerankers (choice) or zero-shot classifiers.
Likely they have some encoder (eg ModernBERT) trained to do late interaction or latent states along the lines of ColBERT, Perceiver IO or poly-encoders.
On the other side, the current administration also has powerful backers whose businesses are complemented by AI and would benefit from it becoming a commodity. Think about who Vance is speaking on behalf of.
So the AI companies need to use fear to drive public demand for their regulation efforts. IMO there's little to no honest discourse and a lot of attempts to manipulate public opinion for objectives that we have little stake in. I personally don't trust any of the involved parties to make decisions for the public good.
Anthropic needs regulation in order to prevent AI from becoming a commodity. This of course does not benefit all tech businesses equally, especially those that are not currently at the AI frontier. So when JD Vance talks about AI, he talks using the mouth of Peter Thiel who may not see benefit from the same policy as Altman or Amodei.
The rest is just public support posturing and most of that is bullshit meant to distract from the high rollers game of winners and losers. The philosophy is money and power, who gets it and who doesn't. Us normies aren't really participants in the game, except where we are being manipulated into cheering for one side or another, and with little stake in the outcomes (although selfishly, I'd be pissed if I didn't have open models to tinker with).
Open source is fundamentally a vehicle for commoditization. This is great if your business is not AI and your business is instead something like GPU hardware or some product that uses AI. But it means that eventually, selling AI is not going to be the money maker.
OSI proliferated open source on a business strategy called "commoditizing your complements". These big companies don't do it out of benevolence. It was pitched to them in a way that FSF did not (which was more about morals and ethics, something business care little about), and it caught on. And the software business became about ads, consulting and cloud services instead.
Imagine the assurances that are made to convince donors it's worth it. Who exactly bears the cost of unpopularity?...
I think about that a lot now, and how my religious upbringing was bombarded with this messaging (in a good way). And how much of the same world that previously claimed espouse it has complete turned their backs on it in such a short time.
Is it dead or did it ever exist at all? At least we don't seem to pretend to have higher values anymore. If we even have leaders at all, they are not measured by the yardsticks we used to make.
Generally though, many of these fairness complaints do go away if there is "3rd party testing". Right now, companies reporting their own benchmarks has all the problems that 3rd party testing resolves in many other industries.
Cherry picking the benchmarks you present is where the falsehoods lie.
So generous.
If you bring something new to the table, then in my experience, AIs are really good at helping you ground it old ideas. If you want to set it and forget it, then you will get the mean. If you want to do something new, in my experience, they are enablers and not blockers.
Skeptical me seriously doubts this is an effective solution for crime. But maybe that's because this country has a history of being willing to do a million expensive and privacy violating things, and only if it's a punitive measure.
90-95% is a very good estimate! That's about what we measure on our test set. I have good news for you, and we will have a blog post about it soon. Because of how our models are built, we are able to optimize for detection accuracy directly by constructing synthetic swipes on each layout for ~50k words, and then testing them through the model. We tested around 800,000 layouts this way.
The biggest issue with QWERTY is that there are far too many words that swipe colinear or obtuse angle letter trigrams. These are both hard to detect and frustrating for swipe users, because you can't clearly indicate the letters you're gesturing. Neural swipe models (at least ours) look for indicators in the gesture pattern that suggests a user was targeting a specific letter, rather than trying to match a gesture shape like algorithmic detection does.
The shape of the keyboard can significantly improve the way the gestures are formed so that there is better indication of letters. The model can still respond to dwell times because unlike shape matching it uses the temporal information. But dwell interrupts flow, and in my opinion should be minimized in swipe layouts.
There's even more options still, especially if you go further back toward more traditional methods. Static word vectors like GloVe or fasttext (optionally more modern equivalents like WordLlama or Model2Vec). Then there's sklearn-style stuff too. Those can be really small/fast but have more accuracy-level tradeoffs.
- Zero-shot encoders like tasksource or GliNER
- Natural language inference: https://huggingface.co/blog/dleemiller/nli-xenc-ways-to-use
- GRPO training
- GEPA prompt tuning Qwen 0.6B (or GEPA, then GRPO)
- Use an embedding model and train a classifier (MLP, logistic, svm)
- Use a larger LLM to generate a synthetic dataset (beware of lack of diversity, mine "seed text" from real sources first)
- Synthetically generate "hard examples" where more than one category may be valid and DPO tune your preferred responses
I think people tend to fixate on the worker-to-worker differences inside of unions. Yes, that is the most visible part of a union when in place, and at least in the US has valid arguments about meritocracy.
What is missed when limiting the scope to just that is the population-level abuses of workers that no amount of meritocracy will fix. When corporations engage in collusion against workers (now common and nearly unpunished in the US) the top-level wages are suppressed industry wide.
The whole pay band alignment that comes out of that undermines the meritocracy argument, and doesn't even begin to address the wage-fixing that has gone almost unchecked in tech for decades[1,2]. As a merited employee, you might have more options to where you can go, but it won't protect you from predatory hiring/layoff cycles and it certainly won't guarantee that you'll receive a truly competitive wage.
On paper, meritocracy sounds great. I have worked many places in tech and never once observed it, personally. Best case, if you have warmed a seat for enough years, then you advance that way. Worst, your employer knows they can just take advantage of you because you're willing to work without a dangling carrot.
As before, either the government frees itself from corruption and enacts justice or unions will come back. That is point we are at.
[1] https://www.npr.org/sections/alltechconsidered/2015/01/16/37...
[2] https://conversableeconomist.com/2025/10/31/the-silicon-vall...
They also seem to have adopted a no-remote hire policy and are in an extreme high CoL location. It’s a truly awful mix for trying to attract outside talent. I don’t know why they even bother.