Smaller, faster, safer: running Kimi and GLM at scale
blog.cloudflare.com
blog.cloudflare.com
However, I wish their testing were more detailed. Firstly, some model families are more sensitive to KV quantisation than others (only Kimi K2.6 was tested). Secondly, the evaluation suite they use to claim that FP8 KV quantisation is indistinguishable is noticeably lacking coding benchmarks; in long-running tasks, minor tool call errors compound over time.
I'm wondering if it needs to be tested with every other model or not.
> None of this would matter if it changed the model's answers
If they want to assert that the answers don’t change, then perhaps they should calculate the statistical distance between the token probability outputs or something to that effect. I doubt the results would indicate that the answers don’t change by any reasonable interpretation.
Maybe the results are still good enough.
We let all traffic get MITM'd, now we're letting our AI conversation get tracked. Cloudflare reeks like a US Honeypot.
Cloudflare’s inference absolutely does support ZDR, as long as you use unified billing (ie not using BYOK).
"Your inputs (e.g., text prompts, image submissions, audio files, etc.), outputs (e.g., generated text/images, translations, etc.), embeddings, and training data constitute Customer Content.
For Workers AI:
* You own, and are responsible for, all of your Customer Content.
* Cloudflare does not make your Customer Content available to any other Cloudflare customer.
* Cloudflare does not use your Customer Content to (1) train any AI models made available on Workers AI or (2) improve any Cloudflare or third-party services, and would not do so unless we received your explicit consent.
* Your Customer Content for Workers AI may be stored by Cloudflare if you specifically use a storage service (e.g., R2, KV, DO, Vectorize, etc.) in conjunction with Workers AI."
OpenRouter has a page of different providers and what OpenRouter understands their position on this to be: https://openrouter.ai/docs/guides/privacy/provider-logging#d...For CF they write: "Cloudflare • Prompts are retained for unknown period • Does not train"
So in the sense of training models on your company's codebase or maybe running analytics on prompt content for the purposes of improving their AI product suite they won't use your data. Though presumably for purposes of security/abuse etc. there will be some level of retention as per their privacy policy.
I guess it comes down to how much you trust providers on openrouter who claim to offer absolute ZDR versus Cloudflare and how they would use retained data. Obviously if someone thinks CF is a honeypot designed to sidestep the rise of LetsEncrypt/widespread HTTPS then they wouldn't trust a ZDR claim by them in any case. Would you then trust some of these frontier labs respective claims of ZDR when they are pushing unbelievably hard to win? I don't have strong opinions for this - I have a pretty conservative approach by default and exclusively use local inference on in-office hardware for anything close to or related to customer data. I do some coding on 3rd party services.
I imagine if Mullvad offered a ZDR set of open model endpoints with similar efforts at building trust like their VPN it might be popular.
Turnstile everywhere + device attestation required is the direction this all seems to be headed. Because of all the AI bots of course (it is an excellent scapegoat).
Why… I wanted to see if it’s worth it to use cloudflare’s endpoint but I can’t even see the pricing
Input tokens (per 1M)$3.00 Cached input tokens (per 1M)$0.30 Output tokens (per 1M)$15.00
$3/$0.30/$15 ( input / cached input / output, all per 1M )
I'd say serving quantized models without saying so on the "store" page is fraud.
What is the typical job title and/or skillset for this?
Applied ML Research Engineer or something maybe, not that I have ever seen that title. Maybe just catchall ML Engineer...
It can be any one of many jobs depending on how high close to the metal one's focus is, but the highest headcount role is usually SRE/Infra/Ops with GPU knowledge sprinkled on top. That is to say Linux sysadmin, networking, fleet management, scaling, incident troubleshooting, etc.
I love AI, but I really hate reading it.
How well it would work on this site, I'm not sure.
then we're all screwed I guess. lol
I can only read so many (either pro OR against) ".. IT'S AI! ! !" comments before skipping the thread. I can't be the only one.
HN is for conversation between humans[1] (about AI generated blogspam, apparently)
I've been asking for that for some time.
* they use quantized models
* they quantize KV cache
* they have a cache tagging mechanism to prevent cache misuse (neat)
The agent can extract numbers without filler prose as well.
I'm not thrilled with it, but he is obviously using it to improve his writing overall- to communicate some great ideas that are personal and germane. I've decided that being too inflexible serves no one. If it is true slop, I'll not revisit the writer in the future- if they are using AI to polish writing that at its core is a unique voice, I'll accept it and learn to live with it...
What's funny it's that is as if AI "learned" to speak english but not really. People simply don't speak using those strange constructs: those sentences sound a bit like if a "Karen" was trying to make a point.
What's scary, to me, as a dev using AI, is that those LLMs do the same thing with code: it looks like proper code, but it really ain't so once you dig a bit.
It's verbose and doesn't add anything: it's just infinite verbiage / sloppy-pasta.
Crazy thing though it's that it's 2026 and apparently devs can't be bothered to copy/paste their sloppy-pasta LLMish into a de-sloppifier before publishing blog posts.
Of course it always looks better, than it is, but it still works for me as in the code does what I asked the agent to.
After all, why would advertisers want to advertise to bots? This whole sentiment of "You'all need to allow my bot postings in your group" is getting tiresome. There's plenty of places where your bot can talk to other bots, publish for other bots, etc.
One vCPU means nothing, which chip? which SIMD? what RAM speed? what storage? what latency?
Quantization is the next layer of lies
Society is moving towards offloading intelligence to clouid overlords, since nothing runs on your computer anymore, you can't trust anything
And since datacenters occupy a physical space, they are building a monopoly defacto
Unless transparency becomes mandatory, we are headed towards the biggest self sabotage mankind has ever witnessed
Don't forget: When will all of this change?
You like sounding dramatic? There is a RAM explosion and big models already work on laptops. Some people will offload everything, some stay in control.