HNHacker News
TopNewBestAskShowJobs

sroussey

5,695 karma · joined December 20, 2009

workglow.dev

Previously: Embarc, Privicy, Wag, Weebly, Firebug, Network54, Kantara

submissionscomments
sroussey··on Gemini 4 Argon
No compelling results because summarization is really hard.
sroussey··on A Privacy Analysis of Web and Mobile Conversational AI Agents [pdf]
It is also super common on e-commerce websites that will then send you an email about a cart you started to the point of adding an email but then abandoned.
sroussey··on Gemini 4 Argon (High): Intelligence, Performance and Price Analysis
Same price as GPT-6.1-sol however.
sroussey··on Gemini 4 Argon (High): Intelligence, Performance and Price Analysis
GPT-6.1-sol costs less than half of Gemini and way less than Anthropic on the cost per task chart of the listed parent page.
sroussey··on A Privacy Analysis of Web and Mobile Conversational AI Agents [pdf]
You open the same convo on another device and the partially written text is there to continue. Does it not do that for you?
sroussey··on Jeff – Jev-compatible 0.8B decision models, trained at home, ~30 ms
Classifiers (instead of LLMs) return results with confidence scores.

Name Entity Recognition (NER) is one example.

So many of them... https://huggingface.co/models?language=ner&sort=trending

Also used to block SSN and CC #s from logs, etc... as small and fast enough to do it. You don't want to call OpenAI GPT-6 and ask it to return your text with the SSN blanked out. I am sure people do though... (SSN is a bit simple, but all kinds of PPI in one model is more likely).

The nice thing about Jev is that people started taking about models that are not LLM text streams again.

sroussey··on Meta's Muse appears to use an OpenAI model labeled muse-special
I’ve had Claude opus do that too.
sroussey··on Allow Carriers on Planes
Um, maybe don't allow babies on planes.

That would make my flying a bit nicer.

sroussey··on California is chasing wealth that has feet
This is a prime example of short term thinking, often espoused by politicians.
sroussey··on AX – Google’s Open Agentic Orchestrator
Why would you need multiple agents and not one agent with multiple repos?
sroussey··on Samsung is expected to more than double output of its HBM4 and HBM4E DRAM
There are AMD and Intel devices on similar process (not talking about A20Pro or M6 which are set to ship later this week), and they do not get the same gains.

And honestly, they have historically had different markets.

When the design is for only one customer, you don't need to generalize things, and those things you generalize to give different customers different options has costs.

AMD will soon be a larger customer for TSMC than Apple (NVIDIA is already there) so Apple's pre-booking new processes is likely to be gone in the near future.

sroussey··on Samsung is expected to more than double output of its HBM4 and HBM4E DRAM
HBM also trades bandwidth for latency, and your regular computing is much more sensitive to latency than bandwidth.
sroussey··on Samsung is expected to more than double output of its HBM4 and HBM4E DRAM
That would be like skipping land line phones for mobile...
sroussey··on Tin: full-text search for Postgres
what do you call bare metal in my office?
sroussey··on Typesafe-computer-use drives a Mac toward a goal for 1/50th of a cent per step
Curious about people’s experience here. I am working on a small model, verify by jev, and escalate to big model. Some cases, the small model is not a model but some regex.

cheap-confirm-escalate

Using jev as the confirm step.

sroussey··on Show HN: Cactus Needle 3: 8-29MB automation models can match DeepSeek V4 Flash
For reference: https://huggingface.co/blog/webgpu-kernels

I think it will be the basis for a rewrite of transformers.js v5, but no need for you to wait as you would likely want direct access. It is also way better than loading WASM, and faster to boot!

sroussey··on Show HN: Cactus Needle 3: 8-29MB automation models can match DeepSeek V4 Flash
Would love to see this implemented with @huggingface/kernels for shader compilation for Webgpu.
sroussey··on Saving another 100TB of RAM
Those machines with GPUs still need RAM of their own, and they generally want large caches to avoid SSD penalties. You even see this spill out in the form of costs for KV cache in <1min, 5m, 1hr rates etc.
sroussey··on Saving another 100TB of RAM
Someone really needed a few hundred TB to waste on inference and went looking under the rugs…
sroussey··on Saving another 100TB of RAM
Having worked in hardware for a moment, everything we do in software is like this. Even C.
sroussey··on A heap overflow and SSO misconfiguration to compromise OpenAI internal repos
Maybe these big ai labs will uses their own devices to find and fix bugs up and down their stack and contribute that back.
sroussey··on A heap overflow and SSO misconfiguration to compromise OpenAI internal repos
Yeah, isn’t that Claude Codes sandbox? That drops and every npm install it taking over the world, lol.
sroussey··on TSMC revealing details about next gen A14 node
Or AMD’s putting sram on a separate chip. I’m surprised they didn’t go there yet except for extra L3 instead of all of it. At some point the costs will shift the decisions.
sroussey··on How Uber Protects Against Retry Storms
Yes, exponential back off and jitter are the first things to work on, and good if you don’t have a better signal (like loss of network).

Also, a simple signal status server or queue system helps to keep global state such that everyone doesn’t retry all at once.

If you have a central error rate server you can skip your retry based on the error rate (100% error rate, don’t retry, etc).

sroussey··on How Uber Protects Against Retry Storms
So many variables, but the simple thing is to set things up like normal rate limiting (which you would want to do anyways). The one generating the errors passes back a retry time. You can add jitter here, tell low priority requests to wait longer, etc.

BTW: do keep track of priority. It’s like having a database that gets flooded with connections and won’t allow new ones in—but will for admin users (btw, it did not used to be that way in the early days of MySQL).

sroussey··on Bonsai 2 27B: Near-Lossless Compression in a 9x Smaller Footprint
Yes! And maybe get a hf fused webgpu runner for that model so it’s fast!
sroussey··on Introducing System One Models and Jev
The hybrid nature of this thing is not a detriment, nor does it make it an LLM.
sroussey··on Introducing System One Models and Jev
I dunno, I would consider Waymo and Tesla to have frontier models.

I think AlphaFold and related are also frontier models.

Being an LLM does not seem like the qualifier for frontier.

sroussey··on iOS 27, iPadOS 27, and macOS 27
I have the same problem when i touch a result and it changes the result line at the exact moment.
sroussey··on Dario, Please
And there was the Tesla thing CNAMEing time server pools and hiring people to pen test, which sent automated attack systems on volunteers servers. Last I heard, Tesla et al didn't even care enough to respond.
Page 1 of 34Next →