HNHacker News
TopNewBestAskShowJobs

fastball

12,756 karma · joined September 1, 2012

I used to argue on the internet too much.

Engineering at Known (https://known.com)

Co-founder and tinkerer at Supernotes (https://supernotes.app)

submissionscomments
fastball··on How Delhi cut electricity loss from 50 to 5 percent
Small tangential note, but I think you might mean "not for the faint of heart".
fastball··on Claude Opus 5.5
It hasn't just fixed it, it has introduced a new paradigm in anti-obscurity.
fastball··on Amazon blocks Meta’s new Muse AI agent from shopping on amazon.com
Amazon is not usually very good anymore, quality has suffered massively.
fastball··on I built non-autoregressive decision models with RL a year ago
Hallucinate tools that don't exist.
fastball··on I built non-autoregressive decision models with RL a year ago
> a bit cheaper

Gemini 2.5 Flash Lite is $500/Gt, Jev is $42/Gt. AKA an order of magnitude cheaper.

> BERT with more data

It is specifically not just that, in the same way that models which have been chat/task-optimized via RLHF (which made these models much more useful for a huge variety of tasks) are not just "the base transformer model with more data".

fastball··on I built non-autoregressive decision models with RL a year ago
I don't think that is an entirely fair comparison. They are comparing Jev to the way people are currently using generative LLMs for things like classifying/tool calling/any kind of structured output.

For example, if you feed in some context to Jev and Claude Haiku and say "make the appropriate tool call based on this context", Claude (or any other frontier LLM) will hallucinate tool calls some percentage of the time. Jev will not. While yes, the "will not" is constrained by Jev's (lack of) capabilities in some sense, this is actually a very real need for a wide variety of use-cases people are currently using off-the-shelf LLMs for at the moment.

Probably the better example is the whole probability thing, where even if you use something like constrained decoding to ensure an LLM only outputs a certain schema, and therefore can't hallucinate a class, if you ask for probabilities, the probabilities output by the model are just hallucinations. Jev meanwhile is outputting calibrated probabilities for different choices based on the actual landscape.

fastball··on Hacking OpenAI
Bounties really ain't what they used to be.
fastball··on TSMC revealing details about next gen A14 node
On mac this is just Opt + A
fastball··on More questions about whether researchers can trust OpenAI with unpublished math
If you have a business account the terms say they will not train on your data, so that seems like the easiest route to avoid such questions for researchers.
fastball··on Neki – Sharded Postgres
Yeah this meta commentary on HN is very confusing for me. The structure of the article is literally

- short paragraph saying "we launched neki"

- short paragraph saying why neki was needed

- paragraphs about what neki is (with a header)

That seems like a very reasonable structure for a blog post like this one. Not being able to get to the third paragraph of a blog post seems like a "you" problem, not a blog problem.

fastball··on Neki – Sharded Postgres
You can also find that in first section of the linked post (after the intro)...
fastball··on A/I shuts down
I appreciate your wanting to introduce nuance into the discussion, but nuance isn't really needed here. Two commenters replied to GP saying "Pro-hamas activists don't exist", which is untrue. That was my claim, and that was basically all GP said as well (with the additional assertion that such people are not voting for Trump). Why they are supporting Hamas is beside the point.

This conversation is basically the "it's not happening" > "it might be happening" > "it's happening and it's good actually" meme, though admittedly not by one person.

fastball··on A/I shuts down
Very clearly untrue. Just search for "people waving hamas paraglider flags" on Google and you will find plenty of sources[1], like this.

Basically every movement/organization on earth has activists who support their activities, why would Hamas be any different? There are people everywhere who don't mind violence/killing and there are people everywhere who are anti-semitic. You think the intersection of those two cohorts is empty?

[1] https://www.bbc.com/news/uk-england-london-68275185

fastball··on A/I shuts down
That's exactly the point though: if you are choosing to platform some and not others, and some (many? most?) of the groups you are choosing to platform are violent groups, you very quickly run out of plausible deniability.
fastball··on A/I shuts down
> A/I’s cadre of radical hackers and tech developers provide a full spectrum of services

Pun intended?

fastball··on GPT-6 Astra
Not really? Nobody is talking about source-available code with a restrictive license. People are talking about having LLMs write up protocols and libraries and firmwares from scratch and making those freely available, as in open source.
fastball··on GPT-6 Astra
How is that relevant?
fastball··on GPT-6 Astra
AFAIK it's just yolo mode, which doesn't actually do any checks like Claude Code's auto-mode (which has a model checking all commands for safety).
fastball··on OpenAI's GPT-6 Astra on ARC-AGI-3
Are we sure an Astra hacker swarm didn't compromise arcprize.org's servers and exfiltrate the private eval set in order to achieve that 99%?
fastball··on GPT-6 Astra
Codex's lack of auto-mode is what prevents me from using it for serious work compared to Claude Code.
fastball··on Claude Fable 5.1 and Claude Mythos 5.1
Output tokens are 5x more expensive than input tokens, so I'm not sure "dominate" is entirely correct.

A conversation with 20 turns, 50k tok growth per turn, 1m tok context at end would price out like this:

Fable 5 ($1/M cache reads) ; cache reads 9.5M tok × $1.00 = $9.50 ; cache writes 1M tok × $12.50 = $12.50 ; output 1M tok × $50 = $50.00 ; total = $72.00

Fable 5.1 ($0.25/M cache reads) ; cache reads 9.5M tok × $0.25 = $2.38 ; cache writes 1M tok × $12.50 = $12.50 ; output 1M tok × $50 = $50.00 ; total = $64.88

So yes, cheaper, but not massively.

fastball··on Should We Have Kept the American Empire?
Say more?
fastball··on Hy4 preview
I wish model providers would stop committing chart crimes in their releases.

- if you're gonna order the rest of the bar chart by rank, order your model accordingly.

- if you're gonna highlight a winner in a table of benchmarks, don't highlight your entire model row in the table.

Etc etc

fastball··on The Harness Is the Thing
Do large development teams care about Firefox?
fastball··on Starbase, LA
The problem is that neither you nor I have evidence that says they should be paying, either.
fastball··on Starbase, LA
Nth order effects are powerful things. A single person could change the trajectory of Louisiana in a meaningful way.
fastball··on Starbase, LA
As much incentive as the oil & gas industry receives?
fastball··on Starbase, LA
hundreds of millions unpaid and zero investigative journalism to try to get at the truth, besides "liens filed therefore Elon Musk is a cheapskate".
fastball··on Building an (almost) fully self-hosted, sandboxed, agentic software factory
You can test for test vacuity with mutation testing, using tools like Stryker (for JS). In my experience this significantly helps models to write tests that actually test for correctness rather than aligned assumptions.
fastball··on AI companies destroy physical books – let's scan rare books before it's too late
So we don't want companies to buy books and scan them and do whatever they want afterwards, but we also don't want to allow piracy of digital copies (the $1.5B Anthropic settlement)? Bit of a rock and a hard place for them.
Page 1 of 34Next →