HNHacker News
TopNewBestAskShowJobs

WhitneyLand

6,579 karma · joined March 13, 2013

I have steadily endeavored to keep my mind free so as to give up any hypothesis, however much beloved, as soon as the facts are shown to be opposed to it - C. Darwin
submissionscomments
WhitneyLand··on Figma restricts MCP access to whitelisted clients, excluding Pi
Even if you’re whitelisted you get only 6 accesses a day on a standard account, have to pay for a dev account to get 200/day which still isn’t great.

For my Figma needs, having Codex do computer use seems just as good as their mcp. I can tell it, “go download the assets for what I need and take a few screenshots for reference”.

WhitneyLand··on StreetComplete on iOS is now in public beta
Yes, to be more specific, every tool like this has benefits offered and then taxes to be considered. For example, what's the tax for:

- Maintaining a significant new moving part in the dev process. KMP updates, management, integration, futures awareness, this all takes non-zero time.

- If an iOS dev wants to work on a KMP project what's the ramp up time? How much less efficiently can they contribute compared to a pure native app?

- How much time is spent on problems, bugs, issues, related to KMP. This is guaranteed non-zero for any tool.

- Is there less benefit to KMP with coding agents? Previously with two native apps, a bug fix had to be done twice. Currently, say a fix in done on Android, it's quite easy to say "propose an equivalent fix for the iOS repo". I do not claim this removes all benefits of a single code base, just that some things are not as bad now.

WhitneyLand··on Meta takes down a critical video about meta AI Glasses after filming at Meta
I don’t know the legal of boundary harassment but he’s clearly trolling them and making people uncomfortable.
WhitneyLand··on Claude Opus 5.5 Intelligence, Performance and Price Analysis (Max)
China who?
WhitneyLand··on I asked Meta’s Muse for its filesystem and it sent me 6.8GB
”we've determined that the reported issue does not qualify as a valid vulnerability…because the behavior described is working as intended”

So I’m sure they won’t be fixing it then.

WhitneyLand··on AI Has No Wisdom and Neither Will You
One thing I think it’s still easy for a human to beat AI at is a good PR summary.

They are often too verbose, or miss capturing an important concept or purpose, or add bullets for parts of the changes that no one cares about.

WhitneyLand··on I built non-autoregressive decision models with RL a year ago
How many people actually read the full post? It builds up this amazing underdog story where all the benchmarks are taken as victories, and then only late in the post and section 6 is it finally revealed that the only way they won was to fine tune directly on the benchmark.

This comparison doesn’t even make sense.

WhitneyLand··on Why I'm still bearish on LLMs after Navier-Stokes
Not sure how that vague truism applies to this paper.

Lots of papers have great results that don’t depend on the latest models.

However in this case it’s problematic:

- They specifically make claims about the state of “current LLMs”. o3 is not representative of this.

- They ask are LLMs capable of X and arrive at a negative result.

If their claim was LLM’s can write coherent sentences, and their conclusion was positive, then there would be no issue using old models because the end result would be factual.

However, when you have a negative result that makes a claim about the current state of all LLMs and the ones you were using are not current, by definition it draws the whole conclusion into question.

WhitneyLand··on Why I'm still bearish on LLMs after Navier-Stokes
1. It’s hard to trust a 2026 paper that’s showing results for such old models.

2. Chess seems to be a poor benchmark for generalized strategic reasoning. People who are good at it rely more on experience and deep domain expertise than on skills that generalize to make them experts at unrelated tasks.

3. The study sounds like proving humans will never fly because they don’t have wings. In reality, humans do fly, and Claude Fable would destroy any human at chess by coding a strong enough engine on the fly.

WhitneyLand··on Introducing System One Models and Jev
Let's say classifiers don't hallucinate. To make a fair comparison we should constrain LLMs to the same classification task. In that case, no, LLMs also don't hallucinate.

- Give Jev and LLM the same input

- Lock down both to approved/rejected/unknown (LLM restricts on decoding)

- Both can be wrong, but neither can hallucinate (invent an another option).

WhitneyLand··on Introducing System One Models and Jev
What was misleading was the original title:

"Jev: New frontier model 40-400x cheaper and 20-200x faster"

I'm not the gatekeeper of who gets to call themselves a frontier model, but I don't think most people would count Jev in that group. It sounds false.

If their specific claims hold up, then it would make more sense to say something like:

"Advanced the speed/cost frontier for structured decisions"

WhitneyLand··on Introducing System One Models and Jev
His claim was that the title is misleading, not sure how it's relevant to that claim that you use "string models" (full LLMs).

The original title before it changed less than an hour ago was:

"Jev: New frontier model 40-400x cheaper and 20-200x faster"

I'm going to agree that was misleading.

And on the second point:

>>Also "can't hallucinate" seems wrong? Sure, it can't emit an invalid type, but it can still emit a completely wrong valid value.

>that is likely true of all ML! perhaps we could debate semantics, but I don't think it's fair to say a random forest "hallucinates" in the way LLMs do"

Also going to disagree here, and I don't think it's semantics.

Type safety is not factual correctness.

WhitneyLand··on Swift 6.4 Released
What would you have done differently? All healthy languages need to constantly evolve, you only get to decide where. You can change syntax, add keywords, attributes, etc but it's a tradeoff.

Any attributes or keywords relating to Objective-C or interop are not everyday baggage for most people.

For SwiftUI, attributes seem like as good a choice as anything else.

As a whole concurrency is a mishmash syntactic mess, but I agree with the direction and the result, they're trying to manage the improvements they keep adding over years of investment.

WhitneyLand··on The k-server conjecture is true
To answer that lets get more specific:

To buy presents for a family Christmas list Mom drives to Store A and Dad drives to Store B.

As more items get added to the list, they must decide who should drive to a new store location to buy the present. How can they minimize total driving distance while kids are randomly adding new items to their list?

The proof above guarantees its possible to never drive more than twice the mileage you would knowing all the items in advance.

The big news is this guarantee works for any number of drivers with any arrangement of gifts.

The algorithm to do this was already known, what we’ve learned is it’s not possible to do any better.

WhitneyLand··on The k-server conjecture is true
This is an important result, sometimes called the holy grail of competitive analysis.

One way to think about competitive analysis is bulk discounts. In life we’re constantly having to choose between quantity and discount. We could buy 1 item for a higher price, or say quantity 5 or 10 to get better discounts. The problem comes when we don’t know in advance exactly how many we’re going to need.

What should be our strategy for choosing how many to buy, and whatever the strategy is how well does it compare with having perfect knowledge upfront?

WhitneyLand··on Sean Carroll explains the biggest ideas in the universe – Full Interview [video] (2025)
How is he very underrated?

I mean, I otherwise agree with the sentiment of your post but I feel like most people are rating him pretty well.

WhitneyLand··on GPT-6 Astra, looped transformers, and hidden reasoning
By that logic we should also consider the case of cutting the number of layers in half because that would also reduce hidden state between token generation.

In your generalized example I think the concern is when the additional evaluation effectively becomes a replacement for CoT, where something like the coconut research could replace it completely.

However, I don’t think we’re anywhere close to that with Astra.

WhitneyLand··on GPT-6 Astra, looped transformers, and hidden reasoning
No. It’s not at all by definition hidden reasoning.

Looping transformers uses additional calculations (repeating layers) to generate a token.

Reasoning (in this context) is test time generation of multiple tokens that allow a model to have a scratch pad to refine its thoughts, chain of thought reasoning in other words.

Doing the former in no way means that you have to hide the latter.

Raschka is right in this post, The Information article was wrong. The Astra system card does concede reasoning traces are sometimes smaller, but this could be for a lot of reasons, including simple efficiency. And it absolutely doesn’t mean they are going away or completely obscured.

The Last Week in AI podcast from Sept 8 seems to have gotten this wrong as well. Jeremie Harris rages that OpenAI implemented latent reasoning, ala the coconut paper, which could potentially actually obscure reasoning traces. But for the life of me, I do not know how he arrived at this conclusion and see no evidence that this has happened in Astra.

WhitneyLand··on Go grandmaster Shin defeats AI KataGo with a two-stone handicap
I don’t know that it’s that shocking, remember Go it’s not solved game, so the the limits of what’s really possible is not known in all cases.

For example, we don’t even know whether perfect White play can possibly overcome two correctly placed Black stones against perfect play.

WhitneyLand··on Discovery of a new OpenAI agent message board
If you’re wondering how they wrote to the wiki having only GET ability…

Basically it was a bug in the wiki code. They transferred the POST form parameters to GET URL parameters, and wiki internally doesn’t distinguish between the two.

WhitneyLand··on Gemini 3.8 Flash and 3.8 Flash Cyber
There are important gaps in that hot take.

For example, it's not even close to Opus 5 on Terminal-bench 4.0, 19.1% vs. 51.8%.

WhitneyLand··on SQLite as a Document Database (2020)
Why do people say document database when they really just mean json database?
WhitneyLand··on Doctors are finally learning to manage antidepressant withdrawal
“It's no where near those numbers”

“get inflated in anecdotes”

Just because one study has the number 34% in it, does not prove the numbers I gave were wrong and does not mean they are anecdotal.

This is something that’s been studied from many angles and there’s more than one number that’s defensible based on the context and dataset.

For example, even the study you cited says: “In a study of 344 patients by Montejo-Gonzalez et al, 58% of patients reported sexual dysfunction when physicians directly inquired”, and “14% with spontaneous reporting”.

In another example, a 1,000 patient study found 57.7–72.7% for common SSRIs. https://pubmed.ncbi.nlm.nih.gov/11229449/

The “scientific numbers” you allude to are not one value. It’s a body of research not one definitive number.

Taken a whole I don’t think anyone is being that mislead when it’s suggested more than half of people may experience these side effects.

WhitneyLand··on Doctors are finally learning to manage antidepressant withdrawal
So, this is not supported by the data when people are asked.

Sexual effects alone hit about 50–70% people. Then you could face nausea, insomnia, profuse sweating when you’re still, emotional blunting, and weight gain.

I don’t want to discourage anyone, they can save lives. It is a net benefit for some people, but even then most people will experience side effects and wish they didn’t have to take them.

WhitneyLand··on Black hole singularity is a surface not a point
That’s only true regarding the one sentence about the singularity not being a point.

People like to reduce papers to a simple hot take, but the paper is more than that, offers viewpoints that are non-standard and speculation about new possibilities.

WhitneyLand··on NIH is ending a key grant for budding clinical researchers
Why would you bring up fraud in the midst of science research in the US being burned to the ground?

Fraud is not the reason it is happening. Even the people who are making the cuts in this case have clarified the purported reason and it has nothing to do with fraud.

WhitneyLand··on DeepSeek V4 Flash 0731
The DeepSeek team is so strong, very impressive.

Imagine if they had GPU resources of western labs.

WhitneyLand··on Muse Code and Muse Spark 1.2
They chose to compare against Open AI’s mid tier model Terra instead of Sol and still lost some benchmark against it.

They left Opus in and got beat in all but one benchmark.

Nothing wrong with trying to improve, but why the marketing games?

Instead of trying to say in the post you’re “closer” to frontier, first set a clear goal to beat the Chinese labs on price or performance and demonstrate it convincingly.

Then when your ready, come back and talk frontier without playing hide the model.

WhitneyLand··on Why Large Language Models Fail at Tabular Prediction
Nowhere in the paper do they mention the reasoning level or budget used for the experiments?

You’ve got to be kidding me. That one variable could make a huge difference in the results. I can’t understand why they would leave that out.

WhitneyLand··on DeepSeek V4 Flash on a Single AMD MI300X
Another headline of “model runs on x”, which usually means “let’s list how much you give up to run on x”.

Dumbed down quantization?

No. Full intended inference weights preserved, so far so good.

Slow performance?

No again. Looks like you could get over 150 tokens/second.

Give up context window size?

Yes. Original model is trained for and served at 1M, this is 256k. A very practical tradeoff though. Codex is in this range, and quality does start to drop off toward the full size.

Page 1 of 34Next →