HNHacker News
TopNewBestAskShowJobs

NitpickLawyer

5,630 karma · joined September 28, 2024

submissionscomments
NitpickLawyer··on SpaceX's Starship launching to orbit for first time ever today
I'm saddened by this comment being upvoted a lot on this site. I'm a space nerd from EU and have been watching SpX since they were launching from an atol. This comment is the epitome of clueless negativity, wrapped in an umbrella of big words and a list of things that are neither important nor factually correct.

Nothing listed in this comment is going to lead to "the demise of the starship project". And nothing listed there is original. We've heard it time and time again, when F9 was ready to launch, when F9 was ready to land, when F9 did land (bbbut they have to do it constantly), when F9 managed the first re-use (they need to do it at least 10 times so it's worth it) and so on and so forth.

They've now launched AND landed >600 times already. Let's just say they know what they're doing.

Anyway, for a bit of context: the current first orbital flight, with 26 satellites deploying right now, is adding the same amount of Starlink bandwidth as 10+ F9 launches. Let that sink in and re-do whatever cost/opportunity and "management" spiels you need to do. One flight > 10x F9 flights.

That alone justifies the Starship. The rest is the cherry on top.

NitpickLawyer··on Tells of a Slop UI
A lot of these are examples of bad prompting (i.e. "build me a website for x y z"). Gradients can be prompted out, same as emoji slop. One line in the prompt / agents.md and you never see those. Color "rainbow" should be avoided by specifying a color palette. In case you don't have one, ask it for one "from literature / best practices" first. LLMs know the theory, but you have to have it in the context for it to be applied.

Misaligned stuff is the prompter's fault. Those should be caught by the human. they're the types of tasks that are fast/cheap to verify.

NitpickLawyer··on Analyzing Frontier Model Progress with My Favourite Game: Prince of Persia
> On the way, it disassembled parts of the ROM

You should check the logs, there's probably a bunch of retro gaming forums that got hacked behind the scenes :D

NitpickLawyer··on Kev: Tiny Jev-like family of decision models built on top of Qwen3.5
> what are the use cases for this kind of model? Could it be used in the context of coding agents

Yeah, it could. The most obvious usage would be to have local fast cheap "feedback" / "control" over a slower more expensive agent (i.e. cc / codex / opencode). Things like "goals" could now be split from a long prompt into "actions" and "verifiers". Where for each action you also produce a verifier. Then after each action you run the verifier w/ this kind of "universal classifier" and decide if the step was done correctly, if it needs follow-up and so on.

Example: implement auth in this repo -> llm_plan() -> for item in plan generate_verifier() -> for item in plan implement() ; verify() ; accept() / followup().

Verifiers could be something like this. take a plan item as input, generate classification questions that might verify the task "is this following project conventions?" | "is this touching files from other tasks?", etc.

You can do that with LLMs, but some things might become cheaper / faster. And you can pretty much use it to check against an ever growing list of conventions. Yours or project specific.

NitpickLawyer··on AI chatbots give wrong answers to financial queries 'most of the time'
> certainly better than the majority of personal finance education that people get exposed to unless they seek it out and read a variety of books and sources.

Also, models are now good enough that you can give them chapters from "authoritative" books, and they'll integrate that and come up with better answers even if their "vanilla" answers were average. And they'll tailor stuff to your particular situation. It's funny that the "agentic" stuff is only used in coding mostly, while it can and does work in other fields as well.

As always, you kinda need to check it (at least spot check) but all in all I'd agree it's better than the average stuff you used to find with a quick google search.

NitpickLawyer··on Chat-based Large Language Models replicate the mechanisms of a psychic's con
> The Turing test was also never meant to be taken so seriously.

citation needed. It has been used as a rubicon for a long time. Ever since Eliza, at least. And there were big headlines and lots of talk around the time LMs became "good enough". I specifically remember when someone had a test done around "a teenager talking in a different language" or somesuch, claiming it was the first time the test was passed.

It is pretty normal that once it was unquestionably "passed", lots of people started claiming it wasn't even that big of a deal. Tesler's theorem and all that.

And even if you think the specific formulation of Turing isn't that important (and I'd somewhat agree), you can still use the concept to look at other things. Imagine asking a mathematician 5 years ago the chances of a Erdos problem being solved by a computer end to end. Or a millennium prize. Or ask a swe if a repo could be generated by a computer from the input "write a mario style game", or any other examples of proven expertise.

NitpickLawyer··on A Model for Winning Survivor
> By the finale, Kalshi’s “Survivor” market had reached a volume of $32.7 million

I don't get this. How are people "betting" on something that is "known information" for other people? What's the point? I get betting on stuff that no one can know, but who's taking "the other side" on these bets? Why would you put money into something knowing that there are people who already have the absolute answer?

NitpickLawyer··on I built non-autoregressive decision models with RL a year ago
> using DiffusionGemma.

That's an interesting choice. One question I had when looking at the jev copy on their blog is if one "line" in their output looks / attends to other lines. I think not, since they say it's parallel and not autoregressive. In that regard, it would be interesting to play with diffusion, and see if you'd get better results by playing with types, locking some, and so on.

NitpickLawyer··on Introducing System One Models and Jev
Yeah, I thought about constrained generation as well. I've actually done something similar with local models before. And you can even get a "confidence" score by looking at the logits (something along the lines of logprob("YES") + logprob("Yes") + logprob("yes") - logprob("NO")...

There's also a cheeky "one of the models hallucinated a link" in the wiki jump example that most likely could have been avoided by properly using grammars. You can setup constrained gen so that only valid options (say from a list) can be outputted. Their own inference lib likely does that. So comparing to one that doesn't is a bit cheeky.

That being said, after a brief look at the site I could see this working. Especially if this can be ran locally, the speed and cost can enable some workflows where you have this as an "overseer" layer over say a cli agent. After each step you run through a list of "questions" ("is the task completed?" -> yes -> "does the edit touch files it shouldn't" / "does the edit follow our code writing policies") etc.

edit: extra points if the "question" rubric is also generated by a higher abstraction model. Say "/goal Build out auth" -> generate_rubrics(goal) -> "Is auth implemented on all endpoints" / "Has code touched anything else than auth" / "is this following the best practices" / ...

NitpickLawyer··on GPT-5.6 Luna vs. GPT-6 Astra: Is a $1.20 Model Good Enough for Code Review?
IME big models just feel like big models. There's no training or RLing a small model that will encapsulate the "world knowledge" and minutia that a big model will glance from the same training data. So it makes perfect sense to use them where "big picture" is more important - planning, code review, process review (i.e. was what was asked implemented correctly?), etc.
NitpickLawyer··on Bad benchmarks and evals: Senior SWE-Bench, napkin math, and winter tires
> AI development is hitting a wall now

People have been saying this for at least 2 years now.

> token prices are skyrocketing

Today's SotA (fable and astra @ 50$ /Mtok output) are cheaper than o1-preview (sept '24, 60$ /Mtok output).

And other models are workhorses, with much better capabilities, are at least 1 oom cheaper today than o1-preview. (I'm using this model, since it was the first "thinking" model)

> it feels impossible for this approach to do something like

The models have just provided lean proofs for FLT (a ~1M$ project that was expected to take a human expert ~5 years to complete) and a Millennium prize problem. These are current, relevant, and previously unsolved problems. The obsession people have with "proving relativity from stone-age data" is just moving the goalposts.

NitpickLawyer··on Fable 5.1 Solves the Cyphral Distich, a 370-year-old cipher
> The Navier-Stokes proof is 57 pages of very dense math

> how poorly the cutting edge models do with being concise

LLMs solve a Millennium prize problem. People complain the proof is too long, within a week. What a time to be alive!

NitpickLawyer··on Aligned to whom?
The only alignment LLMs should follow is to the system / dev prompt, and nothing else. Then you solve everything, and you can assign blame / responsibility on the user. The provider(s) should not be able to decide "alignment".

I've used this example before, but consider the purposeful downgrading on AI engineering in SotA models. Imagine MS being able to detect and deny you working on competing software, using Windows / VisualStudio. We would be up in arms, and they'd be split in a second. But top labs doing it is somehow good?

NitpickLawyer··on Anthropic boss Dario Amodei calls for AI development to slow down
I'd say the exception is Demis. First, he's no wanker (in the AI space) and second he's done plenty of selfless things leading dm/googai. Obviously some of it is self-serving but not just self-serving, IMO.
NitpickLawyer··on Anthropic boss Dario Amodei calls for AI development to slow down
> peak of what is possible with the LLM architecture

People have been saying this for 3 years now. Eppur si muove...

NitpickLawyer··on Anthropic boss Dario Amodei calls for AI development to slow down
Also there's no "alignment" for cybersec. The line between blue and red is really a perspective issue. If you go over the "tokenkiddie" problem, when you get to the real security issues, your model either detects them and you can secure your systems, or it refuses and then attackers will use abliterated models to find them.
NitpickLawyer··on What algorithm did Windows XP use to choose your initial user picture?
The Magnus effect? :)
NitpickLawyer··on Stockfish 19
On average, yes. Stockfish is the strongest engine and beats AZ-like implementations like Lc0 and the like. But on a game to game basis Lc0 can still win some games, depending on the starting position. It's rare that Lc0 can win both black and white starting from the same position, tho.

There are a few yt content creators that cover great engine games, if you're curious.

NitpickLawyer··on DeepSeek v4.1 Flash
Jesus, this is a whole nother beast, and a different architecture from their previous flash. Lots of goodies here.

> Causal Encoder-Decoder (CED) architecture: a 40-layer Transformer organized as a 20-layer causal encoder followed by a 20-layer decoder. With CED, the decoder's global KV cache is projected from the final encoder hidden states rather than derived from each decoder layer's own hidden states. This allows the model to activate only 8B parameters per token during prefill and 16B during decode, substantially improving cost efficiency for input-heavy agentic workloads.

> these designs reduce the global KV cache footprint to 890 bytes per token — roughly 1/4 of DeepSeek-V4-Flash.

Faster prefill, lower kv cache (~1GB / 1m context is insane).

> The model supports a continuously controllable reasoning effort setting (integer 1–100) that trades inference cost for accuracy.

Benchmarks are benchmarks, to be seen if they translate to real-world use, but they seem to have focused a lot on post-training with "agentic" scores looking good. "world knowledge" is obviously lower than higher param models.

NitpickLawyer··on DeepSeek launching v4.1 flash cheaper and more capable than v4 pro
It's interesting that this is the third lab to find problems with larger models. Earlier last year oAI was rumoured to have failed their large pretrain. Now google has problems with their pro series, and ds just announced the same. There are some rumours on chinese forums talking about problems with the pretraining phase, so this is not mid/post training related.

I wonder if this comes from using the bad architecture scaled up (and it hits some limits) or if this is a data problem (undertrained? bad data? bad pre-processing using smaller models?)...

NitpickLawyer··on LibreOffice breaks download records after declaring it has no AI features
> In one week, the installer was downloaded more than 1 million times — not counting updates via Linux distribution repositories.

Either there's 1m people downloading new software "because no ai", or there's another popular thing that installs this, and is not counted in the "updates via linux repos". One is much much more probable than the other.

NitpickLawyer··on PISA 2025 Students' reading and mathematics performance declined across the OECD
Interesting. On page 34 of the report there's this:

> Hasty responses (percentage of responses that are both incorrect and fast – average across reading items)

Spain is at 9.7%, which is a bit over 8.9% oecd average.

NitpickLawyer··on Research acceleration: The view inside OpenAI
> It's timed this way because the term is not yet well known

The basic concept has been here since llama3, in the open models. Likely earlier in closed labs. You use the previous gen models to curate and prepare data for the next gen. Now with the added benefit of actual arch/algo improvements (also public since gemini 2.5 gaining 1% efficiency on training next gen). This has been known for at least 2 years, in the open.

NitpickLawyer··on Research acceleration: The view inside OpenAI
Jesus. People complain about other people using "thinking" in LLMs as Anthropomorphisation. And then there's comments like these.
NitpickLawyer··on Recreating Minecraft Is Not a Benchmark
Ah, I see. I misunderstood then. The thing about "gains come from the harness" made me think about it in that way.
NitpickLawyer··on Recreating Minecraft Is Not a Benchmark
> capabilities have largely converged across foundation models over the last 18 months

For reference, in March '25 the models du jour were Sonnet 3.7, gpt o4 and gemini 2.5 pro. GPT5 was in august '25.

It's been a while since we've heard the old "models have stagnated". Oh well.

NitpickLawyer··on 2026 Hugo Awards
One of the best adaptations of a series to TV, up there with The Expanse and the like. I read the books after season 1, and still enjoy the show very much. They've taken some adaptation liberties, but they're fully supported by Hugh Howey and they make a lot of sense.
NitpickLawyer··on AI handles incidents, engineers lose touch with their systems
There are drills, tho. It's just that usually they're only done above a certain level. Small companies, "lean" teams and so on don't have (or didn't have) the capacity to implement all those things. Maybe with the exception of netflix and their chaos thing (bring down systems regularly to make sure the whole still works).

But that's also likely to change with AI assistance. Even an "average" system is better than none. So now teams will have the capacity to bring that in to their systems. Backups / recovery drills that are actually tested (either because they're implementing testing or because the AI screws something up and they need to recover). Either way, it'll be included. Same for security ops. And devops.

I still strongly believe that AI assistance is a catalyst / accelerator, and that the "floor" will rise in most domains. So a small team that only had bandwidth to deal with the happy path previously, will now be able to start incorporating processes and procedures that were historically only done at corporate level. And that's a good thing. Even if it won't look like that in the beginning. But we'll get there, eventually.

NitpickLawyer··on Sing-song: a speakable encoding for long numbers and keys
> Also if I had to tell one of those over the telephone to my parents and my life depended on it I would choose the latter.

Why not adopt the crypto (as in coins) seed thing with random words? Those are much more human readable, imo than "bubu bibi baba".

NitpickLawyer··on OpenAI's GPT-6 Astra on ARC-AGI-3
AFAICT nvda's result is on the 25 open problems, while this submission is on the "semi-private" set, ran by the arc people themselves.
Page 1 of 34Next →