HNHacker News
TopNewBestAskShowJobs

NitpickLawyer

5,630 karma · joined September 28, 2024

submissionscomments
NitpickLawyer··on OpenAI's GPT-6 Astra on ARC-AGI-3
Since low scored much lower than none, and none scored ~ around medium, could none default to medium in the API? I don't think the new models can even have "instant" via API, unless they train them for that (there was one gpt5 variant called instant or something).
NitpickLawyer··on ChatGPT outage – Resolved
Perl6: say "Fizz"x$_%%(2+1)~"Buzz"x$_%%(4+1)||$_ for 1..100

from here - https://github.com/rsha256/shortest-fizzbuzz/blob/master/Per...

NitpickLawyer··on Gemini 3.8 Flash and 3.8 Flash Cyber
If anything, gemini models are the least benchmaxxed out of any lab, IMO.
NitpickLawyer··on The Emergent Symbolic Structure of Artificial Neural Networks
I think you accidentally a word, there. GP is talking about comprehending the capacity of massively multi-dimensional space.
NitpickLawyer··on GPU World
It really isn't and it's sad seeing so many people say it so confidently on this site. It only detects plain / basic prompted stuff. "write me an essay on x", sure. The moment you prompt it differently, it stops working. Add "output should be STE100 compliant" or something to your prompts and detection goes away. Fine-tune any local llm on real human prose, and detection goes away. Edit 2-3 characters (emdashes, lists, etc) and detection goes away.

And that's for just basic detection. There have been plenty of examples of 100% human written content (either old, or unpublished) that gets falsely flagged as AI.

And these are just technical aspects. The main issue is that pangram and other solutions are being used to summarily judge students work, and that is orders of magnitude more fucked up. Accusing someone of cheating can have devastating effects on their education/career/etc. and they're doing it with snake-oil closed boxes, at scale. We really really shouldn't support this, especially here on a technical site.

NitpickLawyer··on GPU World
I just finished reading Service Model by Adrian Tchaikovsky [1], a really timely novel that deals with lots of open ended questions of AI, robots, humanity, control, self determination, and so on. Really recommend it if you're into these kinds of things.

> Premise Imagine that AI frontier progress stops as of 1 September 2026: AI becomes faster and cheaper, but it never becomes superhuman or improves considerably across the board.

I absolutely love this premise. I wrote a comment a few days ago about this very thing. I've had this "revelation" in early 2024, when using a small local model, that even if models never improve, I'd still have a few years of discovering all the things I could do with the models.

Really cool, hopefully we get some interesting and not overwhelmingly pessimistic stories out of this project. There's enough cynicism and negativity in the world. We could use some funny takes on everyone using openclaw2040 ran by fable. Shenanigans galore.

[1] - https://www.goodreads.com/en/book/show/195790861-service-mod...

NitpickLawyer··on Separating logic and language
First you'd have to come up with a commonly accepted definition. By some ~16 years old definitions from famous experts in the field, we've already achieved it. By today's definition (of the same expert) we haven't. There are people today that think we have it. There are people today that think we'll never have it. It's such a loaded term / topic.

Also, regarding LLMs, 2 years ago another famous expert in the field was adamant that LLMs can't plan, can't do math and they can't work at long contexts. Obviously they can do all of that now. So, really, no one knows.

NitpickLawyer··on Separating logic and language
> plus fMRI on healthy volunteers solving the same puzzles silently

Are they looking at blood flow in areas to map "language network" and other stuff? I remember a few recent papers that found that a) blood flow doesn't necessarily mean "active area" and b) two brands of fmri found completely different "areas" activating for the same patient + same puzzle.

NitpickLawyer··on SK Hynix CEO sees memory chip shortage lasting until 2030
I think that a lot of people miss key aspects of the AI boom. Even discounting the skeptics, and the bubblers, crashers, etc.

There are a bunch of things that happen in parallel to the AI boom:

a) everyone and their mother is building out capacity. For each GB of VRAM installed, you generally want at least 1 GB of RAM, if not more. So just take nvda's numbers and go from there. And that's just for the GPU servers. But you also need ancillary services around GPU farms. You need some compute for filtering, preprocessing, you need data storage, and so on. And they're all RAM hungry processes.

b) RL is giving lots of capabilities improvements for AI, but it needs lots and lots of parallel runners for scenarios. Those runners need RAM. So every bit of non-GPU capacity is also being used, and more installed.

c) With increasing capabilities comes increasing usage. Everyone is deploying "agents" and those need to run somewhere, and they take RAM (especially the badly coded ones). There've also been reports of labs / big players hoovering up every mac available, for "agentic computer control" / training / whatever. That's on top of people buying them to run "24/7" stuff a la claw/hermes/etc.

In all of this, RAM is the bottleneck. And for the skeptics, you don't have to take hynix's word for it. Just look at the sales for CPUs / Mobos. They're down, because building a computer now is bottlenecked at RAM.

NitpickLawyer··on Hy4 preview
Weights are not binary. A model is created at init time, with random values. After that, it is being modified using data. The key point is that the labs modify the models "as weights". That means that weights are the intended / preferred way of modifying a model. Which, coincidentally, matches the definition of source in Apache 2.0. There is no "higher level" place where editing takes place. It all happens in weight space. Through the license you get the same rights as the lab that created it: view, inspect, run, modify, re-release. That's it. That's the only thing a license can grant you.

The rest is semantics, misunderstandings, and FUD. A model released under an open source license is open source. Training data is lab knowhow / IP. Which, historically, has never been required for any open source release.

NitpickLawyer··on EVE Online moves to Python 3
> Does the game have performance issues?

Yes, it has had huge concurrency issues for the entirety of its life. Their solution to large fights has historically been "let us know in advance pls", plus "move systems to beefier hw nodes" and "tidi" which stands for time dilation, where the "tick rate" of the whole server goes down and a fight takes 10-20-100x longer than it should.

It's an amazing concept of a game, but software wise it has been a mess since forever.

NitpickLawyer··on Debian has published the official results for the 2026 GR on LLM usage
> That supports “decisive win” comfortably; whether it qualifies as a “landslide” depends on where the margin threshold is set.

I think the landslide is "pro" vs. "against". The two most "against" options got the lowest rank. Top 4 are all "pro" (with different caveats for each).

  Rank Option 
  θ
  i
  1 5: Responsible Use of Generative AI 1.794
  2 2: Allow AI-Assisted Contributions with conditions 1.430
  3 6: A cautious approach to generative AI 1.424
  4 4: Accept AI contributions for Debian specific work 1.252
  5 8: Avoid LLM use: climate impact 1.007
  6 7: Debian is created by humans 0.881
  7 9: None of the above 0.762
  8 3: Reject LLMs as far as practical 0.683
  9 1: Ban LLM contributions via Social Contract 0.474
NitpickLawyer··on GLM-5.3 is now open-weight
Or if you need stuff that APIs don't / can't provide. Or for future proofing your workflows. Running things locally gets you "the same thing" in perpetuity, while APIs might change, models can be deprecated and features can be removed.

Cybersec is also hit and miss, depending on what provider you choose, verification systems and all that jazz. Also, running locally allows you 100% data privacy, in any situation and for whatever usecase you might have. ~100k for hardware for a small team of devs to code locally is not that expensive in the grand scheme of things.

Lastly, local models allow for training / finetuning on your own data and processes. $/tok is not everything for everyone. Sometimes you can take a hit on value / speed if you get something else that matters for you.

NitpickLawyer··on Autonomous Mathematical Discovery in an Open-World Multi-Agent Environment
> We study autonomous mathematical discovery in the Station, an open-world multi-agent environment in which AI agents from different model families pursue a shared research goal without a central coordinator or scripted pipeline. Agents choose their own research directions, conduct experiments, collaborate, and build a shared scientific literature. Across 12 construction problems from the AlphaEvolve catalogue and two additional case studies, the Station obtained results novel relative to the prior literature on five problems: a new infinite family of finite-field Kakeya sets, new exact 604-point kissing configurations in dimension 11, new records for the discretized Kakeya needle and sign uncertainty problems, and a substantially improved lower bound for Erdős's minimum-overlap problem. Agents also discovered novel infinite families for Book Ramsey numbers. Importantly, the agents produced not only numerical constructions but also theorems and analyses explaining how those constructions work, making the results more interpretable and easier for mathematicians to build upon. We release all raw agent dialogues, proofs, and verification code, providing a transparent record of how these discoveries emerged.

(emphasis mine)

For the last few months, every time a new "famous problem" was solved, there were numerous comments saying variations on this theme: "well, yes, but how about novel stuff, how about new things, original work, yadda yadda". Curious what the "next thing" will be now.

NitpickLawyer··on AI Agent Has Root
> Having a GUI on the thing lets me leave all kinds of things running persistently in the background that might be bothersome if interrupted running on my laptop.

Unless you actually need GUIs, you could just use screen/tmux or the newer versions like zellij/etc.

NitpickLawyer··on Small Models Have Arrived
> But I also think the demand for "fast/cheap/good-enough" models is just about to take off.

There's a sort of "revelation" I had in ~early '24 when I used a 7B local model with a library called Guidance (initially out of MS, then the team moved) to create a flow where the model would receive pseudocode for tests, first write the tests, and once I approved then started writing code until the tests passed. This was before "thinking" models, and yet using that library I was able to "guide" the model in the required "prompt / instruct" context such that it was working towards completion, and I saw the first things like we see now in the thinking traces "oh, test x doesn't pass because blah, I need to..." and so on.

Anyway, the revelation was "even if the models never improve, I'll have years of fun finding out all the ways I can use these things". And, obviously, the models improved a lot since then. But I think that revelation can still be applied, as a sort of "truism". We have, right now, access to things that 10-20 years ago would be considered magic. We are still finding ways of cobbling together systems with glue, duct tape and prayers and find new things they can do.

I think the "good-enough" stage has come not just for API models (cheap, fast, etc) but for local as well. Even if slower, even if clunkier, but they are good enough for a set of ever increasing tasks, and what's more it's incredibly fun to work with them.

NitpickLawyer··on Microduck
Since it's programmable, I guess one could tinker with it to do just that?
NitpickLawyer··on The Hugging Face incident and the road ahead
> a human should've noticed and gotten involved

I think a lot of people miss the fact that the first message board was established during a training run. Those are ran at a scale where it's not feasible for anyone to "notice" or get involved. We're talking tens/hundreds of thousands/millions of scenarios going for hours each. At this scale all they can do is pray that their verifiers work, and the rewards match their intentions. No lab has the capability to "check in" on what the traces look like, unless some system alerts them (loss spike, crashes, etc). Other than that, it's prepare, train, asses, restart.

Then, the hf incident was during an eval run, but the model that was evaluated was trained with the notion that there is a way to communicate between agents, and re-popped artifactory and re-established communication. That phase had more chances of being spotted, but anyway... lessons learned.

NitpickLawyer··on Qwen3.8-Flash-Next
It is 125B A6B. vLLM is already out with support, ngrams can be offloaded to RAM so you only need ~96GB VRAM for nvfp4 w/ full context.

Likely soon we'll see nvme offloading for ngrams as well. They're just an index, so that should be plenty fast for what it does. LLama.cpp support should come soon as well, and they might do some things with offloading first.

NitpickLawyer··on Qwen3.8-Flash-Next
This particular release is interesting because it's a preview of qwen4 architecture. And, while benchmarks are iffy, this is a direct comparison, by the same team, with qwen3.8-27b that was pretty well received for a local model.

This "next" release adds a new concept, first public release with n-grams, I think. And it's in a MoE size that is likely to be very fast and cheap to serve (faster than 27b for sure). It's also well suited for inference on alternative compute (i.e. sparks, macs, etc) so it's relevant to local users.

NitpickLawyer··on Z.ai confirms Ox Alpha is a new GLM-series model and will release its weights
> in which timezone?

Apparently someone working at a 3rd party inference provider also got confused and posted confirmation about it being a glm-flash model, despite having an embargo on that info. Someone jumped in the comments and told them they missed the timezone :)

In any case it should be releasing in a few hours. Timezones are hard.

NitpickLawyer··on More than half of adults in U.S. say they lack basic statistical understanding
I'd say in reality it's way more than that. The first statistics course in uni was a very humbling experience. I realised that while I thought I understood a lot (and I was coming from a CS heavy background, olympiads and such) real statistics is way harder and a lot more counterintuitive than I thought. Granted, this talks about "basic" statistical understanding, but even that is way more complicated than most people assume.
NitpickLawyer··on Qwen 3.8-Flash-Next releasing tomorrow (125B a6B)
A bit of context: 3.5 was the last version where they released their entire suite of models 2b-400b. Then 3.6 got a 27b dense and a 35b moe. Then 3.7 was API only, and 3.8 got only the 27b dense. The devs confirmed on twitter that 35b moe would not come. So that's what I meant by 3.8 is not getting a moe.

3.8 next is not really a 3.8 (but I guess they had to disambiguate from the previous next). It's a preview of qwen4 architecture (and it is an moe + ngram), released early as a preview, and to help the community sort out inference before qwen4 releases.

NitpickLawyer··on Qwen 3.8-Flash-Next releasing tomorrow (125B a6B)
This is what I copied from the en version of the modelscope page, right when they published it:

> Redisgned Multimodal MoE Model: 125B main model parameters, supplemented by an additional 51B N-gram embeddings,and 6B parameters activated per token.

> Efficient Training and Inference: Significantly reduces training and inference costs. At ~1/9th the training cost,Qwen3.8-Flash-Next achieves comparable capability against Qwen3.7-Plus, while being more capable in areas of coding and cowork.

There was another paragraph about a new attention, but I didn't copy that.

NitpickLawyer··on Qwen 3.8-Flash-Next releasing tomorrow (125B a6B)
Their "next" variants are usually undercooked, but useful for the community to verify support for inference stacks. This will likely be the same.
NitpickLawyer··on Qwen 3.8-Flash-Next releasing tomorrow (125B a6B)
They've said no moe for 3.8, and since they're already releasing a qwen4 early preview, they're probably focusing on that arch going forward.
NitpickLawyer··on Anthropic candidates face blunt money question
I think what's inconceivable is that anyone thinks that these kinds of questions really work. The intersection of high-skilled people that can work there with low-awareness people that can't navigate these questions in a "politically correct" (for lack of a better term) way must be near 0.

"Oh, of course I wouldn't mind stopping if we decide that something is too dangerous to pursue. I think the mere fact that we'd get close to knowing how to do it would satisfy my curiosity. Plus, if we'd be close to it, I'd probably take some time to spend with my loved ones, as the competition would likely be close as well."

It's not hard to see these kinds of questions as more than kabuki theatre. I'd put them in the same vein as those "do you intend to perform any terrorists acts while in our country" questions you get while travelling.

NitpickLawyer··on Anthropic's best AI model struggles to attract users as cheaper tools thrive
> The no-ZDR is clearly to permit surveillance.

It's to train better models. The three letter agencies don't need to spell it out in a ToS, they just access it if they want.

NitpickLawyer··on NanoGPT Speedrun Frontier
Cheap, fast and somewhat capable models are insanely effective at highly verifiable tasks, even if their overall capabilities are under SotA. You can leave them banging their tokens against a wall, and come back once their attempts are verified. And youc an always clean up afterwards, once a task is solved, if needed.

I've had a lot of success with dsv4-flash on these type of tasks, where it's easy to set a threshold for the task, and just loop it until that threshold is reached.

oAI's Luna play is really good. They've slashed the prices, the model is somewhat capable, and you can use it both for these kinds of long horizon tasks, or you can hand-hold a bit and get extremely cheap results out of it. And they get to keep devs in their own ecosystem.

NitpickLawyer··on RF Cafe
> throwback to the internet of my youth

It's interesting how simple advertising is as well.

> Please Support My Advertisers!

Some jpegs. No js, no moving shit, just logos, some text, some call to action in a banner format. That's it.

Also:

> A random number generator determines which company's Banner Ad will go in a particular position*. I wrote the code myself and it is very simple and foolproof.

> I do not collect page traffic or link click statistics for advertisements. Your website server statistical analysis is a much better reporter of inbound links from RF Cafe than I or anyone else can provide.

Nice.

> Banner Plan: $400/month

I count 22 advertisers, so there's obviously lots of incentives for advertisers to work with this site. Good for them, and nice to see it can still be lucrative for what I believe to be a single person operation. Quality content still matters.

← PreviousPage 2 of 34Next →