HNHacker News
TopNewBestAskShowJobs

irthomasthomas

3,756 karma · joined December 29, 2019

undecidability.com

crispysky.com

x.com/xundecidability

github.com/irthomasthomas/llm-consortium

submissionscomments
irthomasthomas··on Musk, the Movie
Paypal was invented at Confinity which was part owned by Thiel and had nothing to do with Musk. Thiel bought Musk's x.com in a merger deal that also made Musk ceo of the new company for 6 months before he was fired for incompetence. Beyond the windows server transition he also insisted on rebranding paypal to x.com, the everything app. I don't know where he was mostly right at paypal?
irthomasthomas··on OpenAI is enlisting an influencer army to make it look 'good for the world'
He was kicked out of Canada because they thought the Technocracy movement which he promoted was too communist.
irthomasthomas··on Kev: Tiny Jev-like family of decision models built on top of Qwen3.5
And Claude often identifies as Qwen or Deepseek when prompted in Chinese.
irthomasthomas··on What happened to the Snowden archive
Except for all the exceptions, which are many, like the Jack the Ripper police files which where denied to the public.
irthomasthomas··on GPT-6 Astra Solves a WWI German Radio Cipher
I doub't it. The main issue is not cost, though they do get expensive as context grows, but intelligence. A frontier model like fable becomes as dumb as haiku after 200k tokens. They have been stuck at ~1M context/200k useful context for 18 months, now, with little sign of advancement. A model with a 10M context window that retains it's intelligence up to 2M tokens would be a big breakthrough.
irthomasthomas··on AI safety is mostly a sex cult
I see, thanks. It reminds me a little of those spam emails which are sprinkled, intentionally, with obvious spelling errors. They aren't trying to trick the average person, they are trying to filter for a much smaller, more valuable audience.
irthomasthomas··on AI safety is mostly a sex cult
What are your favourite quotes from it?
irthomasthomas··on Why I didn’t sign the Fields medallists’ letter
It is getting a lot harder for those people to justify using OpenAI to assist such endeavours. Afterall, OpenAI might just front-run you if they hear a rumor you solved some marquee problem that they can brag about in PR campaigns.
irthomasthomas··on Gemini 3.8 Live and 3.8 Live Extended Thinking
artificialanalysis just updated their benchmark after the release of GPT-6. They removed old, saturated benchmarks and replaced them with new, until GPT-6 floated to the top with the cream. One of those new benchmarks is AutomationBench-AA, where GPT-6 had a clear lead. Today that benchmark is topped by DeepSeek v4.1 Flash.

Edit: For those who are not familiar with it, this model is quite a bit faster, and about 100x cheaper, per token, than Fable and Astra.

irthomasthomas··on Mullenweg has returned as CEO after attempted board ouster
It took them two years to finally get him out with the help of the government in Shenzhen

https://www.nme.com/news/arm-china-finally-ousts-rogue-ceo-t...

irthomasthomas··on Real-SWE: Benchmarking AI models on private, real-world, enterprise codebases
> For instance the task naming in the task file starts with an optimistic 1, 2, 3, 5, 5a but then eventually gets to 8a, 8a1, and then ends up with 8b2c2b3 and “8b2c2b2b checkpoint1”. The code that it produced got ever more wild. I don’t want to bore you with what it tried to build, but here are some example pieces of the interpreter changes:

  Hardcoded constants everywhere
  Multiple same-line macro invocations in C
  Random indexes in production code
  Hideous tokenizer code in C
https://lucumr.pocoo.org/2026/9/7/astra-why/
irthomasthomas··on Anthropic boss Dario Amodei calls for AI development to slow down
Astra scores the same on DeepSWE 1.1 (~75%) as Gemini Flash 3.8 and Deeepseek Flash 4.1 So general coding ability has plateaued, for now. Also consider the context windows. 1M token models where a breakthrough two years ago. Today they are still limited to 1M. In fact, if you don't want intelligence to drop off a cliff, you are really limited to 200k tokens.
irthomasthomas··on More questions about whether researchers can trust OpenAI with unpublished math
Is there a reason they scoped that so narrowly to Buckmaster/codex/2 months

two people worked on this for a year before the breakthrough. Perhaps that earlier work reduced the search space sufficiently to brute force the problem with 10,000 agents?

irthomasthomas··on DeepSeek v4.1 Flash
Not on par, but in the same league. Astra is way ahead on visual tasks, but scores the same as gemini and deepseek on DeepSWE.
irthomasthomas··on More questions about whether researchers can trust OpenAI with unpublished math
Openai said that a new model became available to them during this. But that could mean anything from a big new base model to a LoRA, fine-tuned on a few dozen prompts...
irthomasthomas··on DeepSeek v4.1 Flash
hmm I'm hoping there is a bug on their API because my first impression is not good. I asked it to return bash code between <bash></bash> tags. It is failing frequently and writing it's own tool calling format instead.
irthomasthomas··on DeepSeek v4.1 Flash
Quite a flex calling their GPT-6 competitor "Flash"! But it is faster than their last flash model due to a combination of architectural innovations including engrams and a new encoder/decoder design that uses 8B parameters for prefill and 16B for generation.
irthomasthomas··on Qwen 3.8 follows GPT-5.5 Pro reasoning prefills
Has the method for extracting the COT been blocked, now? Otherwise why could we not generate some fresh samples?
irthomasthomas··on The Navier–Stokes Millennium Prize Problem
Chutes.ai models are served from a Trusted Execution Environment, so the GPU owners can't see your prompts.
irthomasthomas··on The Navier–Stokes Millennium Prize Problem
And deliberate or not it is still plagiarism by the sound of it.
irthomasthomas··on On the Navier–Stokes Millennium Prize Problem
The researcher told them it was an independent effort, and they still pushed ahead with it.
irthomasthomas··on Mercury 2.5
This should make an excellent choice for arbiter in llm-consortium, mercury-2 was pretty good. One of the main drawbacks of the multi-model system is the added latency of the llm judge, but having a model run at 1100tps goes a long a way to alleviate that.
irthomasthomas··on On the Navier–Stokes Millennium Prize Problem
If the goal was not to scoop them, why did openai put a massive team on this, working weekends, only after they heard rumors of the solution?
irthomasthomas··on On the Navier–Stokes Millennium Prize Problem
Or they trained a LoRA on the victims chats in order to launder their plagiarism.
irthomasthomas··on On the Navier–Stokes Millennium Prize Problem
It can still be academic plagiarism even if they ticked the box to allow training on their prompts.
irthomasthomas··on On the Navier–Stokes Millennium Prize Problem
Doesn't that count as plagiarism?
irthomasthomas··on On the Navier–Stokes Millennium Prize Problem
"When a further trained version of our internal model became available over the course of the effort, we updated our agents to that model."

woah, this gives some credit to the rumor that openai finetuned a model over the course of a few days for this task, and maybe trained on the Chatgpt/codex history of the authors, including drafts of this research.

irthomasthomas··on Navier-Stokes – Tristan Buckmaster [pdf]
They where working on the problem for a year using codex.
irthomasthomas··on Navier-Stokes – Tristan Buckmaster [pdf]
Attack is the best form of defence
irthomasthomas··on Multi-Agents LLM Financial Trading Framework
Is it opensource?
Page 1 of 34Next →