HNHacker News
TopNewBestAskShowJobs

405error

28 karma · joined August 18, 2026

submissionscomments
405error··on How An AI math breakthrough ignited a controversy
But consider they could decide that they want to scoop more regular research work too. They could automate it with just a few LoC. Even if you opted out in the ToS, you'd have to file a massive lawsuit just to enforce it. And the actual fine would be inconsequential to OpenAI.

I think going forward, any researcher should consider anything submitted to an LLM to be copied/stolen.

405error··on How An AI math breakthrough ignited a controversy
Personally this is a watershed moment for researchers and grad students I know. All of them are close sourcing WIP repos, not putting their progress in LLMs, or have lab level initiatives to self host models.
405error··on The Navier–Stokes Millennium Prize Problem
I can give you some context. 1. Terence Tao's mastodon explains the way this problem was solved does not in itself contribute much. LLMs (and in this case) produce massive, often unintelligible proofs that do not further understanding. It is often that in pursuit of solving these problems, many other discoveries are made. 2. There is a more serious question about scooping. If OAI is using chat data from researchers to make discoveries, essentially every researcher who chats with an LLM can get scooped. You could be 80% of your way to solving a problem, and LLM could solve the remaining 20%, and get all the credit. Years of your work could be scooped in an instant. If you're a PhD student, this is even worse. Here it's a world famous problem. But imagine you're a PhD student, working on your small but extremely career/progression critical problem, and you get scooped by an AI you talk to. No one is even going to care.
405error··on Navier-Stokes – Tristan Buckmaster [pdf]
We know they are training on private chats. It's listed in the ToS.
405error··on Navier-Stokes – Tristan Buckmaster [pdf]
Tristian's allegations are much more serious than academic slap-fighting. If what he suggests is true, every academic using AI is going to get scooped. Yes AI can do non-trivial work, but the situation is that you could be a PhD student 90% of a way to make a major breakthrough. Then OAI scoops up your chats, dumps ten million tokens, and claims it for itself.
405error··on Navier-Stokes – Tristan Buckmaster [pdf]
Many academics and grad students I know have closed source their in progress work, and started being really careful about what they chat with LLMs (or using local ones) because of the drama around this. No one wants four years of their life getting sniped by ten million dollars worth of tokens.
405error··on Navier-Stokes – Tristan Buckmaster [pdf]
It would not be difficult to write a pipeline to remove 99% of low quality posts, especially about specific subjects. It would be very easy to identify accounts as researchers based on their chat logs.
405error··on 4.5B Posts Scraped from TikTok
It's probably AI coded and hallucinated many things. That drumming up the importance of a minor thing is a real tell. Another hallucination - it hasn't found any video APIs (despite statements that it has and uploaded it). It has video metadata.
405error··on 4.5B Posts Scraped from TikTok
More directly, there simply aren't any video files uploaded. It's just parquet files, which contain no video columns (I'm not even sure if it supports it).
405error··on 4.5B Posts Scraped from TikTok
It's that mix of dense, impressive sounding jargon, but even scanning across it raises glaring problems. Like, if you have 4.5 billion videos on HF, and it's 289GB, it's about 60 bytes per video. Checking the column fields as well, there doesn't seem to be any video files*.
405error··on 4.5B Posts Scraped from TikTok
The first immediate smell is that if you have 4.5B rows and 289GB in data, you have ~60 bytes per row.
405error··on 4.5B Posts Scraped from TikTok
I cannot verify whether it is technically correct, but it's about how to defeat Tiktok's bot filters to scrape it.
405error··on Cultivating a state of mind where new ideas are born (2023)
I have a vague recollection of a point someone made before that most ideas are usually wrong or incorrect at the start. For example, Nvidia was built on the thesis that non-triangular polygons would dominate, which was 100% wrong (today all polygons are triangular). But they often develop into something true.
405error··on A quick look at zero-knowledge proofs
Aren't several cryptocurrencies built on ZKPs? Their business model aside, ZKPs do look like one of the rare examples of theoretical elegance and real world use (even if not widespread).