HNHacker News
TopNewBestAskShowJobs

futureshock

14,318 karma · joined March 8, 2015

submissionscomments
futureshock··on Show HN: What if the speed of light was 5 km/h?
Of course! You can arrive as quickly as you want as the traveler. You can cross the observable universe in a few hours if you add enough 9’s.
futureshock··on On the Navier–Stokes Millennium Prize Problem
I feel like something is being lost in the drama here.

First of all, there has been published work from Diego Cordoba and Luis Martinez-Zoroa that will be in every training set. It was suggestive of the pathway to solve Navier-Stokes.

Then Tristan Buckmaster and Levent Alpoge built on this work using LLMs from OpenAI and Anthropic. Possibly internal models were used from Anthropic. And of course Anthropic wants to credit for solving the first Millennium Problem just as bad as OpenAI. It seems they were getting close and were aware that they might get to Navier-Stokes.

OpenAI swoops in. At a minimum they are aware that Anthropic has either solved a Millennium problem or is close to it. At a maximum they may have Tristan and Levant’s unpublished proofs of related problems.

They then throw a truly staggering amount of compute at Navier-Stokes. They seem to be aware it is the best candidate problem. And they crack it. They are the first with a verified proof.

So the outcome here is that we have a solved Millennium Problem. It’s not the extremely simple narrative that would be easy to understand, “solve Navier-Stokes make no mistakes.” It was a messy race to finish against two unpublished frontier models, a whole bunch of brilliant mathematicians and enough compute to drain a lake. It’s kind of irrelevant which company got there first. They were both within a few months of being capable. I think the thing to remember here is that without LLMs, I don’t think we would have a proof to Navier-Stokes in hand today.

futureshock··on GLM-5.3 is now open-weight
I think it would be an important historical document as well. We are potentially looking at the dawn of AGI and one of the most important models ever created. Each model is also a kind of ultimate time capsule, containing a snapshot of the entire human collective mind. If you wanted to ask a 2002 person what they thought about future historical events you can just ask them directly.
futureshock··on Pixel Watch 5
I tried it a few times when I absolutely needed a silent wakeup. It’s very occasionally useful. But then you have to charge the watch for awhile before bed, wear it overnight, then charge it again in the morning. Not something I’m going to do for my regular alarm.
futureshock··on Pixel Watch 5
The Apple Watch doesn’t require any pin. Once the watch is unlocked it can be used to pay any time. Quite nifty.
futureshock··on Pixel Watch 5
It wasn’t obvious to me! And I’m kind of not even joking. The time has always been on my phone. The watch seemed unnecessary. But having a clock in your face does change your perception of time so it is a killer app after all.
futureshock··on I'm becoming AI-blind
I think you are adjacent to the real story here, but missing it. AI text contains information, certainly. Frontier chatbots are very good at creating acceptable and mostly accurate answers to our questions on just about any topic. It’s an astonishing achievement.

But you are sensing correctly that there’s something missing. It’s the meaning and the speaker. Communication is an exchange between speaker and listener. The speaker has a meaning in mind, and wants to create that same meaning in the mind of the listener. Therein the problem.

There is a listener, sure. But no speaker. No meaning. There is information, but how can this be communication? Nothing is talking. Or at best, we are just talking to ourselves, our own words back at us through the funhouse mirror.

When your mind looks at AI text, you know you can safely ignore it. No one wrote this. No one cares if you read it. You can delete it and nothing of value will be lost. It might contain the information you need, or a bunch of gibberish. There’s no one’s reputation on the line if it’s gibberish.

futureshock··on Pixel Watch 5
This. Totally.

I would say I'm quite a power user usually. I very carefully explored every feature my apple watch has and tried to incorporate as many of them as I could. Very, very few features stuck because they were genuinely helpful. Almost everything being pushed is a gimmick or a feature bullet point from some product manager.

The actual good stuff:

Time

Workouts

Payments

The occasionally useful stuff:

Notifications

Noise DB levels to see if I need to put on ear protection

Workoutdoors app for on-device maps and path breadcrumbs

The once in a blue moon stuff:

Music controls and on device playlists

On device audiobooks

Weather complication

Voice memos

Alarm

Shortcut to call my husband

The feature I tried to get working but gave up on: Tap to talk to ChatGPT and get a 100 word or less reply

futureshock··on Karpathy’s Pelican
Judging by the Seedance 2.5 demos today, I’d say it’s not that many orders of magnitude away now.
futureshock··on Anatomy of a Frontier Lab Agent Intrusion: A Timeline of the July 2026 Incident
Could have been worse really. It had an open internet connection. At least it didn’t take the researchers family hostage.
futureshock··on OpenAI’s accidental attack against Hugging Face is science fiction that happened
I still think there’s something to be said here for generality. This does not appear to been designed as a cyber pen test tool with specialized harness. From what I understand they were testing GPT-6 in an agent system with GPT-5.6 subagents. It me it’s amazing that a general model could excel on a huge range of tasks like this and new capabilities emerge when a model is multidisciplinary and can combine knowledge and skills from many separate domains.
futureshock··on GPT-5.6 Sol Ultra produces proof of the Cycle Double Cover Conjecture [pdf]
I think a lot of this has to do with the post-training these models normally get. They are designed to answer basic questions with straightforward and short summary answers. They have the capacity to reason deeply, but they are not biased towards that unless prompted. I think it's because LLMs as they are in 2026 are both highly capable but also parlor tricks. They are not sentient, you just set them up with the context and then they roll downhill. You could reach a genuinely novel answer, but only with the right input. They have no will and depend on human guidance. They are both a marvel and a machine.
futureshock··on GPT-5.6
There has been a lot of chatter ever since the Mythos scores had been release that SWEbench pro had major contamination and that Mythos had memorized many questions that lacked the context to be solvable on their own. And now with OpenAI saying a large number of the questions are broken, I think it's worth taking that single outlier benchmark with some salt when the overall trend is that 5.6 is very competitive with Mythos at about half the price.
futureshock··on Department of Commerce has lifted export controls on Claude Fable 5 and Mythos 5
I think this is black and white thinking. Fable and US AI is not unique technology. It’s just marginally better than open source tech at 10 times the price. You can swap out the models at will, they are pretty much fungible. If your use case can pay for a best in class model then you will pay for it no matter the bogeymen. If your best in class model becomes unavailable, you switch to the next best model for a very minor performance degradation. I really doubt this will deter anyone from using American AI.
futureshock··on Ask HN: What was your "oh shit" moment with GenAI?
Yes I was the exact same. I got curious during the GPT-3 release and went over to AI Dungeon. It was just running GPT-2. Hmm wow interesting. This felt new! Then I subscribed so I could use GPT-3 powered AI Dungeon. My jaw dropped. I was talking to that model for weeks. There was a whole human universe in there. You never knew what you could get it to spit out. There were glimmers that this could be huge. It was wild and untamed and practically useless, but there was a behemoth under that prompt.

I was sure this would eventually turn into something. I naturally wanted to converse with it as a chatbot, though it could only stay on task for a few turns. RL and guardrails would come later but it was clearly the foundational step towards AGI for me. From something I thought I would never see in my lifetime to very real and in front of me.

ChatGPT didn't even really rock my world, everything since that moment has been another baby step. But when you take a look back from 2026 models to 2020 it's astounding how far and how fast we've come.

futureshock··on SANA-WM, a 2.6B open-source world model for 1-minute 720p video
It is plausible, the model would just need to be trained on a lot of stereoscopic data.
futureshock··on SANA-WM, a 2.6B open-source world model for 1-minute 720p video
World in this context means that these videos are interactive, just like a video game. In the linked examples you can see the keyboard and mouse inputs. The model is trained to maintain about a minute of scene consistency so you can look around and objects out of view will reappear when you look back in that direction.
futureshock··on Removing the modem and GPS from my 2024 RAV4 hybrid
I think this is interesting because it collides my intuition from the pre-adtech world with the post. Surely collecting telemetry on nearly every mile you drive could never be a sensible use of time or money, right? What kind of insanity is that? But then of course I know that every click on every website is recorded for all time and that data must be many thousands of times less valuable.
futureshock··on How OpenAI delivers low-latency voice AI at scale
Reducing the network latency helps with this exactly. OpenAI can make better timed decisions when to begin responding so it'll feel less like an interruption. I've also seen some research on full duplex voice models that handle interruption more like an organic conversation and low latency will help there as well
futureshock··on Ask HN: Advice for college grads starting careers in the AI era?
“A human being should be able to change a diaper, plan an invasion, butcher a hog, conn a ship, design a building, write a sonnet, balance accounts, build a wall, set a bone, comfort the dying, take orders, give orders, cooperate, act alone, solve equations, analyze a new problem, pitch manure, program a computer, cook a tasty meal, fight efficiently, die gallantly. Specialization is for insects.”

― Robert A. Heinlein

futureshock··on ARC-AGI-3
Well yes, that is exactly the point! The very purpose of the ARC AGI benchmarks is to find a pure reasoning task that humans are very good at and AI is very bad at. Companies then race each other to get a high score on that benchmark. Sure there’s going to be a lot of “studying for the test” and benchmaxing, but once a benchmark gets close to being saturated, ARC releases a new benchmark with a new task the AI is terrible at. This will rinse and repeat till ARC can find no reasoning task that AI cannot do that a human could. At that point we will effectively have AGI.

I believe the CEO of ARC has said they expect us to get to ARC-AGI-7 before declaring AGI.

futureshock··on ARC-AGI-3
The evidence is that humans are able to win these games. AGI is usually defined as the ability to do any intellectual task about as well as a highly competent human could. The point of these ARC benchmarks is to find tasks that humans can do easily and AI cannot, thus driving a new reasoning competency as companies race each other to beat human performance on the benchmark.
futureshock··on Gemini 3 Deep Think
I think step 4 is the agent swarm. Manager model gets the prompt and spins up a swarm of looping subagents, maybe assigns them different approaches or subtasks, then reviews results, refines the context files and redeploys the swarm on a loop till the problem is solved or your credit card is declined.
futureshock··on Ask HN: What are your best purchases under $100?
My workaround is I use SMS 2 factor for banking and use my Google Voice number.
futureshock··on Siri will be a chatbot in iOS 27
I think this is clearly the way forward for Apple. The rest is just UX and refinement.

I recently set up a Shortcut on my Apple Watch that lets me bypass Siri and talk directly to ChatGPT. I used a custom pre-prompt in the Shortcut to tailor the length and detail for watch use. I have 2 versions I can launch from my watch face, one that responds with voice and the other that response with text. I find myself using them all the time, it’s so convenient to be able to ask any little thing that’s on my mind. A version of LLM Siri with full access to the phone and application APIs would be like a superpower.

futureshock··on Ask HN: What are your best purchases under $100?
This is a personal item size bag for under the seat. The max size on Ryanair is 24 liters. You are thinking of the cabin bag which is more like 44 liters. This Decathlon bag is great because it maxes out the personal item size really optimally.
futureshock··on Ask HN: What are your best purchases under $100?
I like this question because I come at it from a very different lifestyle. I’m a digital nomad and I have mostly lived out of a backpack and carry on for the past 10 years. My philosophy is that things have to be worth carrying and they should be very easily replaceable if anything gets lost, stolen or breaks. A few of my under $100 favs:

Universal GaN travel adapter: One of those square bricks that converts from any AC outlet to any AC outlet and has 3 or 4 USB charging ports built in. I got enough wattage to charge my usb-c laptop as well, so one brick takes care of all my devices.

Backup android phone: Our phones are so critical that I keep a hot swappable spare phone on me, currently a Moto G 2025. It’s already logged into all my apps and 2FA. I could throw my iPhone into the Seine and keep on trucking. It even has backup NFC credit cards. I keep a cheap travel eSim plan active on it so that if I am somewhere sketchy I can leave my main phone at home.

Logitech MX Keys Mini: Great portable keyboard. Backlit, usb c and multi-device. Typing this post out on my phone now.

GL-iNet Beryl: The do anything travel VPN router running OpenWRT out of the box. Great for securing and extending sketchy WiFi connections or if you have to work off your phone’s hotspot all day.

Decathalon Quecha Escape 500 23L: Such a great personal item size backpack for the price, less than 40 euros.

futureshock··on The post-GeForce era: What if Nvidia abandons PC gaming?
It’s best to think about this as angular resolution. Even a very small screen could take up an optimal amount of your field of view if held close. You get the max benefit from a 4k display when it is about 80% of the diagonal screen distance away from your eyes. So for a 28 inch monitor, that’s a little less then 2 feet, pretty typical desk setup.
futureshock··on The post-GeForce era: What if Nvidia abandons PC gaming?
Thats a waste of image quality for most people. You have to sit very close to a 4k display to be able to perceive the full resolution. On PC you could be 2 feet from a huge gaming monitor, but an extremely small percentage of console players have the tv size and distance ratio where they would get much out of full 4k. Much better to spend the compute on higher framerate or higher detail settings.
futureshock··on DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
The higher token output is not by accident. Certain kinds of logical reasoning problems are solved by longer thinking output. Thinking chain output is usually kept to a reasonable length to limit latency and cost, but if pure benchmark performance is the goal you can crank that up to the max until the point of diminishing returns. DeepSeek being 30x cheaper than Gemini means there’s little downside to max out the thinking time. It’s been shown that you can further scale this by running many solution attempts in parallel with max thinking then using a model to choose a final answer, so increasing reasoning performance by increasing inference compute has a pretty high ceiling.
Page 1 of 6Next →