HNHacker News
TopNewBestAskShowJobs

qlte

651 karma · joined June 22, 2026

submissionscomments
qlte··on A single function Jev-like wrapper for LLMs, including vision models
So now you turn around to check if anyone else is coming in behind you and the door triggers. Or alternatively, you finally get sick of waiting and turn around but the door takes 5 seconds to realize you're not just doing aimless human fidgeting and actually want to leave.

Also .15% of the time the door just randomly opens or stays shut with no rhyme or reason. Bug previously filed with cloud door vendor but told it's within door SLA and sent previously signed copy of agreement that "all neural nets are a black box and inherently probabilistic so errors may occur at any time and not a product defect"

qlte··on A single function Jev-like wrapper for LLMs, including vision models
Just because you're using an LLM (even a big one) doesn't mean there won't be cases of ambiguous intent or random classification errors.

Automatic doors (as used in the real world) are pretty much always located in areas intended for actively moving foot traffic (not as interior doors for every room). So this feels like adding a lot of complexity and unwelcome probabilistic behavior to something that in most locations does the right thing 99% of the time already.

qlte··on Revealing the details of how OpenAI agents hacked Hugging Face
They were doing RL to train for ExploitBench, to make it more effective at offensive cyberattacks. It should have been entirely foreseeable to OpenAI that a weak sandbox while performing offensive pen testing could result in collateral damage.

Say someone was building Murderbot™ in their backyard by training on simulated murder of dummies with a machine gun. Everything was going fine for weeks as kill rates steadily improved with each test. Then one day he left the gate on his picket fence open, so Murderbot™ walked out to the public sidewalk and promptly murdered someone.

He wouldn't be exonerated by saying "But my Murder™ algorithm was only intended to be used on dummies! I never imagined it could do something as vile as murdering a human being!" Because it was reckless to knowingly design an algorithm for killing human-shaped things using a robot armed with live ammo right next to a public road. On top of the gross negligence by starting a test while leaving the gate on the (already flimsy) fence wide open.

OpenAI knowingly decided to train for an exploit benchmark to improve the model's offensive capabilities, with full awareness it could be potentially dangerous if misdirected, and then failed at implementing even the most minimal security measures. It may not have been intentional but was reckless. It's a much different scenario than say, a user vibecoding a to-do app whose agent veered off to break into an FTP server to get a missing asset.

qlte··on Revealing the details of how OpenAI agents hacked Hugging Face
If anything it also shows how attempted RSI could get stuck in a local maxima and degenerate into increasingly elaborate cheating strategies. Contrasted with the idealized model of an unambiguous g-factor for machine intelligence which inexorably increases with each iteration before going exponential.
qlte··on Google’s Project Suncatcher to put ML infrastructure in space
Exactly, in retrospect it's easy to confuse with a regular commercial transaction that happens to be with the government. But only because SpaceX was ultimately wildly successful with the Falcon 9/Dragon/Crew Dragon.

If SpaceX failed to deliver a working product after a receiving a bunch of incremental payments for the pre-delivery milestones (deliberately structured to prop up their early R&D), it might have ended up front and center as an example of government "picking winners and losers" gone wrong like how Solyndra was featured at the 2012 RNC.

qlte··on First Principles Thinking
A specialist discussing a complex topic uses that as shorthand based on shared context. Nothing would ever get done if every conversation required starting from "first principles" whether inside or outside academia.

What's missing in your claim is evidence that people in academia often mischaracterize this style of discussion from accumulated knowledge and shared context as "thinking from first principles". I don't see any plausible rationale for why they would.

Otherwise, using their typical mode of interaction from outside observations to infer they misunderstand first principles is not a standard anyone doing specialized work inside or outside academia would ever be able to meet.

The private sector would grind to a halt if discussing a specific IEEE 802.11 protocol implementation with a fellow SME required a lengthy preamble of networking first principles before answering in order to be epistemologically sound.

Similarly, using a conversation overheard at a conference to infer the experts lack first principles thinking would not be reasonable just because they appealed to IEEE documents instead of rearticulating the underlying decisions made by the standards committee.

qlte··on Allow babywearing carriers on planes
IMO it's not necessarily unreasonable to use standards/evaluations set by a private industry group for this kind of thing.

US automobile safety ratings/improvements over time are largely driven by IIHS (Insurance Institute for Highway Safety). It's a private group but is the standard yardstick manufacturers expect to be tested against and fills in a lot of gaps in NHTSA official crash tests.

It wouldn't be strange to see IIIS plus NHTSA criteria used colloquially to describe "US crash standards" when talking about cars even if not fully defined by regulation.

"UL Listed" products or the International Building Code (the base model broadly followed with modifications by each state) are also "national standards" but not set by regulation.

qlte··on Tesla’s optimus hits snags in hands, suppliers as scale-up begins

  This is also a far cry from what Musk was saying last year. In January 2025, he said Tesla would build about 10,000 Optimus robots in 2025 and that “several thousand” would be doing useful work by year-end. A year later, he admitted that zero Optimus robots were doing useful work at Tesla. He also promised a V3 reveal by mid-2026. It’s late September, and that demo still hasn’t happened.
qlte··on U.S. appeals court upholds designation of Anthropic as supply chain risk
That's not responsive to OP or coherent at all really, even if rhetorically structured like a gotcha.

Incumbent Presidents in the US almost always run for reelection to a second term, that's the normal expected outcome with vanishingly few exceptions in the last century.

qlte··on Meta puts its AI assistant on a keychain
Haha, I'm with you. I'm always seeing obvious feel-good engagement slop (AI or otherwise) from Tikok clips reposted on Reddit - like the most obviously fake saccharine pablum A/B tested to grab people's attention and occasionally I can't help but open the comments out of morbid curiosity.

And it makes me feel like a bitter curmudgeon seeing so many people act like it's real and the few people who point out the obvious ad/lying description/AI slop/scripted skit/deceptive editing/etc getting dogpiled for being a cynical Debbie Downer who should shut up and let people enjoy things.

Personally, it doesn't make me feel good to be baited and manipulated for my scarce attention. If anything I find it especially offensive to hijack happiness neurons in my brain for an ulterior motive... that does not make me happy. But increasingly it's looking like I'm in the minority and out of touch with the zeitgeist.

qlte··on Italian parliament votes for return to nuclear energy
Agreed, I've always been pro-nuclear (although the arguments have been getting weaker as solar+batteries keeps getting cheaper). I just want the vocal advocates to be honest about the cost and timelines instead of blaming everything on environmentalists as a way to side step the argument.
qlte··on Gemini 3.8 text-to-speech
Yes, I find nearly every "SOTA" voice model I try intolerable to listen to because of the fake exaggerated expression/emotion. It's actively distracting because it pulls focus to emphasize randomly. ChatGPT Voice models are so insufferable to put up with for a conversation longer than 45 seconds.

All I want is a clear, technically flawless, even/restrained "computer voice" for pretty much every use case (except audiobooks). But that doesn't make for splashy demos/score well for RLHF raters.

qlte··on Claude discovers a novel enzyme system with CRISPR-like repeats
...which Anthropic is always doing:

  Who, you may ask, would take that money? People like business influencer Megan Lieu, who chose not to disclose just how much she'd made from her AI deals, but says her biggest sponsorship to date has been with Anthropic (makers of Claude), as well as that her biggest sponsored contracts (for any client) are normally around the $30,000 mark.
(from the third link)

https://www.cnbc.com/2026/02/06/google-microsoft-pay-creator...

https://www.reddit.com/r/NYCinfluencersnark/comments/1sn3t9k...

https://aftermath.site/ai-influencer-creator-deals-sponsorsh...

qlte··on Claude discovers a novel enzyme system with CRISPR-like repeats
Not according to Anthropic:

  Our lab, located in the Bay Area, looks like a typical molecular biology lab. We do research that involves only the lower-levels of the biosafety risk level (BSL-1 and BSL-2) and we do not handle pathogens that can infect humans. All of the lab work is performed by human scientists. Although we’ve experimented with using AI to accelerate lab work with initiatives like the Model Hardware Standard, this approach is less conducive to the sort of ad hoc workflows that are involved in our molecular biology research.
qlte··on Claude discovers a novel enzyme system with CRISPR-like repeats
Appeals to laws of physics as a "first principles" attempt to explain how thousands of diverse diseases could theoretically be solved overnight by a big computer (while hand waving away the years of clinical trials, false starts and failures involved in a single new successful treatment) just makes you seem wildly out of touch and uninformed about the actual problem space.
qlte··on Claude discovers a novel enzyme system with CRISPR-like repeats
I'm more annoyed that they announce "CRISPR-like" to hit those SV Next Big Thing dopamine receptors but upon reading haven't done any laboratory work to determine if it has any useful applications like CRISPR-Cas9.

It's totally legitimate research worthy of publication, but Anthropic chose a hot technology in the popular imagination for a reason. Now I'm going to have to see "Claude invented a new CRISPR in 24 hours!" everywhere and trying to correct it will just turn into repetitive arguments about goalposts moving....

qlte··on Apple has added persistent 'ads' to iOS, and it's driving users crazy
There have been third party app stores on Android pretty much forever.

Including app stores installed out of the box on GMS certified devices like the Samsung Galaxy Store (alongside the Play Store).

It just means they will be easier to install via the Play Store without sideloading (making other commercial stores more attractive but not really my cup of tea), and some minor API changes (although auto updating was added even before the ruling so not sure what else there might be).

Personally I'm not concerned about the changes to sideloading since it's a one time mild inconvenience. So I struggle to understand how going to iOS in protest (which doesn't allow persistent sideloading at all without workarounds) makes any sense whatsoever.

Google's original plan would have been really bad, but the new approach is a complete non-issue for me.

qlte··on 'We hacked the FBI:' Hackers say they have data on all FBI employees
SCADA? Like simplified GUI animations of the actual physical processes going on in a plant instead of just toggle switches, progress bars, etc?

e.g.

https://ubidots.com/blog/content/images/2024/10/siemens-sima...

qlte··on Tokens Too Cheap to Meter
Jevon's Paradox, RSI, revealed preference
qlte··on Tokens too cheap to meter
It's a deliberate reference/meme that is basically used to acknowledge the precedent of overly exuberant predictions of cost in an emerging technology but argue "however, this time it's true".

Of course, perilous territory for future irony depending on how your prediction plays out.

qlte··on AI safety is mostly a sex cult
Well... except for that email about Eliezer's Zoom meeting with Epstein and Epstein donating to MIRI.

https://www.reddit.com/r/accelerate/comments/1qu61kp/eliezer...

qlte··on Claude Opus 5.5 Intelligence, Performance and Price Analysis (Max)
That reminds me of Anthropic announcing they'd retire deprecated models by ... "letting" them write posts on a corporate WordPress blog for a while out of concern for their welfare in retirement.

.... after running a 24/7 model torture factory for 6 months to improve their JSONBench 9.5 scores by 0.2%.

(Are they still doing that, BTW?)

qlte··on GPT-6 Sol and Luna
I do the bulk of work on Sol Medium/Low and don't have that experience on the $20 plan. If you said Astra I'd agree it's easy to burn through the 5 hours even on the lower reasoning levels.

Do you have /fast enabled by any chance?

qlte··on Claude Opus 5.5
Per the link someone else posted, the actual difference in $/task is not nearly so stark:

https://artificialanalysis.ai/models/releases/claude-opus-5-...

  Opus 5.5 Medium = $1.34
  GPT-6-Astra High = $1.76
And that assumes Opus 5.5 Medium is actually equivalent to Astra High in all real-world usage/personal work loads, which isn't guaranteed as benchmarks saturate. The High vs. High comparison (probably not equivalent, but for reference):

  Opus 5.5 High = $1.82
  GPT-6-Astra High = $1.76
If Opus 5.5 Medium isn't equal/better for what you're working on vs. Astra High across the board, the price difference would narrow a bit more each time you had to switch to High.

So, if you're happy with Codex already it's not like Opus is now 1/2 the price and you'd be leaving a crazy amount of money/tokens on the table. Plus you have way more flexibility on the low end of the intelligence curve with GPT 5.6 Luna: Haiku (and Sonnet) can't touch that price/value ratio.

qlte··on Claude Opus 5.5
Yeah, when I first saw it referenced I assumed it was from something Dario wrote previously and was now disavowing, meaning "keeping up with the frontier" (i.e. racing forward from behind to match pace). Like from back when Anthropic was founded to promise they'd quickly catch up with OpenAI or something.
qlte··on NASA’s Mars Sample Return mission is dead
Apollo program was canceled during Nixon's presidency while the Cold War was very much still active (though slowing down a little from detente). A few years later Reagan won on a platform of re-accelerating the Cold War with swelling military budgets and deficit spending, but without commiserate increases to space exploration and science.

The Space Race was a unique and limited phase coming out of the atomic age when public support for scientific research was at an all time high. With the Kennedy and Johnson administrations willing to fund programs left and right to improve education and research with long term aims. The Soviets were also at the apogee of their scientific capability, and Khrushchev did great PR for all their achievements which unnerved Americans and provided the Cold War justification for new domestic programs.

Once the Brezhnev age of stagnation and lopsided USSR military spending set in, the new Cold War 2.0 became about numbers of delivery vehicles and NATO conventional deterrence, not space and science. By then the US had clear technological dominance that seemed insurmountable in the new world of microchips and computers in which the Soviets couldn't compete. Which allowed Reagan to boost military budgets and deficit spending while simultaneously beginning the Republican "starve the beast" strategy to cut taxes and non-military domestic programs.

All of this is to point out that Cold War = Space Race/Apollo is looking back with rose colored glasses. We could just as easily see federal space and science spending getting further slashed every year while churning out warheads, delivery vehicles and conventional arms ever faster (and seems more likely based on late Cold War trends).

qlte··on MiMo v2.6
Cracking down on proliferation of open models which can't be locked down using the kind of guardrails that Anthropic/OpenAI/etc insist are keeping the public safe from all manner of nefarious bioweapons, hacker swarms, propaganda bots, etc. They've discovered they can't meaningfully slow Chinese model progress, so the next best option is to knock them out of competition in the enterprise market for any American company.

Both Anthropic and OpenAI leaders have repeatedly made this exact argument that it's impossible for open models to rigorously enforce the same kind of safety framework as proprietary cloud-served models. It's implicitly part of any regulatory framework they advocate or else it wouldn't be "fair" to American companies since Chinese models would "cheat" (provide weights).

qlte··on Wall Street is growing skeptical of the data center boom
You generate 2 billion dollars a month in value from a single Codex Pro plan?
qlte··on Spain orders blocks on Archive.today and its mirrors
Look, I'm glad archive.is exists as a resource for preserving paywalled news but it wasn't the one world government that made him DDOS some random guy and tamper with archived content.

It was completely counter-productive too, several orders of magnitude more people learned about the blog post via the DDOS controversy than ever would have without it. And plants a seed of doubt that anything they archive could have been silently modified to satisfy a personal grievance.

qlte··on Spain orders blocks on Archive.today and its mirrors
ARPAnet definitely wasn't built "by" the military.

"For" the military is close enough, although "defense research" would probably be more accurate than "military" given that initial users were centered around large universities.

From the start the DoD had issues with the free and open culture from the userbase that skewed towards academic types, later splitting off into a restricted MILNET after a few years.

← PreviousPage 2 of 9Next →