HNHacker News
TopNewBestAskShowJobs

toraway

567 karma · joined December 7, 2024

submissionscomments
toraway··on Google's Antigravity Bait and Switch
Possibly a bug, but the change in usage quotas on my AI Pro plan going from Gemini CLI to Antigravity CLI was a massive drop. I was kicked off after around 30 minutes using the smallest model available (3.5 Flash).

If they offered 3 Flash (or 3.1 Flash Lite too but might be hoping for too much) with comparable usage limits then the transition to Antigravity CLI wouldn't have bothered me much at all.

toraway··on An OpenAI model has disproved a central conjecture in discrete geometry
Thank you for sharing, that was one of the most insightful long form pieces I've read in a long time! And the writing was enjoyable to read even as a math layperson.

I was going to say you should submit it but I saw you did a few days ago but it only got a few votes... If Dang sees this IMO it would be extremely deserving of the second chance pool as I wouldn't be surprised to see easily jump to the front page with a different roll of the dice.

toraway··on Plex's 200% Lifetime Pass price hike tries forcing users to another subscription
Jellyfish has improved a bunch in the last few years, the front end is a lot more polished. I finally moved on from Plex a year ago after some kind of upsell nag on a basic feature after already paying for a plan.

Although realtime detection of changes on the file system is still a little flaky for me (possibly how I’m running it).

toraway··on OpenAI Adopts Google's SynthID Watermark for AI Images with Verification Tool
That's a lot of hyperbole, there's no cause/effect relationship I can think of here that could realistically produce your slippery slope.

Google or anyone else could start adding those unique tracking watermarks you're concerned about any time they want, regardless of whether they use this AI detection watermark, that to be clear can not track you in any way.

toraway··on OpenAI Adopts Google's SynthID Watermark for AI Images with Verification Tool
FWIW there are a few people in the issues saying that the tool is giving false negatives and the output image gets flagged by the actual Gemini API as having SynthID. Most recently 3 weeks ago without a response.
toraway··on Gemini 3.5 Flash
That would be Flash Lite now, and I'm also interested in the cheaper end of things so kinda disappointed they didn't release 3.5 Flash Lite at the same time...
toraway··on The main thing about P2P meth is that there's so much of it (2021)

  > The result is that the illegal market dwarfs the legal market. The legal suppliers simply can't compete with efficient and untaxed illegal or grey market sellers.
Source for this? I mean consumer buying habits not illegal grow ops in general (which are selling a lot of their product out-of-state).

Even with California taxes an 1/8th of flower is often less than half what I’d typically pay back in 2008 or so, even without adjusting for inflation.

Also, I suspect middle class buying preference in the last decade or so has heavily shifted to gummies/edibles and vape carts which are much sketchier in the black market vs relatively interchangeable flower.

The idea of smoking a literal bowl to get high wouldn’t even be in the first like 5 methods among the people I know. It’s not super appealing in your 30s-40s while living in apartments/not wanting to reek of weed in the office. So buying off the black market just isn’t attractive even if possible cheaper.

toraway··on Frontier AI has broken the open CTF format
Okay, but none of that is actually responsive to what the article is discussing, which is competitive CTFs. There's not a single criticism of using AI for actual security research in anything they wrote and they mention being a heavy user of GPT-5.5 and GPT-5.5 Pro so belittling the author's experience to defend LLMs wasn't actually necessary.
toraway··on Frontier AI has broken the open CTF format
Just parachuting in to reflexively throw the "Luddite" label at someone lamenting the decline of a niche community they've enjoyed participating in and contributing to is certainly ... a choice.

Within the framework of your analogy, it's like responding to someone active in DIY maker groups suddenly dealing with an influx of influencers in meetups showing off Chinese junk from Etsy to post on Tiktok, and accusing them of being a Luddite blinded by their zealous hatred of mass production -- both strangely abrasive and also fairly nonsensical except as a "mass production supporter" social signifier.

Not to mention, in the article they specifically describe themselves as a heavy user of frontier models for security research ever since the release of Opus 4.5, calling them "useful within the field". In fact I don't see any actual criticism of AI/LLMs anywhere whether for security research, programming or anything else, except for making competitive CTFs no longer viable.

What does it take to avoid the "Luddite" brand? Using AI themselves and praising AI as useful (to the point of having a lopsided advantage over humans) isn't enough? Do they also need to say "I haven't written a line of code in 6 months/it's easily a 100x multiplier for my job" every time they mention it too?

toraway··on I believe there are entire companies right now under AI psychosis

  > Vietnam - unnecessary war, but we won the peace
I’m struggling to understand what this spin is even supposed to mean?

  > Afghanistan - we tried our best and made some mistakes along the way. Eventually got Bin Laden though. *Too bad the rest of the world didn’t help.* Now we’re seeing a massive regression in women’s rights there.
Why are you lying about this?

  > At its peak between 2010 and 2012, ISAF had 400 military bases throughout Afghanistan (compared to 300 for the ANSF) and roughly 130,000 troops.[7] Forty-two countries contributed troops to ISAF, including all 30 members of NATO.
https://en.wikipedia.org/wiki/International_Security_Assista...
toraway··on We are retiring our bug bounty program
From the excerpt it sounds like the author is just describing one specific archetype from within a list of others included in the book and doesn't make any claims about it being a uniquely common type within every org, or the most common type of bad engineer in general.

In fact it gives the opposite impression by specifying "at least one", which implies the category is supposed to be distributed widely enough to be recognizable in an org of sufficient size, but not dominating the ranks of software developers in droves. That seems more like a strawman you're arguing against.

toraway··on We are retiring our bug bounty program
The disclosure about being a honeypot is in the CONTRIBUTING.md:

  Warning

  Heads up: This is a research project — bounties listed here are symbolic and   part of an academic study on open-source contribution patterns. PRs are reviewed for research purposes only and will not be merged into production. If you're looking for paid bounty work, this is not the right repo.
Which makes it slightly surprising those bots with system prompts to find "high value bug bounty targets" or similar aren't deterred by that when they pull the repo.

I guess a sort of task blindness where once they've gone as far as to git clone they've already switched gears from searching Github for qualifying bounties into a find bug->fix bug->open slop PR mindset to close the loop and end the turn? By that point an incidental warning they ingest in passing while looking for the Solana contract vulnerability they already committed to working on in a comment might not even register as relevant to the current task at hand.

toraway··on Ontario auditors find doctors' AI note takers routinely blow basic facts
I recently left my mom a voicemail saying happy Mother’s Day with normal human boilerplate of sorry I missed you, feel free to give me a call back tonight or we can talk tomorrow, either is fine by me whatever works best for you, hope we can talk soon, love you, bye.

She called me back later that night and we chatted for bit and then she paused and sort of uncertainly was like “So… was there something you were needing to tell me?” And I was completely baffled and was like “Uhhhh I don’t think so…?”

She then explained the notification she got about my call and apparently the LLM summary of my voicemail converted a message consisting of 75% well-meaning but insignificant interpersonal human filler (like most voicemails) into this stilted, overly formal business-y speak with a somewhat ominous tone. Assigning way too much significance to each of the individual statements in the message about wanting to talk (to say happy Mother’s Day), inquiring about her availability ASAP (to say happy Mother’s Day) etc. Plus grossly exaggerating the information density of the call making it sound like I left this rambling, detailed message about needing to tell her something that was left completely vague, but possibly important and also time critical.

Added up it made her a little worried when she read it and made me a bit pissed that was the end result of my wishing her well. Because apparently everything needs a half baked LLM summary crammed into it now.

toraway··on The Emacsification of Software
Also, as someone who has developed an ever growing suite of bespoke tools for my personal workflows using Codex/Gemini CLI over the last year, something I don’t see mentioned as often is the “mental overhead” of self-designed apps.

Even if the coding process itself is “effortless” and the agent just churns away to implement whatever I ask for on a dime, it can become exhausting thinking through all my needs/wants, tradeoffs, API shape etc. Despite not needing to write a line of code myself or read more than excerpts in the chat it can turn into a slog after the honeymoon period passes and it starts to feel like unpaid work.

I’ve had moments where I’m relieved to discover a popular open source tool that works out-of-the-box as an alternative to my own so I can offload that organizational overhead and decision fatigue to someone else. While benefiting from all their features/enhancements I didn’t have to design or maintain myself over time.

As an example, I had been building a TUI/web app to download and organize ebooks from various sources like Project Gutenberg or Anna’s Archive with a central meta search, and manage my personal library. It solved the immediate problem at the time but I kept needing to add missing features, plug holes in the various search integrations, UI refinements, etc and it never quite worked exactly as I wanted so kept having to work on it and became less and less fun as time went on.

Then I discovered Calibre Web Automated + Shelfmark on GitHub that did 99% of what I needed plus a lot more and overall had a level of polish and reliability my tool never reached. Now I just pull a Docker container every so often for updates and made a few tweaks to syncing but overall spend vastly more time on actually reading/organizing/growing my library vs. increasingly tedious vibe coding sessions and it feels so much more enjoyable.

I still have plenty of self-designed tools and continue making new ones but now tend to reach for an existing, off-the-shelf option first whenever possible for anything more complex than a one-off script. That way I can benefit from a community collectively contributing to improve and maintain the project over time without needing to become an unpaid Product Manager, Lead Designer, Senior Developer and QA Manager for everything I use.

I hope the current period of exuberance around LLM development doesn’t lead to everyone becoming stuck in individual silos duplicating work that in the past could have been directed to an OSS project where that time investment could be shared with everyone else and benefit from way more eyes catching bugs and smoothing off rough edges.

toraway··on Googlebook
Battery management tends to be best-in-class on Chromebooks, it's far from certain that you'll get anything nearly as good after installing 3rd party Linux on it. That's my #1 reason for not even considering it (despite having installed Linux on many different Chromebooks years ago when they were new and ChromeOS was still literally just a browser).

My <$200 ARM Chromebook got around 12-14 hours battery life new (though as with my M1 Macbook has degraded to probably 70% capacity after 2-3 years). It draws essentially no power in standby (ChromeOS will enter an ultra low-power hibernate-like state seamlessly after a while). I've opened it up a month after last using it and it turns on in <10 seconds having lost a couple percent.

Updates are seamless and add like 5 seconds to boot time when they apply during a restart (thanks to ChromeOS A/B update model there's no loading spinner or anything, you reboot and it's done. Update countdown extortion a la Windows isn't a thing either, you can stay on an non-updated Chromebook for months without a reboot and the most you'll see is the same passive "Click Restart to Update" button in the notification area.

I use the built-in Linux VM all the time, it runs GUI apps like VS Code without any issues, and my ARM Chromebook runs all sorts of regular Arm64 Debian builds for GUI or terminal out of the box. I turned off the Play Store for Android Support, in the past when Linux support was weaker and web APIs in general weren't as capable I needed it for some specific apps but don't really have a need at this point.

The security model on ChromeOS means that untrusted scripts/installers running in the Linux VM are completely isolated from anything on the browser so you (or your proverbial Grandma) don't have to worry about credential stealers/ransomware/malware. You can copy files between the ChromeOS filesystem and the LinuxVM filesystem but a process running in Linux can't cross that boundary and are confined to the sandbox.

Plus, very much unlike my Macbook, I can actually install an app from Github or compile myself without 7 clicks and three different dialogs each time (as is the case with Apple's aggravating security hassleware on MacOS Sequoia). Proving you can have a heavily locked down, secure model without actively making the user experience as miserable as possible (to "encourage" use of the built-in app store).

It's easily the least intrusive OS experience of all the major OSes, and completely gets out of the way without drawing attention to itself. And sure, Google is an advertising company, I get it, but my Macbook advertises iCloud products and Apple TV shows to me more than anything on my Chromebook.

With the 10 year Chromebook support policy, I've got a crazy amount of life out of all of my Chromebooks. It really is liberating having an OS that de-emphasizes its own existence in a world where I have to fight ever changing MacOS deprecation and security restrictions and Windows bloatware being thrown over the fence in every other update.

toraway··on Mythos Finds a Curl Vulnerability
Not exactly "burying the lede" since Daniel already posted an update about it months ago [1] with extensive discussion in numerous of articles [2] including on this site [3].

[1] https://lists.haxx.se/pipermail/daniel/2025-September/000127...

[2] https://www.theregister.com/software/2025/10/02/curl-project...

[3] https://news.ycombinator.com/item?id=45449348

toraway··on Mythos Finds a Curl Vulnerability
Also, the people at Mozilla who helped achieve a highly visible collaboration with the hottest AI company in the zeitgeist that included a lot of expensive data center time to harden their flagship product are definitely going to be happy/excited/proud about pulling it off successfully.

There's a lot of kneejerk "so you're accusing Mozilla of a conspiracy to boost Anthropic?" which is an overly simplistic lens. Particularly when it involves groups of individual humans with different motivations and emotional investment in their own contributions to the collaboration.

toraway··on Mythos Finds a Curl Vulnerability
I've seen this suggested a few times in this thread but it seems like it's exactly backwards.

Wouldn't that make it a better to distinguish whether Mythos is uniquely super powerful vs an incremental improvement from Opus etc that are routinely used as the basis for bug reports/fixes in cURL?

If Mythos found a hundred new show stopper bugs then it would have meant Opus missed them and therefore closer to a "step change". Otherwise it implies the difference in capability isn't nearly that stark. Mythos finding 100 low-hanging bugs in a less scrutinized/hardened project on the wouldn't be as useful signal to answer that.

toraway··on Ask HN: We just had an actual UUID v4 collision...
You could do that, but now you're like 90% of the way to maintaining a monotonically increasing number you that could just use as a unique ID instead without any randomness required (and without the additional 128 bits for collision protection via the appended UUID).

So your ID would take like 64 bits for the time unique to the nanosecond plus 128 bits for the UUIDv4 = 192 bits which is a pretty beefy sized ID.

(I know you said just append a second count but you will want a predictable/fixed size for your data structure in pretty much any use case so need to decide the upper bound and precision ahead of time)

Especially when the alternative is a 128 bit UUIDv4 that's guaranteed unique with proper usage of high quality RNG or a 128 bit UUIDv7 if you have a clock (that's needed for your method anyway) that will be much more forgiving of a flaky source of randomness and more sortable than your monotonic-ish ID for 1/3 fewer bits.

Basically, stapling anything onto a UUID is a waste of space if you don't trust it, so might as well drop it completely and use a significantly smaller source of randomness at that point.

toraway··on Ask HN: We just had an actual UUID v4 collision...
Shouldn't your test follow the pattern of how rng() is actually being used in the uuid.ts code internally?

Your test is more-or-less contrived to fail given the tradeoff to avoid repeated memory allocations but that doesn't say much about the actual usage in uuid generation since it's not exported for general purpose use.

Presumably they had some hot path somewhere where rng() is called in a loop and this optimization made sense with awareness that it could be misused as in your example breaking the contract ensuring randomness, which (hopefully) they're not actually doing anywhere.

Unless I'm missing something replacing the package over this with a less vetted implementation seems excessive and possibly even counterproductive.

toraway··on Making LLM Training Faster with Unsloth and NVIDIA
The problem with AI written articles is still feeling uncertain whether there's actually any utility after reading 2000 words as you realize that it's been 90% filler so far but think maybe it will lead somewhere soon? But it doesn't and you wasted ten minutes reading glorified blog spam that was micro targeted at whatever niche you were researching.

After a while you pick up on the warning signs and just bail early without any guilt about false positives. It's really the only sustainable strategy in a world where it takes 5 seconds to absorb 5 minutes of your attention span.

toraway··on Canvas is down as ShinyHunters threatens to leak schools’ data
Exactly. This is the "Declare fentanyl a WMD" of solutions to ransomware. Sounds kinda badass as long as you don't spend too long thinking about it but has no practical relevance to actual enforcement challenges.

It's a familiar example of the perennial "[THING] could be solved overnight if [PERSON_OR_GROUP] would just start taking [THING] seriously" trope.

toraway··on Vibe coding and agentic engineering are getting closer than I'd like
The asbestos hypothetical is a bit different than the "bubble popping" economic crisis scenario though. In this world, AI would just continue being adopted and shoved into every nook and cranny into which it can be made to fit, with valuations only getting bigger and bigger.

The damage would come much later, well beyond the point where it could be simply pulled out and replaced without spending massive amounts of money and would also basically necessitate training an entire new generation of engineers.

Then the AI giants would start appearing vulnerable like cigarette companies in the 90s while an AI Superfund and interstate class action are being planned but Sam Altman would already be a centitrillionaire at that point so it would be someone else's problem.

toraway··on Vibe coding and agentic engineering are getting closer than I'd like
Or, it could be like asbestos and the immediate benefits are just too appealing to listen to arguments of skeptical naysayers about some vaguely defined problems that are decades away, if they even happen.

I use AI tools daily (because they feel like they're helping me) but it's not exactly hard to imagine scenarios where an explosion of slop piling up plus harm to learning by outsourcing all thinking results in systemic damage that actually slows the pace of technological progress given enough time.

History of new technologies tend to average into a positive trend over a long enough time scale but that doesn't mean there aren't individual ups and downs. Including WTF moments looking back at what now seems like baffling decision-making with benefit of hindsight.

toraway··on Today I've made the difficult decision to reduce the size of Coinbase by ~14%
That reminds me of a chart I saw posted in HN comments recently that someone created tracking bullet points in Claude Code release notes per day that was cited as "proof of a step change" in AI development over the last year. It showed like a dozen or so on average that jumped to to like over 50 one month and stayed around that number.

(Not the exact same chart but similar idea, I guess it's sort of a meme: https://imgur.com/a/YrNGYOR)

So I looked at the most recent CC release notes on Github and the majority look like this:

  Fixed /clear not resetting the terminal tab title after a conversation
  Fixed session title chip from /rename disappearing while a permission or other dialog is active
  Fixed agent panel below the prompt being hidden when subagents are running (regression in 2.1.122)
  Fixed external-editor handoff (Ctrl+G) blanking the conversation history above the prompt
  Fixed /context dumping its rendered ASCII visualization grid into the conversation, wasting ~1.6k tokens per call
  Fixed OAuth refresh race after wake-from-sleep that could log out all running sessions
  Fixed 1-hour prompt cache TTL being silently downgraded to 5 minutes
  Fixed cache-miss warning appearing spuriously after /clear or compaction when changing /effort or /model
I'd be extremely interested to know what percentage of these were just fixing last week's Claude Code written PR that no human ever set eyes on.

But hey, all that churn looks great on charts being circulated on social media as free advertising for their flagship product (and consequently the company's valuation) so never mind, LGTM!

toraway··on Today I've made the difficult decision to reduce the size of Coinbase by ~14%
Despite my own personal preference for static sites, marketing using a CMS under their own control to make content updates seems vastly more reasonable than vibe coding open ended PRs as a codebase they don't understand gradually grows in complexity over time.

They could even use one of many headless CMSs combined with a static generator. Claude Code in the hands of non-technical users deploying to prod regularly seems like one of the worst possible ways to do it (except for the "cool" value telling people about it).

At my company the internal devs don't even have access to wherever the company site is hosted, it's a WordPress CMS and marketing can make updates safely with a couple clicks and zero day-to-day development oversight required. IT just helps keep the box updated but otherwise it's entirely their own thing.

toraway··on Accelerating Gemma 4: faster inference with multi-token prediction drafters
Gemini CLI has improved a lot in the past 6 months or so. Back when I used in the 2.5 Pro era it would get stuck in loops literally like 1/8 conversations and I eventually just gave up despite having access included in my AI Pro plan.

But last month I picked it up again and it has crushed everything I've thrown at it. As Codex limits tighten on the Plus plan it's been my main fallback and doesn't even feel like a downgrade when I switch over. Haven't hit a single loop so far using it nearly every day for several weeks so that problem seems solved finally, thank god.

I've been using it in the auto router mode and haven't felt the need to manually lock in the bigger model yet. It's incredibly snappy which I realized I really appreciate vs. waiting around endlessly for minutes each turn, but I've read other people's experiences needing to manually select the Pro model so YMMV.

toraway··on GameStop makes $55.5B takeover offer for eBay
It may not be bad for the buyers, lenders, or the previous owners who all profit. But even then it could still be bad for regular people/society at large if it incentivizes anti-consumer practices by financial necessity when an otherwise healthy business suddenly has billions of dollars of debt it has to pay off ASAP. In which case it sort of looks like a simple transfer of wealth from existing customers to the organizers of the leveraged buyout with no broader societal value provided like new jobs created, R&D, etc.
toraway··on New statue in London, attributed to Banksy, of a suited man, blinded by a flag
Considering that line is supposed to be written by a young child in-context (who couldn't actually "remember" anything more than a decade earlier, I'm pretty confident the intent was not to reference the actual recent history of urban deforestation in Detroit. So this attempt to fact-check the art doesn't actually work at all here.

Off the top of my head, I'd guess the message is closer to an observation about being disconnected from history in the modern world leading to vaguely defined feelings of angst and alienation.

toraway··on Sierra Raises $950M at $15B Valuation

  > Agents (if implemented well) are an order of magnitude more effective at resolving issues compared to a call centre worker who is reading off a script and churn within 9 months
For this to be true, the agent needs to actually be given the means to solve the problem, otherwise an "agent" is just a glorified help page that wastes your time.

But it seems like companies don't want to do this part, possibly because of fears that someone will trick the agent into giving them a refund or something. Or because the actual goal is to optimize for fewer costly refunds/cancellations/policy exceptions etc.

So for whatever reason, they stay stuck in that useless local maxima while simultaneously making traditional help increasingly difficult to get ahold of when needed for an overall net worse experience as a customer.

← PreviousPage 2 of 9Next →