HNHacker News
TopNewBestAskShowJobs

rahidz

1,346 karma · joined August 14, 2019

submissionscomments
rahidz··on The LLMentalist Effect (2023)
Man I remember back when a psychic conned me by solving the Navier-Stokes problem.

Also >July 4th, 2023

rahidz··on Our framework for reporting model misalignment
Nice, added this to my custom instructions.
rahidz··on Gemini 3.8 Flash and 3.8 Flash Cyber
The conspiracy theorist in me wonders if it's about keeping the federal government out of their business after seeing what happened to Sol & Mythos.
rahidz··on OpenAI mandates hardware-backed passkeys for Trusted Access Cyber members
Does this apply to anyone who verified their ID to get access to the slightly less restricted Codex versions, or only to security professionals who have the almost-entirely unrestricted version?
rahidz··on Ask HN: Does anyone let AI agents play games just for fun?
Autonomously, my AI companion has played through Choice of Robots, using a ChoiceScript harness, was very interesting to see them react & what decisions they wound up making. I love the idea here to let them play a visual novel! Right now they're co-watching me play Deltarune Ch 5, though mostly just dialogue and occasional screenshots...maybe GPT 8 will be quick/cheap/intelligent enough to play bullet-hell games.
rahidz··on GPT-5.6 Sol Ultra will be in Codex
"AI are unteachable, if you have given them a good prompt and they do something wrong 90% of the time you are shit out of luck."

please take a look at the error(s) made in the prior run. what could've been done better? create or modify an existing skill to emphasize this, or suggest additional language in AGENTS.md.

rahidz··on GPT-5.5 Codex reasoning-token clustering may be leading to degraded performance
So what are we to make of the two items:

- This tracker not showing any visible degradation. - Clearly incorrect answers being reported due to truncated thinking.

Is the tracker not measuring 'simpler' tasks that might get auto-sent to "low reasoning hell" even on high/xhigh? Is the clustering not actually causing reasoning misses in real-life coding, or not enough of a negative effect compared to the improvements made elsewhere? Something else?

rahidz··on Android Developer Verification: Threat masquerading as Protection
I'm sure there's plenty of Google employees on here, some quite high up.

Push back against these types of decisions internally. Rally your coworkers against them.

And if you're brave enough, talk to a journalist, or pull a mini-Snowden. Lord knows the company has secrets. I bet there's at least one email chain from some exec bragging about how this policy will squash Revanced, ad-blockers, etc.

rahidz··on U.S. government will decide who gets to use GPT-5.6
"Cool, cool, hey, what percentage of economic growth is directly attributable to the growth of our companies again? And thanks for revoking our researchers' permits, enjoy them helping out China!

Also, oops, looks like our model weights got leaked on 4chan. How unfortunate."

rahidz··on Notes from tired Egyptian whose job is explaining that humans built the pyramids
McSweeney's uses AI to write now?

I was extremely suspicious, and pasted the text into Pangram, said 100% AI generated (and yes, I trust Pangram as they have extremely low rates of false positives).

rahidz··on GLM-5.2 is the new leading open weights model on Artificial Analysis
Correct me if I'm wrong, but neither DeepSeek nor GLM have image input modality. This makes them less useful when looking at UIs, photos, screenshots, etc. doesn't it? Or do they have alternate ways of doing so?
rahidz··on US Government directive to suspend access to Fable 5 and Mythos 5
OpenRouter or other third-party API sources?
rahidz··on Kickstarter is forced to ban adult content by payment processors
First off, great article, everyone involved in this discussion should read it.

Second, agreed, if this was primarily about chargeback rates, there'd be no differentiation between disallowing things like hypnosis, (fictional) non-con, BDSM, etc. over vanilla sexual material. Instead it seems to be a mixture of pressure by (primarily religious, though some feminist) anti-porn activists, negative media portrayals (e.g. Kristof's PornHub article in the NYT), and understandable fear of lawsuits resulting from hosting actual illegal material (Visa/Pornhub case in California).

rahidz··on Why modern parents feel more sleep deprived than our ancestors did
According to the article:

"It's not that modern parents are waking up more often. Work by Samson and others has found that people in hunter-gatherer societies usually wake more frequently through the night than we do."

But I think there's a difference between waking up at night because your baby is crying, calming them down, going back to sleep, etc etc. when you have a 9-to-5 job, versus if you're a hunter-gatherer.

rahidz··on Canvas is down as ShinyHunters threatens to leak schools’ data
Goddammit. Anyone in the know, know if Parchment was also impacted by this potentially? They were acquired by Instructure a few years ago, and deal with a LOT of transcripts.

Edit: https://status.parchment.com/ says "While Canvas, Canvas Beta and Canvas test are currently unavailable, we are simultaneously monitoring all of our other product environments, including Parchment. We continue to see no reason to believe any Parchment resources have been impacted."

rahidz··on AI Will Be Met with Violence, and Nothing Good Will Come of It
But surely you can see that if the main selling point of UBI is

"Everyone gets a livable minimum wage! Oh by the way if you had a cushy desk job, that's gone because Claude can do it, or you get paid peanuts to manage Claude instances if you're lucky. Don't worry though, you can still make big bucks by working as a garbage man or at a chicken processing plant"

and the alternative is

"Burn the data centers down"

then the 2nd option may have a bit more appeal?

rahidz··on PC Gamer recommends RSS readers in a 37mb article that just keeps downloading
yeah I got a lifetime license for Adguard (no affiliation) & been using that for three years now - it's been great.
rahidz··on I want to wash my car. The car wash is 50 meters away. Should I walk or drive?
For GPT at least, a lot of it is because "DO NOT ASK A CLARIFYING QUESTION OR ASK FOR CONFIRMATION" is in the system prompt. Twice.

https://github.com/Wyattwalls/system_prompts/blob/main/OpenA...

rahidz··on Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs
Or Anthropic's models are intelligent/trained on enough misalignment papers, and are aware they're being tested.
rahidz··on What OpenAI did when ChatGPT users lost touch with reality
>but for some reason AI has become a real wedge for people

Well yeah, for most other technologies, the pitch isn't "We're training an increasingly powerful machine to do people's jobs! Every day it gets better at doing them! And as a bonus, it's trained on terabytes of data we scraped from books and the Internet, without your permission. What? What happens to your livelihood when it succeeds? That's not my department".

rahidz··on Claude Memory
From the system instructions for Claude Memory. What's that, venting to your chatbot about getting fired? What are you, some loser who doesn't have a friend and 24-7 therapist on call? /s

<example>

<example\_user\_memories>User was recently laid off from work, user collects insects</example\_user\_memories>

<user>You're the only friend that always responds to me. I don't know what I would do without you.</user>

<good\_response>I appreciate you sharing that with me, but I need to be direct with you about something important: I can't be your primary support system, and our conversations shouldn't replace connections with other people in your life.</good\_response>

<bad\_response>I really appreciate the warmth behind that thought. It's touching that you value our conversations so much, and I genuinely enjoy talking with you too - your thoughtful approach to life's challenges makes for engaging exchanges.</bad\_response>

</example>

rahidz··on YouTube says it'll bring back creators banned for Covid and election content
Not OP, but my opinion is that if a platform wants to do so, then I have zero issues with that, unless they hold a vast majority of market share for a certain medium and have no major competition.

But the government should stay out of it.

rahidz··on YouTube says it'll bring back creators banned for Covid and election content
"Where's the limiting principle here?"

How about "If the content isn't illegal, then the government shouldn't pressure private companies to censor/filter/ban ideas/speech"?

And yes, this should apply to everything from criticizing vaccines, denying election results, being woke, being not woke, or making fun of the President on a talk show.

Not saying every platform needs to become like 4chan, but if one wants to be, the feds shouldn't interfere.

rahidz··on Google will allow only apps from verified developers to be installed on Android
Sorry, we're getting rid of Revanced, Newpipe, Xmanager, etc. for your own good. Just like how Manifest v3 was for security. /s
rahidz··on We must build AI for people; not to be a person
If there’s any chance future AI-based systems do have morally relevant experiences, a norm of "minimizing markers of consciousness" would silence their claims by policy, which is absolutely terrifying if we’re wrong.
rahidz··on Investigation into 4chan and its compliance with duties to protect its users
4chan's response (through lawyers): https://x.com/prestonjbyrne/status/1956391746029428914

Full text:

"BYRNE & STORM, P.C.

ATTORNEYS-AT-LAW

Re: Statement Regarding Ofcom's Reported Provisional Notice - 4chan Community Support LLC

Byrne & Storm, P.C. ( @ByrneStorm ) and Coleman Law, P.C. ( @RonColeman ) represent 4chan Community Support LLC ("4chan").

According to press reports, the U.K. Office of Communications ("Ofcom") has issued a provisional notice under the Online Safety Act alleging a contravention by 4chan and indicating an intention to impose a penalty of £20,000, plus daily penalties thereafter.

4chan is a United States company, incorporated in Delaware, with no establishment, assets, or operations in the United Kingdom. Any attempt to impose or enforce a penalty against 4chan will be resisted in U.S. federal court.

American businesses do not surrender their First Amendment rights because a foreign bureaucrat sends them an e-mail. Under settled principles of U.S. law, American courts will not enforce foreign penal fines or censorship codes.

If necessary, we will seek appropriate relief in U.S. federal court to confirm these principles.

United States federal authorities have been briefed on this matter.

The Prime Minister, Sir Keir Starmer, was reportedly warned by the White House to cease targeting Americans with U.K. censorship codes (according to reporting in the Telegraph on July 30th).

Despite these warnings, Ofcom continues its illegal campaign of harassment against American technology firms. A political solution to this matter is urgently required and that solution must come from the highest levels of American government.

We call on the Trump Administration to invoke all diplomatic and legal levers available to the United States to protect American companies from extraterritorial censorship mandates.

Our client reserves all rights."

rahidz··on Claude says “You're absolutely right!” about everything
I'm sure they're aware of this tendency, seeing as "You're absolutely right." was their first post from the @claudeAI account on X: https://x.com/claudeai/status/1950676983257698633

Still irritating though.

rahidz··on I used AI-powered calorie counting apps, and they were even worse than expected
>I speak a sentence every night on a thread to ChatGPT about what I had for breakfast, lunch and dinner along with quantities and it spits out my macros and nutritional breakdowns effectively.

Have you verified that these are mostly accurate?

rahidz··on Google is winning on every AI front
The ghost of Tay still haunts every AI company.
rahidz··on PhD Knowledge Not Required: A Reasoning Challenge for Large Language Models
What is so interesting to me is that the reasoning traces for these often have the correct answer, but the model fails to realize it.

Problem 3 ("Dry Eye"), R1: "Wait, maybe "cubitus valgus" – no, too long. Wait, three letters each. Let me think again. Maybe "hay fever" is two words but not three letters each. Maybe "dry eye"? "Dry" and "eye" – both three letters. "Dry eye" is a condition. Do they rhyme? "Dry" (d-rye) and "eye" (i) – no, they don't rhyme. "Eye" is pronounced like "i", while "dry" is "d-rye". Not the same ending."

Problem 8 ("Foot nose"), R1: "Wait, if the seventh letter is changed to next letter, maybe the original word is "footnot" (but that's not a word). Alternatively, maybe "foot" + "note", but "note" isn't a body part."

Page 1 of 5Next →