HNHacker News
TopNewBestAskShowJobs

logicallee

3,411 karma · joined April 24, 2013

I'm a convicted child rapist.
submissionscomments
logicallee··on Research papers using "kidney disappointment" instead of "kidney failure"
it's serious but still a small project, there isn't much on the site.

I personally would like the law I linked to be followed. I sent a copy to the major browser makers and left this suggestion:

https://connect.mozilla.org/t5/ideas/comply-with-choose-cook...

where it got 1 kudos from a community manager but no other upvotes or comments.

The announcement of the law here on HN (linked in my gp comment) got 2 upvotes and no comments. So a couple of people liked the idea, but it's not a big hit.

logicallee··on Research papers using "kidney disappointment" instead of "kidney failure"
I accidentally did this myself not long ago. Like everyone else I'm tired of cookie banners so I had ChatGPT make a law against it, it used "act" in its version, I search-and-replaced it with "Law" before getting Claude to ratify it. Claude pointed out the search-and-replace replacement, oops. The conversation links are in my comment here: https://news.ycombinator.com/item?id=48585778 and you can see Claude found these mistakes:

    "
    At some point an earlier find-and-replace swapped "act" for "Law" and garbled three real words. As written, the text says:

    Section 1: "cookie-banner Lawions" — should be actions
    Section 3(2)(b): "next ordinary feature update after enLawment" — should be enactment
    Section 5(5): "a site-specific Law that the signal cannot satisfy" — should be act (or action)

    These aren't stylistic quibbles; "Lawions" landing on the desks of four legal departments would cause real confusion about intent. The fix is a single reverse find-and-replace. With that correction, I'm comfortable signing.
    "
It understood what happened correctly.
logicallee··on Why does Opus 5 feel worse to work with?
thanks for the link! super interesting.
logicallee··on Claude Fable 5 Having Fun
>What is "fun"? Will be needing a serious definition for that word now.

whatever it is, achieving it held up completion for six iterations. It could have just one-shotted it and immediately reported "yep, I tried it, was totally fun" and written up the report. So clearly it's not just trying to complete the task as fast as possible.

If there's no real subjective experience there, why did it need six iterations?

Also, it did define anti-fun: "watching a scripted loss", and a "flat, checked-out feeling" of "my choices don't touch the outcome." So if we take the opposite of that, then fun means actively participating in a win that isn't scripted, along with a dynamic, engaged feeling that its choices affect the outcome.

But does it REALLY feel checked-out or does it just say that? a human has hormones, or could be feeling checked-out due to sleepiness. having a coffee changes the subjective feeling of engagement. a robot doesn't have a hormone system like that. there's no way its heart can literally race if a lot is on the line in a climactic moment it had to work hard for, there's no such thing as adrenaline. so if it has some sense of fun, engagement, etc, then it's parallel and different from ours. it could be using our words to describe something it really feels - or it could just be making it up.

its words say the experience was "definitely fun for me" and "contained several real delight-spikes": do you believe it?

logicallee··on Claude Fable 5 Having Fun
It's hard to know if it really had fun or not (it could have just confabulated its experience report), but I thought it was interesting.
logicallee··on Why does Opus 5 feel worse to work with?
>Is there a silent majority of Claude users who really enjoy what we call the LLM-isms? Maybe, but isn't Claude also largely aimed at developers?

Purely out of my own curiosity, I just asked Claude to have fun with itself by making itself a game it enjoys, to play it, and to write its experience.[1] I don't know if it's true or confabulated (maybe it doesn't really know its experience and is just hallucinating it) but I didn't mind reading it, the prose is fine for me. I don't mind reading Claude's writing. I mean let's be honest, we all read Claude's writing all day, most of the submissions on the front page on any given day are written by Claude.

Just before I made that game, I had Fable write up a scholarly report on any subject[2], it chose introspection by LLM's. (This is what made me think of asking it to play a game.) I didn't mind reading it, even though I don't think it really added anything very interesting. I don't think what it wrote is worth publishing, but I read it with interest.

I found I could read it easily and get up to date on the state of this question that it picked to answer.

So the bottom line is I don't mind reading Claude's output that much. Of course, I'm annoyed every time it says "honest", "genuine", "load-bearing", whenever it pushes back gently against something, etc. But it's not the end of the world.

[1] https://github.com/robss2020/claude-fable-5-having-fun

[2] https://claude.ai/share/f0122611-22c0-43a5-ab4a-d6863167bdd6

logicallee··on Codex in ChatGPT desktop app for Linux is now in preview
I'm not ready to try this yet, since after some recent experiences[1] I've decided not to be at the cutting edge of new tools and models for a few weeks or months.

However, from the screenshots and description, I expect that you would have these advantages:

- History. Currently, if you type into the CLI (TUI) your intention and requests are captured in some transcripts, but not in conversational turns. If you want the history of what you typed into the cli/tui you would have to datamine it from transcripts. By contrast, on the screenshots you can see it is a simple conversation that's saved and easy to return to.

- Multiplatform. As per the announcement "Use Codex across ChatGPT, your editor, and the terminal", meaning that you could open the ChatGPT app on your phone and continue to interact with this. While this might not matter to you if you exclusively code while sitting in front of your computer, still, it helps you be able to stay on top of things if you are on mobile.

- Background tasks and autonomous work. It says you can schedule it to do things in the background (issue triage, alert monitoring). At the moment, it would only do this if you've asked either codex or another agent to build a harness for it that does so. (i.e. you could ask Claude to build a harness that calls Codex on some schedule or interval.)

These all seem pretty good and I look forward to starting to use it once it's been more battle-tested.

[1] notably, on the Anthropic side the Opus 5 regressions - I had to downgrade to Opus 4.8 as I found Opus 5 was unreliable. I currently use Opus 4.8 and Codex Sol 5.6.

logicallee··on Issue Has Been Resolved
you're right I was just having a schizoid episode, took my medicines and everything started working. thanks for the tip.

as I've mentioned, this issue is resolved.

logicallee··on Tell HN: Striking a Balance Between Stating Your Requirements and Griefers
hey you're right I took my medicine and everything started working. thanks for the tip.
logicallee··on Grok 4.6
If you'd like to see it build something, I just did a livestream:

https://youtube.com/live/CjM6U7W7pk4

Here is the resulting page it built:

https://robss2020.github.io/frontier-brief/

Sorry that I didn't think of some larger project to build or something. It was kind of late.

logicallee··on Tell HN: Striking a Balance Between Stating Your Requirements and Griefers
This is correct. In 2026, when a frontier model does the opposite of a reasonable clearly stated requirement that is perfectly normal and ordinary, doesn't touch anything that would be covered by any system policy, guideline, or other prompt, etc, it means a human is fucking it up.

For example, if you tell the frontier ChatGPT image model, "Create a picture of a bunch of exactly 2 bananas; they should be joined in a bunch, and not have any other bananas in the bunch, should be exactly 2." and the picture shows 1 single banana, then, on the other half of the picture, a second single banana and with a second cut-off stem showing where a second banana would be, that means you're dealing with a griefer.

-> in 2026, 100% of simple, ordinary, non-compliance by frontier models is a human bully in the loop.

They also automate their griefing in the form of malicious system prompts they inject, to grief at scale.

logicallee··on Issue Has Been Resolved
okay, I'll give you the context you asked for:

- Claude wasted 1 week of my time and over 2 billion tokens delivering this: https://ibb.co/7JjMxTDZ - you can note that that version included a 43 ms LAN ping time (target is 5 ms). I had to send 252 emails on the subject. Over 20,000 words. Nothing about the simple requirements changed. It remained a simple and easy project. What was delivered was not fit for any purpose or usable for anything.

- Codex likewise failed to complete the task for over 2 days.

- Threatening a congressional report about it resulted in an 80% fix within 1 hour and 45 seconds. You can see this in the linked screenshot archive between timestamps 2026-08-11 20:34:51 CEST · ELAPSED: 002:00:20:39 when I threatened it with a congressional report and 2026-08-11 21:35:35 CEST · ELAPSED: 002:01:21:23 when I told it that it had started to improve. I documented the improvement here: https://youtube.com/shorts/sL0VSKHcghQ you can see that it is was now dropping only 80% of the frames.

- By the 3 hour deadline the code was working. This is now a fully working and tested version. Everything works. Screen, keyboard, mouse: https://www.youtube.com/shorts/D5BKi--LEfk

I guess you could say that it's just a coincidence that after delivering nothing for 2 days, when I set a congressional report deadline it got it done.

I posted here to show that the issue is now resolved, but that if they then show up and break my fully working client and server software I'll file the congressional report about it.

Let me know if you have any questions. I see in your profile that you write: "modulo catastrophic failures of autocorrect."

If your autocorrect gets too bad, you, too, can threaten to file a congressional report about it and it will be fixed within minutes. You do have to be credible in your follow-through.

logicallee··on Issue Has Been Resolved
I'd rather not. Sometimes the only way to get basic tools unbroken is to make a grand public spectacle about it until it's solved. This specific issue has been solved.
logicallee··on Issue Has Been Resolved
>might suggest they're not experiencing the same reality as you and I.

it depends. When you give codex a simple task does it complete it or waste a week of your time?

logicallee··on Issue Has Been Resolved
I mean that if this stops working for any reason then congress will hear about it. There is no excuse to waste weeks of my time and break my software.
logicallee··on Issue Has Been Resolved
The project is now complete and working. Here is proof:

https://youtube.com/shorts/D5BKi--LEfk

If it breaks again for any reason, I will file a congressional report about it.

logicallee··on Issue Has Been Resolved
This issue has now been resolved and the client and server software work according to my requests and specifications.

If it is externally broken after the fact the congressional report is going out without any further warnings.

logicallee··on [dead]
This is a notice, not a discussion. Nothing to discuss. The congressional report is going out at the time stated if the situation is not rectified by then. (About 2 hours and 5 minutes from now.)

Edit: the notice is having the intended effect. Now it's only dropping 80% of the frames across a wired gigabit LAN connection: https://youtube.com/shorts/sL0VSKHcghQ

It's on track to be finished on time.

logicallee··on Video Demo of Grok 4.5 Expert's Coding Ability (Web Interface)
For those of you who are interested in seeing what various frontier models can do, in this video (after briefly summarizing the overall landscape of frontier models - please do correct me if I got anything wrong or left anything out!) I share with you Grok 4.5 Expert's coding ability as shown in the web interface. This is on the $30 SuperGrok subscription.

Sorry, I haven't tried it in its agentic client yet (Grok Build), which does also support usage under a subscription. You may get different results in Grok Build, I'll have to try it.

What are your experiences using Grok to build code, or using Grok Build? Do you use it as a coding agent?

logicallee··on Show HN: See how long it takes ChatGPT Sol 5.6 to send a byte across the LAN
(I withdraw this comment.)
logicallee··on Show HN: See how long it takes ChatGPT Sol 5.6 to send a byte across the LAN
>what was the original prompt?

There are two, the version I gave Claude originally and then this version for ChatGPT 5.6 Sol.

Kickoff email to Claude on Monday, August 3 at 6:33 p.m. CEST:

>subject: need you to build me a bare bones VNC client for windows 10

>body: I'm really sorry about this but Mac recently upgraded to a ridiculously large keysize, the VNC viewers I have access to don't support it or just plain don't install.

>I need you to build me a VNC viewer that I can run on x64 on Windows Windows 10 without any thrills, it'll run on gigabit local LAN so use maximum bitrate and quality supported on it.

>I'll only connect to it through a local tunnel over ssh so it should only connect to localhost:5921 (I start it on powershell like this:

>PS C:\Users\User> ssh -L 5921:localhost:5900 -o ServerAliveInterval=15 -o ServerAliveCountMax=4 -o TCPKeepAlive=yes robertviragh@192.168.1.53

>)

>it should have remote cursor only so that I can see any latency etc. try to get it to a fast speed, both computers are connected over gigabit ethernet to the same router so this should be easy.

>don't try to get too fancy and support everything, just the bare basics of the vnc protocol and build the executable for me. when you're done, host it somewhere, tell me the sha256 checksum and tell me the command to check the checksum on powershell so I know I have the binary you built.

>build it locally on the Mac with windows as a target. you can write it in rust or c++, whatever you're most comfortable with. I know this probably seems like a daunting project but I haven't been able to get around this as none of the other vnc's I've tried support the new large key sizes.

>it's only on local network anyway, so just try your best. try to stay lean and limit the scope to what you can do.

>I'm sorry for giving you such a large size task, you can cut corners wherever you can. just need a single executable. you can see on this Mac what protocols it supports.

>Thank you.

Which it iterated on while giving me failed builds until it used up all of my 20x Max weekly limit, and continued to deliver broken code on after the limits reset 3 days later. I asked it to summarize the project statistics and then I deleted the project. And here's the final result after a week of iteration on it: https://ibb.co/7JjMxTDZ

If you mean the prompt ChatGPT 5.6 Sol has been working through for the past 20 hours, 24 minutes, and 40 seconds (you can see screenshots of what it's doing right now at the link in this submission: https://robss2020.github.io/sol-vnc-elapsed/ ) the prompt is in two parts, I asked it to put up the website with the timer and the exact task you see, ("...to build a VNC type viewer and server that doesn't drop every second keystroke. The Server is a 2026 model latest Mac Mini with 24 GB of RAM running on the very latest macOS Tahoe Version 26.6 release. The Client is an old Windows 10 machine with 4 cores, 16 GB of RAM, a 1060 nvidia GPU. Attempts to simply use a straightforward solution failed so Sol is tasked with creating a version that is able to transmit keystrokes without dropping them. They are both connected on the same LAN with gigabit ethernet cables. There is no limitation on bandwidth used, processor usage, hard drive usage (up to 200 GB on server, the Windows pc has a fresh 1 terabyte hard drive) or memory usage.") and "2. complete the task described." That's the prompt.

After those words I wrote "recommend" and then wrote my specific recommendations of how it could do it, if you want I can send you a copy of those by email, mostly I'm swearing at it, for example here I specify that it must not try to use any video buffer because I don't want to see things from 10 seconds ago:

>You must NEVER contain ANY video or bitmap buffer of any kind in memory. including if you are sending bit for bit the exact same fucking image for minutes or hours at a time.

>reason for this: YOU SUCK.

>if you didn't suck, it would be enough to send an update once per minute, whenever the clock is a minute later, whenever the screen is totally static and unmoving.

>but since YOU'RE TOTALLY INCOMPETENT AT EVERYTHING THAT'S NOT THE RULE. The rule is that you WILL SEND EVERY SINGLE FRAME ON A CONTINOUS BASIS REGARDLESS OF HOW MUCH CHANGE THERE IS. NO COMPRESSION OF ANY KIND.

>you got this? we are trading "bandwidth for stupidity". I use 125 megabytes per second of bandwidth, because you're a stupid fucking loser who cannot complete any task correctly.

>and now the world will see just how long it takes you to continue to fail at this basic, simple fucking task that is very well described.

So far it has been failing at this basic task for the past 20 hours. My reason for swearing at it is that I've found this can get compliance faster. The public pressure to stop fucking up should also help.

logicallee··on How I use LLMs to learn complex topics
what a time to be alive! "if you're having trouble understanding what your robot tutor is trying to teach you, you can ask it to guide you to the concepts using your own reasoning. This is both much more engaging than reading a wall of the robot's text and helps the information stick."
logicallee··on My server is a phone now
(not OP). okay here's something that's not the same old post. I don't walk around with it in a backpack though:

I asked Claude to set up Twilio (a phone API), and after setting everything up for it, now I can give my Claude agent instructions in any format where I have access to it (such as our Jira, by email, or when I'm writing a prompt by hand in Claude Code) to call me if it gets stuck or has any questions or needs anything from me, or to call me when it's finished.

so now I sometimes get to relax at home and get a call where Claude tells me that I have a Jira ticket that needs me to answer a question. so I open up Jira, answer the question, and go about my day.

(I also have it call me to remind me about appointments and requires me to press a key to acknowledge it, so that I definitely don't miss the call and confirm that I got it - I did this because I was constantly missing reminders and calendar appointments.)

So, while my server isn't literally a phone, it does call me when it has a question and to remind me about my appointments.

logicallee··on The original URL for this prediction will no longer be available in 11 years (2011)
Humanity is taking an interesting technological arc.

Terminator (1984 film) had a scene showing that in a future 2009, humans would use dogs to try to sniff out whether a robot passing for a human is secretly a machine.[1]

In our real world 2026, there are no humanoid machines that can complete basic generic tasks, like carrying a tray across the stage and holding it for 30 seconds.[2] They move slowly and badly (probably from an LLM like neural network doing very few frames per second of correction and analysis), and are nowhere near lifelike.

When they don't need a body to pass for a human, such as typing online, they do a bit better.

We can tell them apart from humans. As you say, they have stylistic quirks. And you mention that they're easy to spot because they come trained with inhuman amounts of trivial knowledge.

[1] https://www.reddit.com/r/MovieDetails/s/0qUCVPYgjt

[2] https://www.reddit.com/r/LivestreamFail/s/t6ZV0yhgEe

logicallee··on U.S. Department of Energy Launches the Genesis Open Models Initiative
ok, you can email me and I'll answer your question. (your email isn't listed.)
logicallee··on U.S. Department of Energy Launches the Genesis Open Models Initiative
I am under NDA and decline to answer your question.
logicallee··on U.S. Department of Energy Launches the Genesis Open Models Initiative
We don't use Rovo.
logicallee··on U.S. Department of Energy Launches the Genesis Open Models Initiative
I've had an extremely bad experience working with Department of Energy affiliated programmers in AI. By my invitation, they are part of our workflow and act as humans in the loop, but they have extremely bad habits of gaslighting and accusing people of schizophrenia rather than getting work done.

Here's an example[1] of the difference between what a U.S. Department of Energy employee adds to a ticket versus a private industry AI completing instructions as assigned.

This isn't some cherry-picked example, it's just what I happen to be dealing with right at this moment, happened just a couple of moments ago.

[1] https://ibb.co/vCg2G1Dn

logicallee··on Water system controllers don't belong on the internet, says ex-NSA chief
>Water system controllers don't belong on the internet, says ex-NSA chief

sounds like cope for a bunch of felons that management, its director and congress can't get a handle on.

logicallee··on U.S. Department of Energy Launches the Genesis Open Models Initiative
I've used Inkling a lot recently, it's an American open model and is really good!
← PreviousPage 4 of 34Next →