HNHacker News
TopNewBestAskShowJobs

rectang

16,680 karma · joined October 5, 2014

submissionscomments
rectang··on Trump fires NSF's oversight board
Neither the Butler, Pennsylvania nor the White House Correspondents Dinner incidents were staged — that's conspiracy theory nutter nonsense.

(Not that there aren't elements within the Trump administration who might contemplate a false flag operation, but it will be obvious if they try it because they aren't clever enough to pull it off.)

Furthermore, both Ceaucescu and Mussolini were extrajudicial executions.

Nixon is a better precedent for the USA because it demonstrates that we're already capable of ejecting a criminal president without bloodshed. It wasn't a perfect process (the Ford pardon was terrible) but in the end Nixon still lost power.

rectang··on Trump fires NSF's oversight board
With your focus on death as punishment — the day after an apparent assassination attempt no less — it seems we have less in common than I supposed. I’m not with you after all.
rectang··on Trump fires NSF's oversight board
I'm very much with you for "punish" and "jail", but if you are insisting on execution like commenters upthread, we will part ways there.
rectang··on Trump fires NSF's oversight board
Trying to find a silver lining and think positively...

Will a future administration have an opportunity to build something new and better from scratch which would not have been possible due to institutional resistance before it was all burnt down?

rectang··on I cancelled Claude: Token issues, declining quality, and poor support
Commoditization will be complete for my purposes when an LLM trained on a legitimately licensed corpus can achieve roughly what Opus 4.5+ or the highest powered GPTs can today.

I anticipate a Napster-style reckoning at some point when there's a successful high-profile copyright suit around obviously derivative output. It will probably happen in video or imagery first.

rectang··on I cancelled Claude: Token issues, declining quality, and poor support
I have ethical qualms to varying degrees with most LLMs, primarily because of copyright laundering.

I'm a BSD-style Open Source advocate who has published a lot of Apache-licensed code. I have never accepted that AI companies can just come in and train their models on that code without preserving my license, just allowing their users to claim copyright on generated output and take it proprietary or do whatever.

I would actually not mind licensing my work in an LLM-friendly way, contributing towards a public pool from which generated output would remain in that pool. Perhaps there is opportunity for Open Source organizations to evolve licenses to facilitate such usage.

For what it's worth, I would be happy to pay for a commercial LLM trained on public domain or other properly licensed works whose output is legitimately public domain.

rectang··on I cancelled Claude: Token issues, declining quality, and poor support
I feel like I'm using Claude Opus pretty effectively and I'm honestly not running up against limits in my mid-tier subscriptions. My workflow is more "copilot" than "autopilot", in that I craft prompts for contained tasks and review nearly everything, so it's pretty light compared to people doing vibe coding.

The market-leading technology is pretty close to "good enough" for how I'm using it. I look forward to the day when LLM-assisted coding is commoditized. I could really go for an open source model based on properly licensed code.

rectang··on US special forces soldier arrested after allegedly winning $400k on Maduro raid
The broader theme of antagonism to Black success motivating the thoroughness of the destruction is a common observation about Tulsa.
rectang··on French government agency confirms breach as hacker offers to sell data
Seeing another one of these breaches had me returning to look at local-first software. https://lofi.so

I feel like if we're going to make progress in preventing wholesale data breaches it will be through architectural innovations that attack the problem of why a trove of concentrated data needs to exist. Even if the government needs to be a central authority, are there ways to house the data that limit the blast radius?

I'm sure there are innumerable arguments why this can't help, but when the mainstream alternative is despair and helplessness, progress will be made in the margins.

rectang··on Claude Code to be removed from Anthropic's Pro plan?
I see lots of speculation that Anthropic needs to cut usage because they are compute constrained. If that's the case, will they be focusing on reducing compute costs for their models?

From what I can tell Opus 4.7 is more resource-intensive than Opus 4.6 is more resource-intensive than Opus 4.5.

rectang··on Cal.diy: open-source community edition of cal.com
Who says I believe it? ;)

I'm just choosing to focus on the substance of the argument itself, which I think is risible regardless of who makes it and why.

rectang··on Cal.diy: open-source community edition of cal.com
I'm unpersuaded by the assertion that closing the source is an effective security bulwark.

From that page:

> Today, AI can be pointed at an open source codebase and systematically scan it for vulnerabilities.

Yeah, and AI can also be pointed at closed source as soon as that source leaks. The threat has increased for both open and closed source in roughly the same amount.

In fact, open source benefits from white hat scanning for vulnerabilities, while closed source does not. So when there's a vuln in open source, there will likely be a shorter window between when it is known by attackers and when authors are alerted.

rectang··on Changes to GitHub Copilot Individual Plans
I've started messing with this and the experience seems pretty similar.
rectang··on Changes to GitHub Copilot individual plans
My experience is similar. Opus, especially Opus 4.5, understands my intentions better even when poorly phrased, and more consistently follows my instructions to do only what's necessary and no more.

As far as I can tell, the distinctive feature of my workflow is that I'm giving it small, contained single-commit-sized tasks and limited context. For instance: "For all controller `output()` functions under `Controller/Edit/` and `Controller/Report/`, ensure that they check `Auth::userCanManage`." Others seem to be taking bigger swings.

rectang··on Changes to GitHub Copilot individual plans
Welp. I already added a $20 Claude Pro subscription to complement my $10 Github Copilot Pro subscription and $10 DuckDuckGo Plus. That was partly to show support for Anthropic after the OpenAI/DOD episode, but also because I've been using Opus 4.5 exclusively with Copilot and I figured I should try Claude Code eventually.

Now it's going to cost me an upgrade to $39 Github Pro+ to keep using Opus, and even then it's with much higher multipliers. I don't fully understand the extent to which this reflects actual costs for Opus versus Microsoft leveraging network effects to discourage the usage of a competitor.

I didn't really want to wander outside of VSCode just yet because I was happy with VSCode/Copilot/Opus-4.5 and I don't want to spend all my time experimenting when stuff is changing so fast. But I guess my hand has been forced.

rectang··on Deezer says 44% of songs uploaded to its platform daily are AI-generated
I suppose that my generalization was too broad and that LLMs can be either good or bad at writing tests depending on your workflow and expectations.

I'm closely supervising the LLM, giving it fine-grained instructions — I generally understand the full interface design and most times the whole implementation (though sometimes I skim). When I have the LLM write unit tests for me, it writes essentially what I would have written a couple years ago, except that it tends to be more thorough and add a few more tests I wouldn't have had the patience to write. That saves me quite a bit of time, and the LLM-generated unit tests are probably somewhat better than what I would have written myself.

I won't say that I never see brain-dead mistakes of the "5-vertex square" variety (haha) — by their nature, LLMs tend towards consistency rather than understanding after all. But I've been using Claude Opus exclusively for while and it doesn't tend to make those mistakes nearly as often as I used to see with lower-powered LLMs.

rectang··on Deezer says 44% of songs uploaded to its platform daily are AI-generated
44% of uploads are probably not created by 44% of "artists". The core of people who are looking to exploit the system are going to be good at gaming the recommendation algorithm — they're specialists in it solely for the money who don't need to trouble themselves with artistic concerns.
rectang··on Deezer says 44% of songs uploaded to its platform daily are AI-generated
> I recently deleted a whole bunch of automated tests because if the AI is going to write most of the code then I should test it to make sure it's good!

??

You say you deleted the tests, because you "should test it"? The logic seems inconsistent.

Sanity checking LLM-generated code with LLM-generated automated tests is low-cost and high-yield because LLMs are really good at writing tests.

rectang··on Anonymous request-token comparisons from Opus 4.6 and Opus 4.7
Are the benchmarks being used to measure these models biased towards completing huge and highly complex tasks, rather than ensuring correctness for less complex tasks?

It seems like they're working hard to prioritize wrapping their arms around huge contexts, as opposed to handling small tasks with precision. I prefer to limit the context and the scope of the task and focus on trying to get everything right in incremental steps.

rectang··on Anonymous request-token comparisons from Opus 4.6 and Opus 4.7
I've used Sonnet a lot. It is not as good as Opus at understanding what I'm asking for. I have to coach Sonnet more closely, taking more care to be precise in my prompts, and often building up Plan steps when I could just YOLO an Agent instruction at Opus and it would get it right.

I find that Opus is really good at discerning what I mean, even when I don't state it very clearly. Sonnet often doesn't quite get where I'm going and it sometimes builds things that don't make sense. Sonnet also occasionally makes outright mistakes, like not catching every location that needs to be changed; Opus makes nearly every code change flawlessly, as if it's thinking through "what could go wrong" like a good engineer would.

Sonnet is still better than older and/or less-capable models like GPT 4.1, Raptor mini (Preview), or GPT-5 mini, which all fail in the same way as Sonnet but more dramatically... but Opus is much better than Sonnet.

Recent full-powered GPTs (including the Codex variants) are competitive with Opus 4.6, but Opus 4.5 in particular is best in class for my workflow. I speculate that Opus 4.5 dedicates the most cycles out of all models to checking its work and ensuring correctness — as opposed to reaching for the skies to chase ambitious, highly complex coding tasks.

rectang··on Anonymous request-token comparisons from Opus 4.6 and Opus 4.7
https://github.blog/changelog/2026-04-16-claude-opus-4-7-is-...

> Over the coming weeks, Opus 4.7 will replace Opus 4.5 and Opus 4.6 in the model picker for Copilot Pro+[...]

> This model is launching with a 7.5× premium request multiplier as part of promotional pricing until April 30th.

rectang··on Opus 4.7 to 4.6 Inflation is ~45%
For my working style (fine-grained instructions to the agent), Opus 4.5 is basically ideal. Opus 4.6 and 4.7 seem optimized for more long-running tasks with less back and forth between human and agent; but for me Opus 4.6 was a regression, and it seems like Opus 4.7 will be another.

This gives me hope that even if future versions of Opus continue to target long-running tasks and get more and more expensive while being less-and-less appropriate for my style, that a competitor can build a model akin to Opus 4.5 which is suitable for my workflow, optimizing for other factors like cost.

rectang··on Anonymous request-token comparisons from Opus 4.6 and Opus 4.7
For now, I'm planning to stick with Opus 4.5 as a driver in VSCode Copilot.

My workflow is to give the agent pretty fine-grained instructions, and I'm always fighting agents that insist on doing too much. Opus 4.5 is the best out of all agents I've tried at following the guidance to do only-what-is-needed-and-no-more.

Opus 4.6 takes longer, overthinks things and changes too much; the high-powered GPTs are similarly flawed. Other models such as Sonnet aren't nearly as good at discerning my intentions from less-than-perfectly-crafted prompts as Opus.

Eventually, I quit experimenting and just started using Opus 4.5 exclusively knowing this would all be different in a few months anyway. Opus cost more, but the value was there.

But now I see that 4.7 is going to replace both 4.5 and 4.6 in VSCode Copilot, and with a 7.5x modifier. Based on the description, this is going to be a price hike for slower performance — and if the 4.5 to 4.6 change is any guide, more overthinking targeted at long-running tasks, rather than fine-grained. For me, that seems like a step backwards.

rectang··on Google broke its promise to me – now ICE has my data
Sounds like promises are worthless and only capabilities matter.
rectang··on Volunteers turn a fan's recordings of 10K concerts into an online treasure trove
Cassettes are a pain. Head alignment is extremely important for analog tape fidelity, and it's always off for home recordings.

With pro analog tape recordings (e.g. 2-inch 24 track, half-inch 2-track), you record alignment tones onto the tapes to capture the state of the recording device, and then later calibrate the playback device to the particular tape so that playback alignment matches recording alignment. But this is essentially never done with cassettes, so you have to earball it.

Cassette players for mastering studios actually have alignment options (e.g. adjustable azimuth) that aren't present on consumer devices. But without the tones, you have to guess.

The problem with starting from a digitized source is that it may have been digitized from non-aligned playback. Ideally you want to go back to the analog originals - but old cassettes are rarely in perfect condition.

rectang··on HBO Obtains DMCA Subpoena to Unmask 'Euphoria' Spoiler Account on X
To be clear, I'm not sitting in judgment of you or any of the other spoiler trolls, not back then and certainly not now. This is an instance where the Potter-philes couldn't fight back, and to my mind that's inevitably going to bring out the worst in human nature.

Elsethread, you mention that "some people really got upset.". In some sense, the more upset they get, the more successful the troll and the funnier it gets, right? At least, it feels funny to me, at the same time as it also feels bad to imagine upset kids, at the same time as feeling that upset kids learning that other humans are cruel is a necessary part of growing up.

rectang··on HBO Obtains DMCA Subpoena to Unmask 'Euphoria' Spoiler Account on X
There's never going to be philosophical consensus on the "good/evil/social/antisocial" debate because the human impulse to self-justify and believe that you're the "good guy" is extremely powerful. Those of us who seek to understand human nature have to proceed without consensus as a goal.
rectang··on HBO Obtains DMCA Subpoena to Unmask 'Euphoria' Spoiler Account on X
A mild reminder that humans tend towards inflicting pain on their fellow humans when there are no consequences.
rectang··on Microsoft is employing dark patterns to goad users into paying for storage?
Selective incompetence for fun and profit:

Pricing mistakes which make the supermarket money are unfortunate but low priority. Pricing mistakes which cost the supermarket money must be fixed immediately.

rectang··on IPv6 is the only way forward
If UTF-8 represents the triumph of a design prioritizing backwards compatibility with an existing standard (ASCII) to facilitate a transition, then IPv6 is the cautionary tale of a design which could have made the transition simpler but did not.
← PreviousPage 5 of 34Next →