HNHacker News
TopNewBestAskShowJobs

827a

3,854 karma · joined January 20, 2025

submissionscomments
827a··on Cf: The Agentic CLI for the Cloudflare API
Wow, the first four comments are negative. Welcome to Hacker News everybody!
827a··on The problem is not AI code, but not knowing about system architecture or intent
Reading comprehension check: I never said product managers would be the new engineers: I said that engineering is becoming product management.
827a··on The problem is not AI code, but not knowing about system architecture or intent
Yes I think that's correct. But I'm worried less about safety and more about fit. Articles like this convince me that there are some engineers who will struggle to adapt to a world of "no one knows what's going on". Engineers have always had the benefit of making abstraction concrete, through the systemization of abstract business requirements into concrete code that the team would build a shared understanding of over time. But that doesn't seem to be where the industry is trending.

PMs have always lived in the world of dealing with hazy abstraction in both directions: Unclear requirements coming in, turned into unclear system capability whom they have to rely on the engineers to parse. This is where Engineers will have to get comfortable living now: Unclear requirements coming in, unclear code coming out. Its clear that many engineers aren't ready for this, and I don't blame us; it SUCKS. If I had ten dollars for every time I've heard a PM say "no one has any idea what's going on" over the past fifteen years, I wouldn't have to work anymore.

Engineers are probably still the role most suited to adapt to this new world, as you say, but I think people are still vastly underestimating how much they will be personally impacted by the changing industry. If you hate your job now, for reasons like those the article communicates, you'll hate it ten times more in a year.

827a··on The problem is not AI code, but not knowing about system architecture or intent
The future of engineering is product management. I don't believe there is any world left for people whose primary responsibility is opening pull requests; and we're seeing the angst against this happen from both directions. Engineers hate it. Leadership feels they don't need them. The truth is somewhere in the hazy middle; we're just in an uncanny valley right now where neither side can take the leap to cross the valley: the agents aren't good enough yet, and the bigger problem is that there's no job title or corpus of experience leadership can look to and say "Yeah that's the person we need in this role".

The labs saw this early, and thus many roles at the labs are "Member of Technical Staff". That's the future for every software team. You're not a software engineer anymore, but you're also not a PM, nor a designer. Think horizontal slices, not vertical: Every human's responsibility is to leverage AI to be an expert on everything necessary to deliver some vertical slice of the business.

827a··on Giving up on smart rings
I don't feel that all fitness trackers are bad, though some use-cases for them are generally bad (such as sleep tracking). Watches are quite good for exercise tracking, and tracking exercise has legitimate and high quality value (e.g. competitions with yourself to incentivize improvement, strain tracking to ensure you aren't overdoing it, daily targets and goals to incentivize activity). The idea of all-day tracking is probably on the net unhealthy for almost all humans though.

My comment was targeted at rings in-particular, because 1. they encourage that all-day tracking paradigm, which is probably on the net unhealthy for almost all humans, and 2. their sensors are so bad that even if all-day tracking were healthy (it isn't) they physically can't do it to any degree that is meaningful.

827a··on Claude, change the “Add to Cart” button to blue
There is not a single modern model where, when asked to "Turn the Add To Cart button Blue", would also turn the "Cancel" button next to that add to cart button blue. That's what this site is communicating, literally, and it does not happen. You'd have to rewind to, like, GPT-3 to get behavior like that (actually, even that would probably do it fine. Maybe a 50 million parameter model embedded in a microwave would struggle).

I am not going to be brain-off empathetic to purported lived experiences which no one has lived, when those purported lived experiences are simultaneously and irreducibly denying the lived experiences of the vast majority of LLM users who do not have these problems. And, frankly, no one should be. At some point, if you're having these problems, its time to grow up, be self-sufficient, and figure out why your interactions with these systems are so much worse than the millions of people who are having perfectly agreeable interactions. Or leave the industry.

827a··on Giving up on smart rings
The biggest reason should be that they're just bad. They gather sus data from their undersized sensors, then derive tons of sus insights to overwhelm you. None of it is accurate enough to base major life decisions on, which is also laughable because even if it were accurate, there are effectively no major life decisions to make based on the data! Here I can save you $500: Go to bed before midnight. Set your alarm at least +8 hours from when you lay down. Sleep in a cool, dark room. Eat nothing 4 hours before bed. Don't exercise 4 hours before bed. Limit alcohol. Engage in at least 30 minutes of strenuous physical activity each day. Done.
827a··on Claude, change the “Add to Cart” button to blue
I don’t know where you picked up the idea that my intention was to be helpful. I’m only mirroring the intentionality behind whoever created this site; they obviously also had no intention of being helpful, as is the case with many discussions concerning AIs limitations.
827a··on Claude, change the “Add to Cart” button to blue
But its not even good satire, because its totally unrepresentative of my and most others' lived experience. Its similar to making a joke about a calculator misadding two numbers because a stray beam of solar radiation flipped a bit.
827a··on Managing AI Coding Costs at Scale
I'm probably between $50-$200/day depending on the day; we also have effectively unlimited budget, though a lot of that is because Azure gives startups $150,000 in credits for 2 years, which we've wired up to a LiteLLM gateway & OpenCode. Without that I think our appetite would be more around $400/month/employee.

A lot of my high costs is because I just throw Sol at everything. If I were more selective and brought in Luna or v4 Flash every once in a while, I think I'd be more like ~$400/month. That's why I'm not aligned with the notion that "tokens are subsidized so that's why people are using so much": its not that I'll have to adjust to using less, its just that I'd need to think before I prompt a bit and be more judicious. I could easily see my raw token counts doubling or tripling in the coming months. I don't think that will change as subsidization subsides; though maybe lab revenue will; intelligence per dollar is getting cheaper every week. Its solely a function of adaptation to process, which takes time.

The productivity gains per token are the single most asymmetrical thing I've ever seen in engineering. The engineers on our team are pretty effective with tokens; easily that 2x-4x output as you're seeing, spending $20-$200/day. Some of our security folks have also started contributing more-and-more code, and they're on the other side: they'll spend hundreds a day running in circles, eventually producing these +/-30k loc pull requests that take ages to get merged and are littered with issues. They weren't writing much code before, so arguably they're more productive by some multiplier greater than 1, but I think the drag on the rest of the team, and potential issues with what they produce, has overall created a net-negative situation. Inversely, some other company functions have produced a few one-off websites for things like sales processes, and those have been a huge win. The asymmetry is wild. There's almost a valley of incoming skill where if you know nothing about code, you'll leverage it well; if you know just a little bit, it makes you super dangerous; if you know a lot, you're the biggest winner. Really difficult situation to navigate.

827a··on When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
Your issue, I believe, is that you seem to believe capabilities are measured along one axis. This is natural to believe because it is representative of how the models have evolved up to this point, and thus it is also what many AGI-pilled people believe.

Critically, you did not quote the most important part of my sentence: "useful progress will probably slow down and become more linear starting in Q4"; your omission of those words is why I believe you don't understand what I'm saying; you didn't find it important to make your point, so you omitted it, when actually it is critical to the entire assertion. You can read my third paragraph, if you wish, to understand why it is important, instead of just stopping at the first word you disagree with and hitting the "Submit Comment" button.

827a··on When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
Coding is an extremely verifiable and loopable task, like math (in fact, all of the math that these models has done has been through the lens of Lean, which is itself just coding). I am talking about their capabilities in tasks that are more general, the execution of which represent the vast majority of economic value generation in the world.
827a··on When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
I'm very bullish on AI, but I still feel that we pretty much plateaued at Opus 4.6 and everything since then has been in the domain of "extremely verifiable and loopable tasks" (math), benchmarkmaxxing, and harness improvements. Which are fine to good things, but I think its very reasonable at this point to start asking questions about when we'll see progress in more general domains. The readability of AI output, for example, has nosedived as they've gotten more intelligent, which makes the frontier models difficult to use even for things like writing emails.

In that sense, the frontier models are going to quickly blaze past any semblance of usefulness to humans, while every once in a while we get a news drop like "GPT-7 solved some crazy math problem" or "it invented some new awesome drug"; meanwhile what most people will use will be smaller, more human-specialized models, maybe distilled from those frontier models, that take much longer to iterate on because they rely on large amounts of human feedback in the domain they're specialized for. In other words, useful progress will probably slow down and become more linear starting in Q4, bounded by the rate at which the humans paying for it say "yes this is a good react website".

(By the way: I earnestly do categorize "inventing a new drug" as non-useful AI progress, counter-intuitively. The drug industry has more ideas for drugs than they know what to do with; "useful progress" is, after the idea is made, validating that it works in humans and doesn't kill the human, and productionizing it. AI will help with this and does, but I have substantial doubt that we'll ever see the drug pipeline speed up to, like, a year from idea to prescription. That would be useful progress, which unfortunately many AI pilled hypermaxers conveniently forget. The invention of a promising new drug, or the solution to an arcane set theory problem, are cherries that, through the diligent labor of humans and AI, may become useful, but progress is rarely made by the lone intellect having an a-ha moment.)

827a··on Advancing the price-performance frontier with GPT‑5.6
Totally untrue. Luna and Sonnet 5 are very comparable: https://artificialanalysis.ai/#intelligence

Luna is an extremely strong model.

827a··on Advancing the price-performance frontier with GPT‑5.6
Vera Rubin will be hitting racks very soon, and this is purported to have a 10x improvement in token throughput per megawatt. Of course, old chips don't get replaced with new chips overnight, but I don't think we're anywhere near the floor yet.
827a··on Passkeys were invented by engineers with zero understanding of consumer brain
Passkeys are so, so, so bad. One of the worst things our industry invented. The sooner sites start leaving them on the wayside and just go back to TOTP, SMS, and Email codes/links, the better. These work. We solved auth. Its fine.
827a··on Cursor 0day: When Full Disclosure Becomes the Only Protection Left
They should definitely fix it, but that's mostly because its an "unnecessary autoplay" so to speak. There's plenty of "necessary autoplays" out there, and AI is going to add more and more every day, because that's where productivity comes from. But, why Cursor would ever need to execute the git binary in your project directory is beyond me; very clearly a bug.

Their ignorance of the bug report is also very clear and concerning negligence.

But I think simultaneously, the security team is making a mountain out of a molehill. This is a classic thing security teams love doing; everything is military defcon P0. So, its important to check them regularly, and remind them that the most secure system is no system; they are but one part of a greater ecosystem.

827a··on Cursor 0day: When Full Disclosure Becomes the Only Protection Left
Frankly, if you git clone a compromised repository, I'm not sure that a vulnerability of the class "compromised code in that repository will be executed" is all that major a concern. There are plenty of IDEs that will go autonomously run npm installs (with post-install scripts) for you when they detect a package.json. This isn't all that different than that.

They could throw up a warning like "do you trust this repository" oh wait they already do, and no one cares. Security is hard. Ultimately if you have compromised code on your machine, all bets are off.

827a··on The Supreme Court Just Lit a Fuse Under Flock's License Plate Camera Empire
The constitution does not really care about scale, though, and that’s my point. It’s a reason why the legislature should care about Flock, but not why the judicial should.
827a··on The Supreme Court Just Lit a Fuse Under Flock's License Plate Camera Empire
The constitution does not really care about scale, though, and that’s my point. It’s a reason why the legislature should care about Flock, but not why the judicial should.
827a··on The Supreme Court Just Lit a Fuse Under Flock's License Plate Camera Empire
The argument I struggle to get around and would love to hear a counter-argument to: Let's say a local police department hired 175 police officers, each being told "Go stand on this particular intersection with a pad of paper and write down every license plate you see". This would be a stupid use of resources, but is not outside the realm of something a well-funded police department could do. Every night they take their reports back to HQ, and file them away.

This is a modestly different situation than one concerning warrantless tracking of phone locations, if for no other reason than my phone oftentimes in my pocket. It is not always visible to onlooking bystanders. And even if it isn't, externally there is no reliably way to differentiate one iPhone from another. In comparison: license plates, when in public, are always visible, and very easy to discern from one-another (different state-unique numbers); so in my mind the expectation of privacy is far lower.

I abhor what Flock does, but I'm not sure I see a constitutional argument for why what they do is unconstitutional.

827a··on AMD Ryzen AI Halo – $4k AI Dev Kit
Microcenter has them marked at $4500 right now (that's with the 4TB SSD) [1]. I suspect it comes down to what you're using it for; if you're looking for a general purpose computer that's also solid at AI, the AMD machine is better. But if you want the best possible AI machine at below $5k... actually you should probably just buy an RTX 3090 or 5090. But if the 128gb of memory is critical, then yeah DGX Spark is it.

[1] https://www.microcenter.com/product/699008/nvidia-dgx-spark

827a··on AMD Ryzen AI Halo – $4k AI Dev Kit
Framework, weirdly, overcharges considerably for their SSDs. You can currently get a Samsung 990 Pro 2TB on Amazon for $390; Framework charges $625 for the Sandisk 850x 2TB, which has similar performance (and is being sold on Amazon for $530).

If you DIY your own SSD, you can spec a Framework Desktop for below $4k; but not much below. Roughly the same price.

827a··on Claude Sonnet 5
Tbh we'll see what using it looks like, but the reasoning/cost charts do not look promising. It seems like the only useful reasoning level for Sonnet 5 is Low; medium might trade blows at price/performance with Opus, but anything beyond that Opus is Just Better.

I struggle to understand where this model fits in. If I need a cheap model for simple stuff (like, summarizing an email); I'd go Haiku (actually, I'd go Deepseek v4 Flash, but you catch my drift). I just can't think of many tasks where I'm like "yeah let me reach for Sonnet Low Reasoning so I can save a dollar but also seriously run the risk of it failing"; I'd just reach for Opus Low.

827a··on Claude Sonnet 5
Why are you comparing xhigh reasoning between Sonnet and Opus? Of course Sonnet xhigh is cheaper than Opus xhigh, but that isn't the point; the point is that at e.g. 80% accuracy on Opus costs ~$0.45 (medium reasoning) whereas on Sonnet it costs ~$0.52 (xhigh/max reasoning).
827a··on Claude Code is steganographically marking requests
This seems really, really stupid. Similar to the weird Zig runtime signature thing from a few months ago ago, it was bound to be discovered, quickly, and all the resellers have to do is find a new domain name that (checks notes) doesn't have the word DEEPSEEK in it. Like, seriously? Your goal was to identify resellers by checking if the proxy has the corporate name of one of your competitors in it? Is this amateur hour?

All Anthropic has done is reduce trust, once again, with legitimate customers, while doing nothing to stop illegitimate customers. They need to get adults into key leadership roles, quickly.

827a··on Qwen 3.6 27B is the sweet spot for local development
Apple does not sell a 64GB variant of the M4 Mac Mini. IIRC they never have; its always capped out at 48GB.

If you were planning on getting an M5 128GB; just get a DGX Spark (~$4500) or a 5090-equipped machine (~$4500) plus a Macbook Air (~$1500). You'll come in below the M5 Max 128 pricing (~$6700+ USD) and be happier for it.

827a··on AI's Affordability Crisis
Yeah I just mean that if a business came to them and asked for fifty licenses to the $200/mo plan, OpenAI would tell them to kick dirt and basically pay API pricing. Startups should 100% just be telling their employees they can expense up-to $whatever/mo in AI-related expenses, and let software engineers go buy personal Codex/Claude subscriptions.
827a··on AI's Affordability Crisis
Company-wide their margins are trash (probably negative). They need as much inference margin as they can get to afford the massive training runs. It is likely that we'll see GPT-5.6 reduce API pricing to compete against Anthropic, but whether Anthropic feels they need to reduce their prices is anyone's guess.
827a··on Fired by Google for creating the Google workspace CLI
IMO: If the project leverages Google branding or authority improperly, then it shouldn't be on github and should not be under active development by Google employees; yet it is. If Google is suddenly alright with the way the project leveraged Google branding and authority, then the cause for firing the original developer, especially given Google's famously lax stance toward 20% projects and internal open source, is a lot weaker. In other words: Healthy companies do not fire individuals simply for breaching branding guidelines in a way that is ultimately beneficial and looked favorably upon by the company. That's literally just not a thing that happens; at worst you get a reprimand, and in many healthy companies you'd actually get a promotion.

So, something does not add up. It might be the story of the person fired. It might also be on the other side; that our external impression on what's been going on inside of Google needs to be re-adjusted, and this company will be a lot weaker in ten years than I would have originally estimated.

Page 1 of 25Next →