3,854 karma · joined January 20, 2025
PMs have always lived in the world of dealing with hazy abstraction in both directions: Unclear requirements coming in, turned into unclear system capability whom they have to rely on the engineers to parse. This is where Engineers will have to get comfortable living now: Unclear requirements coming in, unclear code coming out. Its clear that many engineers aren't ready for this, and I don't blame us; it SUCKS. If I had ten dollars for every time I've heard a PM say "no one has any idea what's going on" over the past fifteen years, I wouldn't have to work anymore.
Engineers are probably still the role most suited to adapt to this new world, as you say, but I think people are still vastly underestimating how much they will be personally impacted by the changing industry. If you hate your job now, for reasons like those the article communicates, you'll hate it ten times more in a year.
The labs saw this early, and thus many roles at the labs are "Member of Technical Staff". That's the future for every software team. You're not a software engineer anymore, but you're also not a PM, nor a designer. Think horizontal slices, not vertical: Every human's responsibility is to leverage AI to be an expert on everything necessary to deliver some vertical slice of the business.
My comment was targeted at rings in-particular, because 1. they encourage that all-day tracking paradigm, which is probably on the net unhealthy for almost all humans, and 2. their sensors are so bad that even if all-day tracking were healthy (it isn't) they physically can't do it to any degree that is meaningful.
I am not going to be brain-off empathetic to purported lived experiences which no one has lived, when those purported lived experiences are simultaneously and irreducibly denying the lived experiences of the vast majority of LLM users who do not have these problems. And, frankly, no one should be. At some point, if you're having these problems, its time to grow up, be self-sufficient, and figure out why your interactions with these systems are so much worse than the millions of people who are having perfectly agreeable interactions. Or leave the industry.
A lot of my high costs is because I just throw Sol at everything. If I were more selective and brought in Luna or v4 Flash every once in a while, I think I'd be more like ~$400/month. That's why I'm not aligned with the notion that "tokens are subsidized so that's why people are using so much": its not that I'll have to adjust to using less, its just that I'd need to think before I prompt a bit and be more judicious. I could easily see my raw token counts doubling or tripling in the coming months. I don't think that will change as subsidization subsides; though maybe lab revenue will; intelligence per dollar is getting cheaper every week. Its solely a function of adaptation to process, which takes time.
The productivity gains per token are the single most asymmetrical thing I've ever seen in engineering. The engineers on our team are pretty effective with tokens; easily that 2x-4x output as you're seeing, spending $20-$200/day. Some of our security folks have also started contributing more-and-more code, and they're on the other side: they'll spend hundreds a day running in circles, eventually producing these +/-30k loc pull requests that take ages to get merged and are littered with issues. They weren't writing much code before, so arguably they're more productive by some multiplier greater than 1, but I think the drag on the rest of the team, and potential issues with what they produce, has overall created a net-negative situation. Inversely, some other company functions have produced a few one-off websites for things like sales processes, and those have been a huge win. The asymmetry is wild. There's almost a valley of incoming skill where if you know nothing about code, you'll leverage it well; if you know just a little bit, it makes you super dangerous; if you know a lot, you're the biggest winner. Really difficult situation to navigate.
Critically, you did not quote the most important part of my sentence: "useful progress will probably slow down and become more linear starting in Q4"; your omission of those words is why I believe you don't understand what I'm saying; you didn't find it important to make your point, so you omitted it, when actually it is critical to the entire assertion. You can read my third paragraph, if you wish, to understand why it is important, instead of just stopping at the first word you disagree with and hitting the "Submit Comment" button.
In that sense, the frontier models are going to quickly blaze past any semblance of usefulness to humans, while every once in a while we get a news drop like "GPT-7 solved some crazy math problem" or "it invented some new awesome drug"; meanwhile what most people will use will be smaller, more human-specialized models, maybe distilled from those frontier models, that take much longer to iterate on because they rely on large amounts of human feedback in the domain they're specialized for. In other words, useful progress will probably slow down and become more linear starting in Q4, bounded by the rate at which the humans paying for it say "yes this is a good react website".
(By the way: I earnestly do categorize "inventing a new drug" as non-useful AI progress, counter-intuitively. The drug industry has more ideas for drugs than they know what to do with; "useful progress" is, after the idea is made, validating that it works in humans and doesn't kill the human, and productionizing it. AI will help with this and does, but I have substantial doubt that we'll ever see the drug pipeline speed up to, like, a year from idea to prescription. That would be useful progress, which unfortunately many AI pilled hypermaxers conveniently forget. The invention of a promising new drug, or the solution to an arcane set theory problem, are cherries that, through the diligent labor of humans and AI, may become useful, but progress is rarely made by the lone intellect having an a-ha moment.)
Luna is an extremely strong model.
Their ignorance of the bug report is also very clear and concerning negligence.
But I think simultaneously, the security team is making a mountain out of a molehill. This is a classic thing security teams love doing; everything is military defcon P0. So, its important to check them regularly, and remind them that the most secure system is no system; they are but one part of a greater ecosystem.
They could throw up a warning like "do you trust this repository" oh wait they already do, and no one cares. Security is hard. Ultimately if you have compromised code on your machine, all bets are off.
This is a modestly different situation than one concerning warrantless tracking of phone locations, if for no other reason than my phone oftentimes in my pocket. It is not always visible to onlooking bystanders. And even if it isn't, externally there is no reliably way to differentiate one iPhone from another. In comparison: license plates, when in public, are always visible, and very easy to discern from one-another (different state-unique numbers); so in my mind the expectation of privacy is far lower.
I abhor what Flock does, but I'm not sure I see a constitutional argument for why what they do is unconstitutional.
[1] https://www.microcenter.com/product/699008/nvidia-dgx-spark
If you DIY your own SSD, you can spec a Framework Desktop for below $4k; but not much below. Roughly the same price.
I struggle to understand where this model fits in. If I need a cheap model for simple stuff (like, summarizing an email); I'd go Haiku (actually, I'd go Deepseek v4 Flash, but you catch my drift). I just can't think of many tasks where I'm like "yeah let me reach for Sonnet Low Reasoning so I can save a dollar but also seriously run the risk of it failing"; I'd just reach for Opus Low.
All Anthropic has done is reduce trust, once again, with legitimate customers, while doing nothing to stop illegitimate customers. They need to get adults into key leadership roles, quickly.
If you were planning on getting an M5 128GB; just get a DGX Spark (~$4500) or a 5090-equipped machine (~$4500) plus a Macbook Air (~$1500). You'll come in below the M5 Max 128 pricing (~$6700+ USD) and be happier for it.
So, something does not add up. It might be the story of the person fired. It might also be on the other side; that our external impression on what's been going on inside of Google needs to be re-adjusted, and this company will be a lot weaker in ten years than I would have originally estimated.