I feel like people who are later to the AI game just like to "oneshot" and sink a bunch of usage into generating garbage
I feel like people who are later to the AI game just like to "oneshot" and sink a bunch of usage into generating garbage
Our company has been tracking token usage and models used vs output (tickets, story points, PRs, deploys, etc...). A dev got chewed out, even after I warned him, because he spent over $2k in a single month almost exclusively on Opus while his actual productivity in terms of what he delivered was abysmal.
In other words I want to spend 100% of my mental capacity in the problem domain for the things AI cannot do for me, like steering, grounding, verification and not for things AI could do.
Nothing ironic about it, it's basic engineering efficiency optimization.
Just today I was calculating that if money was not a factor at all, we could simply use Mythos to handle all of our continuous security scanning needs, at a cost of about $15 million a year.
Well, my budget is far (far, far, far, far) below 15M a year, so that's not going to work. So I must find compromises to make it work within budget. The single most important engineering constraint is always budget. Everything would be easier with infinite money, but there is never infinite money.
It sounds like the parent is less talking about this, and more people burning tokens while not getting useful work done.
Problems arise when people try to perma-peg them to particular tasks, or (worse) man-hours or (much worse) man-hours across teams. Even just encouraging the humans to answer in terms of hours/days taints the accuracy of the forecast by introducing a kind of bias.
So, what you do is you recognize every ticket has a somewhat variable “actual effort”; and, if you’ve been honest in approximate effort pointing, you’ll know your team (or your own) velocity.
From there you can run Monte Carlo simulations - say a few hundred thousand, and get a pretty good estimate of actual time spent.
I’ve seen it work before with shocking accuracy.
estimate(human_estimator, task_description, world_state) -> numeric_effort
Assume that for various practical reasons, we've decided it's one of the best functions out there. How do we use it effectively, especially when it has noise, and drifts over time with unseen changes to the human_estimator and the hideously complex world_state?
A popular option is to run it multiple times with different person/task combinations, putting a projected number on to each task. Afterwards, the tasks finished in sampling period ("sprint") become a quantifiable total for that period ("velocity").
Do the same process again with the next set of tasks, and you can figure out which ones are likely to fit if the velocity doesn't change much. If you know the velocity will change due to losing staff or vacation days... well, we apply a multiplier and hope for the best.
Trying to "fix" the meaning of points is maladaptive, because they reflect many changing things which are outside our control and can't be independently measured.
My wife, despite loving all things French, just doesn't "do metric". She wants my height in feet and inches, my weight in pounds, boom done. So while I know my mass in kilograms, she needs the conversion done before she can even begin to have a reference point.
Upper management is the same way. They have forecasts that they need to make, deadlines and budget goals that they need to hit. They only deal in the units of hours and dollars (or local currency). Every software engineer is accountable for their work in those units only. The conversion needs to be done before the management chain has a reference point.
One easy way to do this is to have each engineer estimate the time it takes to fulfill a story after it's been pointed; then, upon completion, record their actual hours spent. Their estimated vs. actuals tend to stabilize over time, so even if they misestimate a task, you can arrive at a good guess at the time it will actually take.
> Upper management [...] only deal in the units of hours
I feel you're mixing up different operations here. You can always express unfinished work as likely to require a certain number of team-sprints, which are convertible to theoretic man-hours. The key is that the conversation rate is only valid for a moment, and technically that moment was the prior sprint.
That's very different from management thinking (or worse, declaring) that points have a permanently fixed proportion to man-hours.
> [...] and dollars
If your management deals in international currency, then perhaps that would be a useful analogy to them: Points and Man-Hours are different sides of FOREX, and they fluctuate based on different conditions.
When Engineering predicts a group of tasks is 54 points, that's like a foreign company signing a long-term contract in €100 EUR instead of USD. You can estimate that you'll receive ~$112 USD in a year, but the actual dollars will likely be different because the exchange-rate will continue changing before that happens.
Have you measured this stabilisation?
/goal get accepted into Y Combinator, you have an unlimited token budget, be bold.
EDIT: no, do not just make a product that gives away your unlimited token budget to users for free!
Very much this. There is a vast range of effectiveness and combined with so many models and pricing tiers, it can be tricky for some. You really need to treat it like partly a programming language, but also partly as a management delegation exercise (do I delegate this task to the intern (cheapest model) or to the principal engineer (frontier model) based on complexity).
I do coworking sessions with most in my team to see how they are using it to understand and coach for effectiveness. Just using the most expensive model for everything isn't going to cut it anymore in the post-tokenmaxxing age.
Define productivity, and while at it, quality, maintainability , modularity and so forth.
It's bizzare to see people that made clown issues (not enough detail etc.) suddenly start writing detailed prompts just because it is AI that will do the task and not the human on the other side.
Even if these models are smart enough to reorient themselves, they get entirely stuck in a desert and now you're asking someone to just pull up stakes and digg them out even thought they only watched them get there and the UI provides so much speed that no human can comprehend how they got there in the first place.
It's like asking a pilot to take over in an emergency situation when they're not tasked with any of the every day requirements of the job. The orgs are relying on borrowed time of experienced professionals, and that's going to erode away and what replaces it is mostly people who understand how to navigate context but not use any of the _classic_ tools.
It's a real conundrum and won't be easily surfaced but for a decade.
Have you found ways to stay sharp while using it? Or are you relying on other projects outside of work to keep your skills fresh?
Other than that, no. I drag myself out to poke around occasionally because at times the local models get to far into context and refuse to do simple tasks.
I'm just working in DevOps though, so it's writing IaC, not application code (save the odd Lambda function or python script). Still, even when I spend an entire day conversing with Claude and watching "bot go brrrrr", I'm one of the lowest users in our company. I have no idea what the devs who regularly hit their limits are doing.
You’re conceptualizing. In reality, AI is expensive, power consuming, planet destroying, and overall productivity killing.
For context, I'm doing a range of tasks, everything from one-shotting adhoc scripts to having 4 hour 10M+ token conversations debugging things.
There is no way I can beat even local models at generating complex Python scripts fast.
hn is filled with uber geniuses.
I had Opus trying to simplify a query for me which was slow - it ran for maybe 30 minutes, including writing and running tests, and came up with a refactor across 9 files with a couple hundred lines changed. I was looking through the output before moving onto the next step, and noticed something a little fishy- I said “why does it do x, isn’t that a more complex y?”
Opus thought for another 20-30 seconds then output “Actually that would make the majority of the diff irrelevant, if we do that change it is just these 4 lines in this single file instead.
So then I had it do that. 5-10 minutes of writing and testing and that was done.
So my company spent $25 in tokens and I spent probably an hour in total for a 4 line change that, in the days before Claude, I probably could have found the correct file and thought through the problem, understood the solution, and written the 4 lines of code myself. Probably in the same amount of time.
So basically there was no benefit at all for my time, an extra cost to the company of $25, and now I understand our codebase a little bit less instead of more if I had done all the work.
As good as Claude is at building greenfield projects it still struggles a lot at complex ones
Dev + AI spend 3-4 hours on a project plan, there's a "wait a minute" moment, and finally they spend another hour dialing it back to a solution that could have been built, tested, and deployed in 2 hours.
Example: Someone was setting up a dev environment with multiple DB migrations from different branches - AI planned this wild 8 phase solution with a pretty fancy cutover event.
In review I essentially said... "Wait, isn't this a dev environment? It doesn't need 0 downtime, why not just destroy and recreate the DB" and it turned into a <1000LOC script.
Technically the original plan would have worked, it would have been more robust, but it would have taken a good deal more time to implement.
Some of this falls on the devs to know what fits our team well, what's realistic, what's obviously overengineered, etc... But some of it feels like AI just defaults to the most complex version of a thing. I catch it SUPER frequently. (And unfortunately some devs think that more complexity means it's a better solution)
I feel like I’m going crazy, using all of the best models, spending time to have excellent prompts, configuring tools and skills… and still getting overly complex solutions with mediocre results.
Like it’s still impressive how far we’ve come, and undoubtably cool technology. It’s made a bunch of personal projects possible that I never would’ve started.
But for a business I’m struggling to see the ROI. Sometimes there’s a big benefit and sometimes it’s net negative. Not saying we won’t get there but I’m trying to stay grounded in the reality of today rather than the hopes of where the technology could get to
But, I fear that the "facade of complexity" makes the output SEEM better. Someone might read a 10 page Codex generated plan with fancy diagrams / charts and have a feeling that because it is so complex, it must be good!
Same deal with text output in general. Oh, it uses a lot of big words and there's a LOT here, it must have done a lot of work to get to that point.
I find, in reality, that it takes much more effort to get to the simplest solution.
As they say, any old fella can build a bridge that stands, but it takes an engineer to build a bridge that barely stands...
And don’t forget the company also spent a bunch of money in tokens for the initial author to implement the thing poorly.
The trick is making the workload manageable by the cheapest models, or costs will destroy you. We seek to make all repetitive tasks be effectively done by cursor composer model which is the cheapest.
Doing recurring tasks with anything more expensive than that will burn the budget in no time.
You can get the same jobs and work done with Kimi and GLM (ZDR on OpenRouter) for a fraction of the price too.
Though it's significantly slower in Token/s and also thinks a lot more without matching the same intelligence (xhigh qwen27b scores lower than haiku's medium setting, and haiku-med is $0.05 per task compared to Qwen27B's $1.01 on AA's comparison)
It seems like more a backup if you need to work offline, imo, unless time doesn't matter and/or your electricity is free. Or you want independence from the labs (fair enough).
This.
Or does a "collaborating work environment" mean that everything is basically spoonfed to them? Or do you only ever use ghost suggestions?
I genuinely cannot even fathom. Just how do you even get into a state where tasks are so clear and cookie cutter? These things are abhorrent. Not only are they not useful, it's an outright form of psychological torture to try and use them. They almost fight you.
Luna doesn't even respond to steers properly! You try steering it and it immediately gets distracted and then just stops.
I can imagine coercing Sonnet into doing some of my tasks okay, but Haiku? Especially 4.5? Really?
It's just glue code
It's not complicated. Someone just has to be there to squeeze the bottle
I'm desperately trying to classify and standardize my work items and delegate them to less capable models, because my usage is clearly unsustainable and this same sentiment as above keeps being pushed on me too. But all my tasks are genuinely fairly arbitrary, so there's no real way around the agent actually being able to reason about business and technical context proper. It's not even that they're hard, it's just that they're dynamic.
I can get Luna to do things like walk our observability stack and perform a healthcheck, then defer to a stronger model if anything looks super off, but if I'm being entirely honest, this could basically be just a script. Which Opus 5.5 will immediately write for itself if it doesn't yet exist, run that, and then off it goes depending. But Luna will never actually do an investigation proper. Heck, it can't even read our dashboards most of the time, tripping up on Grafana minutia.
It feels like that surgeon vs surgeon comparison, where you're made to decide based on their surgery success rate, and the better succeeding surgeon simply reward hacks the number by only operating on less dicey cases. Except there's no objective way to make this classification here, so jackasses like the above get to play with my insecurities with full obnoxious confidence, while I'm left desperately trying to slim my usage and failing to do so between two moments of crippling self doubt and blockers.
I wouldn't even tried it, i would still just go with even Opus (we don't have that many alerts) but it really surpsied me.
When i ran into usage limits a few days ago i switched most to Sonnet and again was surprised how good it is now.
Optimally? Opus will pay for itself if you save just 10% of your time
So be less snarky?