One real-world constraint is that end-users won't always do what you intend for them to with the thing you engineered.
6,701 karma · joined September 9, 2017
One real-world constraint is that end-users won't always do what you intend for them to with the thing you engineered.
Those are tactical objectives, not strategic aims. The US is very good at winning tactically, but losing strategically. This is yet another example.
They lost the plot long ago. They're firmly in extraction mode now: how much value can they get from end-users?
No job will ever tell you you've done enough and so you should do something else. If you don't make that call yourself, whose life are you actually living?
>And why not take the alternative approach of identifying the subset of people who have indeed found solid uses and spread their best practices around?
A bottom-up approach has a far better chance of finding those particularly good use cases, and if you lean on the people how found those fits, they're more persuasive than top-down edicts. They actually know what they're talking about. If the point is to leverage AI for better work outcomes, someone with your experience is far more valuable than "here's a dashboard, make the number go up," which seems to be what's going on at Amazon.
The top-down approach to encouraging (mandating?) AI usage strikes me as infantilizing to the workers, who are perfectly capable of choosing which tools they use and when.
The sustainable alternative I've arrived at is private group chats with notifications disabled. Engagement is entirely on my own terms (i.e., I have to manually open the thing to see any updates), and there's no formalized moderation or amplification because it's human-scale enough to let shared social norms govern it.
It's a possibility, but it doesn't eliminate the possibility that it's hype. If these claims were indeed serious, they would submit it for independent analysis somewhere.
This isn't some crazy process. Defense contractors are required to submit their systems (secret sauce and all) for operational test and evaluation before they're fielded.
It's the tyranny of the marginal user. How I wish YouTube (and generally other platforms for user-generated content) would have fine-grained search and filtering controls that let me specify exactly what I want, no recommendation algorithm trying to guess what I actually meant. But such a feature won't attract and retain the least interested people, so we'll never get it.
An automated fine is the least painful way to enforce that.
This is magical thinking. "Presence" and "time cost" are inextricably linked. You can't have one without the other.
When you use AI to decouple them, you're telling your audience/colleagues/attend that you want them to listen to you but not the other way around.
I did exactly what I said I did. I'm using these systems the way they're designed and advertised. I'm following the happy path with tasks that are small, trivial, and easy to check. This is the charitable approach. Yet the system creaks under the lightest load. If Google wants to put on a better show with stronger models, then they should make those the default.
You don't need to make excuses for shoddy engineering from multi-billion dollar corporations. And you're quite welcome to run the same prompt on ChatGPT and evaluate it on your own time.
#8 has an incorrect answer (3 appearances according to Gemini, 2 according to reality https://en.wikipedia.org/wiki/Bowl_championship_series#BCS_a...)
So it works well 95% of the time for literally a trivial use case. Imagine if any other tech tool had that kind of reliability: `ls` displays 95% of your files, your phone successfully sends and receives 95% of text messages, or Microsoft Word saving 95% of the characters you typed in. That's just not acceptable.
If the purpose is indeed software development with review, then there's nothing stopping multi-billion dollar companies from putting friction into these sytems to direct users towards where the system is at its strongest.
I stress test commercially deployed LLMs like Gemini and Claude with trivial tasks: sports trivia, fixing recipes, explaining board game rules, etc. It works well like 95% of the time. That's fine for inconsequential things. But you'd have to be deeply irresponsible to accept that kind of error rate on things that actually matter.
The most intellectually honest way to evaluate these things is how they behave now on real tasks. Not with some unfalsifiable appeal to the future of "oh, they'll fix it."
Pedestrians regularly wave acknowledgement or even say "thank you." Some other cyclists (especially on e-bikes) just blast by with no warning.
Engineers should be honest that everything is a tradeoff. For the up-front convenience you get with phone tickets, you impose additional failure modes, dependency chains, and accessibility issues that simply weren't a problem with paper ticketing.
The "phone-ification" of everything will probably bite us in the behind in the future, just like the buildout of out car-centric environments does now.