1: "Oh, well, this is not a mundane detail, Michael!" https://www.youtube.com/watch?v=3fGHaVn5rGo
1: "Oh, well, this is not a mundane detail, Michael!" https://www.youtube.com/watch?v=3fGHaVn5rGo
There's another service that says "ok we take the 100 bytes from A, and we take the $17 SKU from B, and this should equal $X".
It's the third service that multiplies these things that failed. Where are the tests for that?
I’m totally guessing though.
In a computer system, dont those categories cover pretty much everything except a meteor strike?
Indeed.
>> To me this sounds like a human-or-LLM-driven error.
In fact, half of what’s implicated in that very sentence is not an LLM.
If you support multicurrency billing, then have the tests bill a test customer in XTS (ISO 4217 code for test currency).
Nobody is pricing their services in test currency
They are either pricing in $ or in "units" and then converting upstream
This is an even bigger foot gun than what's happening here
So how is it a "foot gun" to add XTS as an additional currency, for internal use only?
I presume AWS sets its prices in USD, and then converts them to the other currencies using the relevant exchange rate - the services themselves don't know about UAE dirham, but the billing system does. So XTS just becomes another exchange rate. You could even fix it as 1 XTS = 1 USD, although choosing a different exchange rate than parity is likely to surface more bugs.
Of course. I'm not saying that it doesn't.
Read it again
I'm saying, the conversion to (any) currency needs to be done downstream of the service (to a general billing service). The service needs to bill "credits".
> presume AWS sets its prices in USD, and then converts them to the other currencies using the relevant exchange rate
Yes that's what I wrote in the first post
When talking about e2e tests, it matters where the e's are. If they are at the public interface of "the service", then, indeed, you'd test against "credits".
But ideally, there'd be e2e tests that test the consumers' experience, where the 'e's are the web interface, emails, CLI and other things consumers click on, read, enter etc. This is where a "test currency" makes sense.
You can have your billing service billing in the test currency for e2e (and probably with a value that does not match the USD) so you can test "ok this should cost 1 credit so in the end I should have 18 $currency"
Now I’m picturing the 3am end of quarter fire drill in finance when they discover the company has accounts receivable of fourteen billion $XTS and it’s appearing in the quarterlies.
Unfortunately "automatic generation" of unit tests have made a big part of the value of unit tests ("I am telling you what I expect here") disappear.
You don't have a non-production environment where you can have billing that doesn't involve real money?
Like maybe if the bill amounts increase by like 10M% there should be someone that looks into it
You'd be better off letting the heart attacks happen and take the 3am on-call and be the hero instead. It would be good promo doc material, and being a hero is extremely good insurance against getting kicked out of the country (via the PIP->H1B grace period expiry mechanism).
The point was not that their opinion is suspect, the point was that they are former because people who care about the customer get fired and/or that everyone who cared is former, so nobody who is left cares.
Or in my case, actively ignore any and all recruiting from that sesspool.
Every once in a while (~1:12) you get one that sounds like a Mormon missionary praising how its not that way and AWS is perfect.
The strange thing is that I only come across those missionaries online and never in person (although I did go to a AWS event once [never again] and met a bunch of them, so seemingly they are actually real, to my surprise).
The other factor to add here is that, with some exceptions, the whole company feels like a Rube Goldberg machine and very few people care about what happens outside their cog (because they’re not incentivized to do so).
I got put on Focus because my "contributions were not coming through" to leadership.
There is literally no fucking reason to try to improve your skill. Any IDIOT with AI will do an OK job.
And no one is shooting for better than OK.
If it takes no skill or intelligence then nobody can get left behind because it's very fast to get back in.
In fact, there are regular AWS-wide meetings where L10 technical staff will randomly pick and review reports from across the organization. Getting picked for one of these is not a fun experience.
COEs are such a huge annoyance for teams that they create a strong incentive to be proactive in preventing issues like this from happening. One of the rules when it comes to writing COEs is that they are not the fault of individuals but processes; but in reality, no one wants to be the cause of one.
Depending on if you're a cost cutting team, fixed expense team or organization, if you're a revenue driving team, or if you're a core team, or the very many other splits you can come up about the relationship between the expense/balance sheets and the team itself...there are very very different attitudes towards COEs and leadership principles.
There's just extreme variation in the quality of the management, the quality of the engineers, the operational/development role split, the on-call schedule, and the development and testing methodology.
If you find a problem like this thread’s hypothetical, the process stops being an annoyance just to line level managers, and something that directors and vice presidents need to handle by changing strategic priorities within their organizations.
That entails a real loss of face for them, and because they are the ones who actually run the show, it would will only happen if you have one that is naïve or a masochist. In either case that moves them out of management.
From the perspective of people you manage, it's a very different picture.
We (I say "we", because I was an IC) sit under you, and every year at performance review time you're effectively required to put some percentage of us in the "LE" bucket. Never mind that we could theoretically all HV3+ if you measure by "normal" peoples' standards, your manager isn't going to let you mark all of us as HV3 at the performance meetings. I know this, because I've been there as well at those meetings where truly high performing people were downrated to fit a distribution.
So what happens? When I see a peer's critical lurking bug, I have no incentive to fix it for the sake of prevention. If I fix it quietly so that it never surfaces, it looks like I haven't done any work for the week, or have done un-impactful work, and I get the stick from you. Preventing fires doesn't look like work, to non-technical eyes.
The only "safe" way to play this game of "survivor" is to let that bug surface eventually, then when the SEV comes up, I jump in and fix it, earning your approval, skip approval, VP approval, as well as potentially the other person gets the stick from you, because you have to give the stick to someone anyway, you get a reason to stick it on them. At least it's not getting stuck on me.
I'm sorry if this comes off as shocking to you, but it really shouldn't; the incentive structure is NOT set up for teamwork, plain and simple. If "putting customers first" is a value, then it absolutely needs to start from systematic changes of how people are managed.
> this is not the case
> [describes all sticks and no carrots]
Absolutely not my experience at AWS.
All the teams I was on treated them as "not a big deal", kind of a non-punitive exercise in technical writing, and the COE was always assigned to be written by an engineer who was not involved in causing the COE.
Also, the kinds of issues that did or didn't lead to COEs appeared to be largely random. I was considered to be an extremely good operational trouble-shooter on the team where I spent most of my time at AWS, and I was never able to predict what an L7-8 manager would decide was COE-worthy.
I spent the slight majority of my time at AWS in RDS.
And this surely works out great and completely prevents any issue, and this "inaccurate estimated billing data" issue this submission is about is both the first and last time this issue ever happens at Amazon?
Ah yes, the good old ITism "Everything's good, what are we even paying you for?" followed by "Everything's on fire, what are we even paying you for?"
I moved out of it largely for that reason, am now an infrastructure/IT project manager, quite refreshing actually.
I learned about <https://en.wikipedia.org/wiki/2011_T%C5%8Dhoku_earthquake_an...> from alarms like this, as sales in Japan almost entirely stopped.
I've been told a tale of another incident where some customer ran some huge cpu-intensive workload that didn't do any networking. It caused various alarms to fire because it "looked like" a part of the network was idle (potentially indicating some sort of networking failure)
It's generally (in the broad sense) easy to add alarms for things going wrong, but in my experience anomaly detectors are just as likely to fire from other weird things like that happening.
Well you’re half right. Either there wasn’t monitoring for this or if there was monitoring in place that’s not what caught this, because the page originated from a customer support ticket.
Earlier this week my slopservant implemented several comprehensive changes to a codebase.
It also wrote extensive tests to verify the correctness of the changes.
A few days later I was working on something else and realized, everything had been implemented backwards, in a way that was nonsensical and also completely pointless.
The many tests it had written were just confirming the LLM's idea of correctness, which turned out to be... completely incorrect.
I laughed when I realized, if I had been using Rust, or indeed, formal verification, that wouldn't have helped at all — it would have just written a mathematical proof, proving the correctness of the wrong thing!
Not sure what lesson to take from that (except read the damn diffs, obviously — it was a hobby project okay ;), but it seems like the more reliable this stuff gets, the more we expect it to work properly, the more risky it becomes.
The more features I add now will likely make the inevitable refactor much harder, but adding new features is so easy that I want want extend this illusion of productivity just a little bit longer!
The argument in favour of "read each delta" (aside from catching them doing stupid shit) is to keep your mental model synced. But I already didn't understand my own code before AI, so that's moot too!
If you write the tests the agents have an easier time writing the code!
Then it can write tests based on the interface, and that is usually a lot more solid.
With a skill like this, which I have written, the LLM will end a session with "wrote 24 tests, 5 failing due to API contract mismatch" or something, and leave it at that. I can then decide if that's a bad interface or bad implementation.
This is a well-known issue in formal methods, often referred to as the "formalization gap". It's pretty much an intractable problem, i.e. all we can do is manage it, not solve it.
In fact this issue exists in many fields, often under the name "specification gap" or "specification problem". You'll find versions of that discussed inv economics, architecture (of buildings, not software), etc. In architecture and construction, a high percentage of mistakes that result in lawsuits turn out to be due not to construction, but to specifications that were incorrect in some way.
The bottom line is that there's no perfect formal/mathematical way to go from intent to proof.
There's a similar thing with physical processes. My wife works in geointelligence ground processing and encountered this when her system was expanded to support SAR collections instead of just visual spectrum. Software that passed all of its internal tests was passing nonsense collection parameters because the developers didn't understand the difference between energy collected to a sensor cell reflected from the sun versus energy reflected from your own active scan. The capabilities and limitations aren't the same. You can perform the same processing. The bits on disk won't care. But if the process you encode doesn't accurately correspond to the physics of reality, you're producing nonsense. Memory-safe, syntactically-valid nonsense, but nonsense nonetheless.
This seems to happen quite a bit with software companies.
Maybe some rare few are, but by and large they generally stink.
They tend to give that impression because they hire tons of auditors and panjandrum, to hassle their suppliers, but internally they are winging it like the rest of us.
So while the actual "on the ground" picture may look suboptimal at any given point, overall it does make for security that in practice is much better than average.
This is reflected in the actual security compromise statistics. Your money in a bank is a lot safer than, say, your credit card deals on file with many retailers.
They transfer money to me. Their mistake: there is nothing I can do. Rules for me but not for thee?
[edit] Testing your tests, like testing your backups, is a good idea
You just have to trust your colleagues won't edit the tests in such a way that vandalizes the invariant. That's why you can never trust an LLM to edit a test. They're notorious for tweaking tests such that they pass.
[edit] This may need to be an integration test to be effective, there is a certain peril to mocking that could bite you here. But that's fine, we have the technology.
There's a lot of hypotheticals in there, so who knows if testing could have caught this specific problem. But I see what they're saying in principle.
insufficient tests that dont assert on the right things?
the existence of a test doesnt mean it catches the right thing
based on the description, id bet the COE action item will be to do a migration that enforces units are passed at the billing service level
theres no good reason for the billing service to make up its own units.