704 karma · joined May 11, 2018
I would try to solve this by making the market structure reflect the underlying difficulty: we have to decide what capacity to produce years in advance, to construct the memory fabs. So this should be a futures market, and a capacity crunch would affect short-term-futures, but leave full term futures at the same price. Because the companies supplying the memory can just construct more capacity to fill those futures at the same cost regardless of the AI demand.
1. Your service that retries should have some retry budget. This is a good place to be "smart", because you can reason entirely locally instead of turning it into a distributed systems problem. The best library I've seen for this was doing Exponential Moving Average of requests per second sent down that pipe (not counting retries) and only allowing 20% more requests per second as retries, total. Each individual request could be retried 3 times. This was critical as it bounds the additional load from retries.
2. Whenever a service retries but has to give up, the error it sends to its callers should never be retried. There has to be some agreement that that HTTP code will never be retried. This prevents the multiplicative factor of retry on top of retry, which is why those storms can generate so much load.
Everything else is nice-to-have, but those two alone should bound the total requests you get in a retry storm.
It used to be that when someone else at your company was asking for something that wasn't a priority, you would erect bureaucratic roadblocks to protect your time. Now, the new normal is to just forward their questions to AI and sling the slop back over to them.
I'm not really familiar with this stuff, but the example uses what it calls a "pin" (which in their docs is a type of "binding") on ECX before calling CPUID.
You and I aren't in those business channels, and we're not being given anything with a hash. There's simply no hash to collide with?
FWIW both the company I was leaving and the company I was joining were startups selling dynamic seat pricing systems to airlines. Your call if that's a conflict of interest :).
> Opus reads base.py, signing.py, the test file (twenty cards of gray), then writes its plan and leaves. And what's the first thing Flash does with that beautiful document? It re-reads base.py and the test file, because a plan is not a file and you cannot edit prose. The gray reads just keep stacking, first at Opus prices, then again at Flash prices. There is no version of this where a second reader is the cost optimization.
The author is probably only evaluating their own harness with different models, so Claude-Code-specifics like memory aren't part of their evaluation.