2,191 karma · joined June 12, 2012
You can contact me at sethev@gmail.com
I could make all sorts of claims on the spot here. It doesn't create a duty for people reading this thread to go investigate them.
They acquired Cerner, which had ~30k employees.
At the time "generating a program from a spec" was an idea floating around that you could come up with a "spec language" that was easier than regular programming languages but somehow still had the same power and could be compiled directly into a program. That's the crackpot idea that Joel is referencing - but that's not what a spec language used with an LLM is doing.
[1]: https://www.joelonsoftware.com/2000/10/02/painless-functiona...
A bell curve tracks the distribution of a single random variable. You're mixing statistical metaphors.
That's describing something that's not a world war, though. The Russian invasion of Ukraine is already far worse than what you're describing as WW3. (setting aside nuclear escalation)
All of this to say. I suspect a lot 10k person companies made up of white collar workers could significantly cut their staff and still survive. By the time you get to that size, there's a large middle management that is constantly looking for reasons to increase their 30 person org to 40, and who will be overbooked whether they have 20 people or 100.
The registry is just a big list of names and addresses.
- The accuracy of detecting a mass
- The true distribution of masses in the population
- The likelihood that of falsely detecting a mass in the same place twice (you seem to implicitly assume that false detections are uncorrelated with each other)
- The likelihood that a real mass is cancerous (you stipulate that this is 95% in your scenario, but you don't say what other factors are used to determine this - as opposed to just knowing that there's a mass that grew.)
- The positive effect of treatment in the case of true-positives.
- The negative effect of treatment or further diagnostics in the case of false-positive.
Saying that doctors are lying about over diagnosis to cope with the fact that diagnostic techniques are too expensive is absurd. They have to actually make decisions in the real world, where your two neat little categories can't be known even if they hypothetically exist.Forget about agents or AI: the amount of money that it makes sense to spend on software engineering for a particular company is highly dependent on the specifics of that company.
Perhaps for them this number makes sense, but it's kind of crazy to extrapolate that to everyone as some kind of benchmark. It would be far more interesting to hear how they place a value on the code produced.
I have a harsher take down-thread, but the simulation testing (what they call DTU) is actually interesting and a useful insight into grounding agent behavior.
Setting aside the absurdity of using dollars per day spent on tokens as the new lines of code per day, have they not heard of mocks or simulation testing? These are long proven techniques, but they appear bent on taking credit for some kind revolutionary discovery by recasting these standard techniques as a Digital Twin Universe.
One positive(?) thing I'll say is that this fits well with my experience of people who like to talk about software factories (or digital factories), but at least they're up front about the massive cost of this type of approach - whereas "digital factories" are typically cast as a miracle cure that will reduce costs dramatically somehow (once it's eventually done correctly, of course).
Hard pass.
These examples are closer to control loops, where a decision is made and then carried out or finalized later. This kind of "eventual consistency" is pervasive but also significantly easier to reason about than what people usually mean by that term when talking about a distributed database, for example.
To expand on the 24/7 grocery store example: if the database with prices is consistent, you will always know what the current price is supposed to be. If the database is eventually consistent, you may get inconsistent answers about the current price that have to be resolved in the code somehow. That's way harder to reason about then "the price changed, but the tag hasn't been hung yet". The first case, professional software engineers struggle to deal with correctly. The second case, anyone can understand.
For example there's a class of join algorithms called 'worst-case optimal' - which is not a great name, but basically means that these algorithms run in time proportional to the worst-case output size. These algorithms ditch the two at a time approach that databases typically use and joins multiple relations at the same time.
One example is the leapfrog trie join which was part of the LogicBlox database.
(map v [4 5 7])
Would return you a list of the items at index 4, 5, and 7 in the vector v.