3,002 karma · joined June 26, 2023
The world has standardised methods of accounting. Not only do Anthropic and OpenAI avoid using those methods, they both use the same phrase “annualised revenues” to describe two radically different accounting processes.
They’re both also leaking those annualised numbers slowly to the press at irregular intervals, which hints that they’re disclosing new numbers in the days after a big sale lands. So you see “$30bn annualised” because they managed to land a $1bn contract the week before, bumping the annualised figure up by $12bn compared to the start of the previous month, and the end of the next.
>What the test measures: A model is given a passage and a fixed set of questions with short, checkable answers — a date, a name, a count.
So, a model is given content which is especially amenable to compression, and asked to reproduce it under certain constraints, like...
>Why isn’t the plaintext baseline 100%? Answering questions about an uncompressed passage in plaintext scores ~91%.... a correct answer worded differently scores as a [failure]
Models can (and do) give objectively correct answers, but are penalised for not having some kind of omniscient knowledge of the implementer's phrasing preferences.
If this phenomenon is emergent in models, this benchmark is not proof of it in any meaningful way.
I think the drivers have some extraordinary capability for trust, because the chain of people who need to have got their job correct for that brake pedal to work in that moment is in the dozens, and the consequence for any of them getting it wrong is a particularly violent crash.
This isn't really the case anymore. There's a budget cap, and all of the teams are financially sound most of the time. While the poorer teams might bring fewer upgrades in a year, they still bring upgrades all year long, and R&D is a continuous process in every F1 factory these days. Smaller teams also get additional R&D, compute, and testing time compared to the big teams.
Most of the real low hanging fruit was picked up by humans years ago. When doing automated scanning, the majority of stuff is overly-verbose nonsense which takes hours of expert human labour to understand, test, and discard.
Reading through a Claude generated false positive is absolutely excruciating, because it is absolutely determined that what it’s found is justified. Often you’ll receive very long accompanying “proof of concept” code which demonstrates absolutely wild scenarios. It’s especially frustrating when you’re volunteering your time for a project, and a well-meaning contributor submits the report without the technical nous to understand why you’re rejecting it.
I see absolutely no distinction between the two, aside from minor technical approaches to gathering the content.
Lots of hires, especially in places big enough to have dedicated teams for platform work, are for empire building, rather than to support business needs. In these situations, companies hire engineers because a higher-level manager needs to increase their headcount, in order to pad out their CV for their next move.
Without introducing an unpopular tax regime, the French govt has found a way to progressively increase prices for higher mileage motorists. Supplying most motorists at a reasonable price, while having high-mileage motorists pay the unsubsidised price for most of their petrol.
Especially expensive when you take into account the amount of that code which must have been boilerplate & meta-code in nature, meaning it should have been straightforward to move.
AI abstracts effort and cognitive load away from code at a heafty rate, but it doesn’t abstract liability away from code at all.
My business is paid to produce artefacts for which it has liability in the case of error, so we need to do additional work to mitigate and eliminate the liability risk introduced with language models. So far I’ve not found a better way to do that than a plan/act/assert type approach on every feature.
They've not got a pressing need for the cash, so can afford to take a multi-decade view of SpaceX. If you look at other investors who take a view across that time horizon, you'll see some pension funds and huge trusts putting money into SpaceX for the same reason.
It's overpriced today on fundamentals, but in 50 years time, today's price will look cheap.
This isn't such an unexpected outcome. Junior & senior dev roles rarely prepare staff for the type of meta responsibility that's involved in managerial work, and there's almost no exposure to the majority of your manager's duties. It's typical that the only exposure a non-manager gets to their manager's role is a 20 daily minute standup, some tickets shuffling around a Jira board, and a fortnightly 30-45 minute catch up. Then when your mastery of .NET/python/K8s/whatever is sufficiently well established, the business takes that as evidence that you'll be an adequate manager of humans, and promotes you with little additional training, and you often stop doing hands-on dev work.
Meanwhile, if you're a junior staff member at a supermarket, you're doing managerial-type work daily, and throughout your time, you get exposure to low-stakes managerial responsibility. You work alongside your manager for 9-12 hours/day, and you see pretty much everything they do. Over time, you're given more and more of the meta-managerial work, and when you finally get promoted, it's because both you and the business already have evidence that you can do the managerial work. You then continue stacking shelves 9-12 hours per day, as a manager.
I think there's something the software industry could learn from the supermarket industry's approach. They have a much smoother gradient between non-manager and manager.
Blink twice if you need help
> a story
I think this is the first time I’ve knowingly laughed at a model’s joke.
So, the precursor to online media has already gone through this paradigm shift.
I downloaded the app to listen and write constructive notes for people, but I can’t enter the app until I’ve given microphone access and recorded a 30s clip. I’m on the bus, and can’t provide that audio now, and I’ll almost certainly forget about the app before I get 30 free seconds this evening. I think you’ll drop a lot of first time users here.
I do a bunch of work in remote areas with low-ping/low-bandwidth networks. The “hold on a skimmer while we round trip to Google” dark pattern makes the service unusable for me when I’m on-location.
Vulcan Materials Company has produced some of the strongest and most resilient economic gains for investors at around 25% gross margins for 50 years.
Vulcan’s business is crushing rocks, and then driving those rocks to where people need crushed rocks.
I mention it because Vulcan isn’t a sophisticated business at its core, but its economic returns are exceptional, because they do their unsophisticated work exceptionally well, and exceptionally efficiently.
Across the economy, most economic gains are created by companies like Vulcan, who do boring repetitive work exceptionally well and efficiently.
I expect this to hold true into the era of AI. Most stuff probably doesn’t need an exceptional model, and paying for an exceptional model to do unsophisticated work will leave your business vulnerable to competitors who take time to find the most efficient model for the task, and undercut you.
I subscribe to YouTube and my partner subscribes to Prime. We have a joint cable-like tv package too, and one of us pays for Disney Plus. In the past we’d have been consolidated into one cable subscription consumption, but now we appear in the stats as lots of distinct consumptions.
I think all-in, streamers plus other subscription based entertainment services are nearer to the $140 figure per household than the individual figures describe.
I'd assume from the rest of your post that you'd all become olympians, astronauts, high ranked politicians, or founders of mega-corps. These are all fairly run-of-the-mill careers your family entered into, and plenty of people (I'd guess close to 100%) enter these roles without the type of abuse you suffered.
Hopefully the mortal disappointment you project onto your kids is a little lighter than that your parents projected onto you.