In my own company, the last 2 months, I feel like we've made more feature announcements to our internal staff than they can handle. AI has genuinely made us that much faster.
In the past, making one of these announcements every month and the company celebrated. Now we're making making multiple each week.
Not only that, we're making so many tiny improvements and bug fixes that improve the experience but we don't even bother to make those announcements anymore. They don't feel "grand" enough anymore. The goal post has shifted a lot in the last 6 months.
You could've stopped with just the first sentence and I've would learned just as much as I did reading that comment to the end.
Our own Github commits volume seem to follow quite closely with token usage on Open Router:
Do you also have a graph for more useful metrics
Yes. I wrote about it in my original post. In my own company, the last 2 months, I feel like we've made more feature announcements to our internal staff than they can handle. AI has genuinely made us that much faster.
In the past, making one of these announcements every month and the company celebrated. Now we're making making multiple each week.
But I suspect you want our project management pipeline? Maybe I can just ask our coding agent to search and summarize all the features and fixes for you and then build a dashboard for you. Better yet, my email is in my profile. Email me, we'll get on a call, and I'll show you. /sLet me ask you. What are you doing such that your velocity hasn't been greatly accelerated in the last 6 months? Can you prove that it hasn't been accelerated with facts?
Now my question is, if your internal staff are not requesting these features and are struggling to adapt fast enough, how useful are they
Some are internal staff requested, some are customer requested, some are PM requested.have you considered the consequences of that or are you still drunk and thinking that this is a good thing?
And your solution to staff being unable to handle the number of feature announcements be like..?
What exactly would AI have to do in order to not be called a bubble?
As things mature there will be a correction, ie the bubble will pop.
I would recommend to read « Boom and Bust: a global history of financial bubbles » https://pure.qub.ac.uk/en/publications/boom-and-bust-a-globa...
And today these data centers are fully utilized. OpenAI tweeted today that they may need to disable new signups for the Pro subscription in the near future due to capacity constraints.
A year from now, who knows what the situation is going to be like. It seems quite possible that robotics, self driving, research, etc. drive even more demand and revenue.
Stating with any certainty that allocating capital to build infrastructure is a mistake and that there is a correction coming seems unserious.
That cuts both ways, we are building datacenters for an immature technology that is quickly evolving. We have no idea what AI will look like in the next 5-10y. Everything that is planned to be built is based on the demand we see right now, not what it will be in the future. That means different GPUs that require different cooling systems, different power supplies, etc. NVIDIA already broke backward compatibility with their new cards, which requires a different infrastructure.
What is unserious is the opposite position: believing that we already know what will be valuable in the future and bet the entire economy on it, without any proof of positive ROI.
It's a reasonable assumption that data centers that are set up for large power usage and cooling will be valuable.
Claiming the opposite based on, well, nothing at all, in order to forecast a correction, seems less reasonable.
What does the depreciation curve look like for nvidia cards purchased today? How long will it take to recoup the investment on this buildout? Will those datacenters pay for themselves before they're scrapped?
Let's say we serve a Fable class model on 8x B300.
From Kimi K3 metrics, with 8x concurrent streams, we would achieve 55-60 tok/s per stream, matching Fable 5.1 throughput.
432 tok/s x 3600 => 1.555M output tokens/h x 50$/M API price = $77.76 revenue per hour.
Assuming total API billing at 2.06x output token bill = $160.2 / hour or ~$20 per B300.
A server with 8x B300 could be $461.5k.
At an obviously unrealistic 100% utilization we would look at 4 months of revenue to match the cost of the server.
About how model serving works at scale and actual utilization I know little.
And for all we know Anthropic could serve their model with 64 streams on the same hardware instead of the 8 we assumed here.
Just because the bond markets 1/2/3/5-year bonds are bloated does that mean the bubble will pop. It just increases the risk.
The current prices the largest players set for their models are not profitable, they bleed money. Eventually they will "fix" it. It could end up making their services less affordable and it could cascade other businesses and services that are dependent on them go out of business
How are the open-weight Chinese models staying ~6-12 months behind on widely distributed / commodified hardware, and serving for even lower prices?
If you want check a example company from the .com days check cisco, their stock peaked at 75 then crashed hard and only managed hit that again thanks for the AI bubble.
That being said I think LLMs are impressive, still.