"Holds up"
"not a toy, but not a sprawling project either, and ideally.."
"And that’s the trap"
ctrl F "real" -> 6 usages
ctrl F "genuine" -> 4 usages
2,645 karma · joined August 23, 2012
"Holds up"
"not a toy, but not a sprawling project either, and ideally.."
"And that’s the trap"
ctrl F "real" -> 6 usages
ctrl F "genuine" -> 4 usages
As a serial successful field-hopper, I agree that I'm not the right person to be making these estimates.
But the external view is that college courses roughly expect you to do what I'm claiming, in roughly the time investment that I'm claiming -- and undergrads are typically in 4+ classes at a time. So is it that the whole educational system is delusional? (I fully acknowledge it might be so!)
I think if you want to give your reader a quick intro to, e.g., what is the Adam optimizer, a simple link to Wikipedia is fine. No need to copy-paste an AI tutorial on Adam into the blog post.
This throughput assumes 100% utilizations. A bunch of things raise the cost at scale:
- There are no on-demand GPUs at this scale. You have to rent them for multi-year contracts. So you have to lock in some number of GPUs for your maximum throughput (or some sufficiently high percentile), not your average throughput. Your peak throughput at west coast business hours is probably 2-3x higher than the throughput at tail hours (east coast morning, west coast evenings)
- GPUs are often regionally locked due to data processing issues + latency issues. Thus, it's difficult to utilize these GPUs overnight because Asia doesn't want their data sent to the US and the US doesn't want their data sent to Asia.
These two factors mean that GPU utilization comes in at 10-20%. Now, if you're a massive company that spends a lot of money on training new models, you could conceivably slot in RL inference or model training to happen in these off-peak hours, maximizing utilization.
But for those companies purely specializing in inference, I would _not_ assume that these 90% margins are real. I would guess that even when it seems "10x cheaper", you're only seeing margins of 50%.
Here's one way they could get around their own privacy policy: keep track of what % of Claude-generated code is retained in the codebase over time (as an indicator of how high-quality / bug-free the code was); A/B test variations of Claude Code to see which variations have higher retention percentages.
No usage data is retained, no code is retained, no data is used (other than a single floating point number) and yet they get to improve their product atop your usage patterns.
Here's another idea: use a summarization model to transform your session transcript into a set of bits saying "user was satisfied/dissatisfied with this conversation", "user indicated that claude was doing something dangerous", "user indicated that claude was doing something overly complicated / too simple", "user interrupted claude", "user indicated claude should remember something in CLAUDE.md", etc. etc. and then train on these auxiliary signals, without ever seeing the original code or usage data.
I think it's quite plausible that Anthropic is bleeding out ~100/month on token costs per $20/month user, and even at 80% margin, this is just merely breakeven. Their limited capacity also means that they are _losing_ the opportunity to sell the same capacity at a per-token marginal profit. I think the only plausible endgame here is that Anthropic uses the usage data to RL-finetune Claude Code to the point where it is actually worth a $200/month subscription.
Enjoy the $20/month Claude Pro plan while it lasts; I don't really see it sticking around for more than a year at best.
I think showing the raw reasoning text is not quite the right UI; maybe highlighting the specific text in red and showing a suggested correction would work better?
It's also a little awkward that the conversation is live; I don't really have any breathing room to read the reasoning traces on what mistakes I made / could have done better. I hung up the first time I tried to figure out how to pause.
Tldr: most applications of free energy have capital costs that far outweigh the free energy harvest potential.
See my source code here:
https://github.com/brilee/modern-descartes-v2/blob/master/ma...
Includes:
1. RSS feed
2. Blog listing pages ordered by date
3. Tagging system
4. Localhost dev server with file-watching recompilation step.
I happen to have recently written up a longer history of Go AI. If you're wondering about what is special about Go in particular or what generalizes to other problems, give it a read.
Another purely speculative example - Geosmin is the human-recognizable indicator of rain, and given how intensely sensitive we are to this molecule, I wouldn't be surprised if many ecosystem participants also sniff for this molecule.
Why did Brain Exist? https://www.moderndescartes.com/essays/why_brain/ Who pays you? And why? https://www.moderndescartes.com/essays/who_pays_you/
Free lunch briefly existed for a small lucky few in 2017-2021, but today there is definitely no more free lunch.
In Operation Choke Point, the regulators were actively moving the lever to debank morally repentant industries. In "Operation Choke Point 2.0", the complex system of regulatory guidelines and actors seemingly self-coordinates to debank crypto, and the failure of anyone to intercede is painted as a willful neglect. Those on the regulator side say, "I didn't do anything", and those on the crypto side say, "you have every power to do something".
Yes, I can do this. I don't understand why people are so keen on gatekeeping perfect pitch.
Very, very obviously a Chopin Waltz. The chord choices, usage of triplets/mordants, and the suspended pedal tones is characteristic. It's a mix of [Prelude Op 28 no 11](https://www.youtube.com/watch?v=si5aT6FDPZ0)'s ephemeral brilliance and [Waltz Op 34 no. 2](https://www.youtube.com/watch?v=PGdpRmL2XUc)'s lilting, moody style.
My "edition" is definitely not an [urtext](https://en.wikipedia.org/wiki/Urtext_edition). I made minor simplifications to the descending line at measure 8, following Lang-Lang's performance (which I believe is the correct decision - the arpeggiated diminished chord F-D-B-G#-F-D sounds better without any gaps, compared to what's in the scan, B-F-B-G#-F-D.), and added some phrasing where I thought it was obvious and perhaps went missing over the ages from the raw scan. The ornamentation in measure 20 was probably modified by Lang-Lang, but I think it fits the piece better than a plain mordant, so I notated it as played.
If I were to judge based purely on the music - minus all of the contextual clues like paper, ink, backstory - the probability of it being fake is ~10%. I say this only because it is shockingly similar to 34-2 in harmonic and stylistic elements, which is exactly the kind of thing an AI trained on a not-big-enough dataset would do. While AI utterly fails at longer pieces, it could plausibly render a coherent 24-measure piece in the style of Chopin, and DeepMind could plausibly be working in stealth on a really good music-composition AI. But in the end, the piece is too tightly composed, and I trust NYT's decision to trust the historians who are familiar with evaluating such artifacts.
https://www.moderndescartes.com/essays/chopin_waltz_posthumo...