67 karma · joined April 27, 2017
For fast inference you typically keep all experts in memory (or shard them), so VRAM still scales with the total number of experts.
Practically, that’s why home setups are wasteful: you buy a GPU for its VRAM capacity, but MoE only activates a fraction of the compute each token, and some experts/devices sit idle (because you are the only one using the model).
MoE does not make batching more efficient, but it demands larger batches to maximize compute utilization and to amortize routing. Dense models batch trivially (same weights every token). MoE batches well once the batch is large enough so each expert has work. So the point isn’t that MoE makes batching better, but that MoE needs bigger batches to reach its best utilization.
If you try to run GPT4 at home, you'll still need enough VRAM to load the entire model, which means you'll need several H100s (each one costs like $40k). But you will be under-utilizing those cards by a huge amount for personal use.
It's a bit like saying "How come Apple can make iphones for billions of people but I can't even build a single one in my garage"
Dropping the novels into a machine‑learning corpus is a fundamentally different act. The text is not being resold, and the resulting model is not advertised as “official Harry Potter.” The books are just statistical nutrition. One ingredient among millions. Much like a human writer who reads widely before producing new work. No consumer is choosing between “Rowling’s novel” and “the tokens her novel contributed to an LLM,” so there’s no comparable displacement of demand.
In economic terms, the merch market is rivalrous and zero‑sum; the training market is non‑rivalrous and produces no direct substitute good. That asymmetry is why copyright doctrine (and fair‑use case law) treats toy knock‑offs and corpus building very differently.
I've seen this firsthand multiple times: people who really don't want it to work will (unconsciously or not) sabotage themselves by writing vague prompts or withholding context/tips they'd naturally give a human colleague.
Then when the LLM inevitably fails, they get their "gotcha!" moment.
All I need now is a tool that will generate the YC application
He spends a lot of his money on giving back: 42, a free programming bootcamp (https://www.42.fr/), Station F, a huge co-working space in Paris which hosts entrepreneurs for free (https://stationf.co/)
He also has a very active VC/Business Angel fund (he funds 2 early stage startups every week out of his own pocket). His telecom company (free) helped divide by 2 the average price of mobile plans in France.
All in all, his impact seems very positive, I'd love to understand why you think he is evil