Like a Carbon tax, the money doesn't disappear. But to whom it gets distributed, that's another story...
2,169 karma · joined October 1, 2011
https://github.com/RaphaelJ/
https://datethis.app https://noisycamp.com
Like a Carbon tax, the money doesn't disappear. But to whom it gets distributed, that's another story...
Obviously Starlink can and will growth. I'm just pointing out how insane the market cap is, when compared to similar scale "connectivity" businesses.
But IIRC, in Belgium at least, these plants are also remunerated on "secondary" markets for non-productive tasks. These secondary markets are mostly for grid balancing.
Same legislation as the non-plug&play inverters.
With a e-meter, you will get compensated when you're generating a surplus (+/- €0.04/kWh last time I checked).
However, thanks to the battery, I'm self-consuming almost all that electricity, saving around €0.30/kWh.
Expect 800kWh of annual production per 1kW of panels.
Buying supports for PV is actually less economical than buying additional PV panels.
--
I bought a 1600Wc + 1.9KWh kit (Ecoflow Stream) for +/- 1300€ last summer. It took us about 2h to install (we had to setup a new plug outside), and I already saved 200€+ since July. I am expecting to save about 350€ per year.
Also, as u/jstch said, it's extremely fun to setup and generate your own power!
That being said, the total cost per kWh could well reach 20c/kWh, which is ridiculous. It's not only not competitive against renewables, but also not competitive with natural gas (CCGT are probably around 10-15c€/kWh).
It might not be about bringing more revenues but retaining market share.
On some tasks like build scripts, infra and CI stuff, I am getting a significant speedup. Maybe I am 2x faster on these tasks, when measured from start to PR.
I am working on a HPC project[1] that requires more careful architectural thinking. Trying to let the LLM do the whole task most often fail, or produce low quality code (even with top models like Opus 4.5).
What works well though is "assisted" coding. I am usually writing the interface code (e.g. headers in C++) with some help from the agent, and then let the LLM do the actual implementation of these functions/methods. Then I do final adjustments. Writing a good AGENTS.md helps a lot. I might be 30% faster on these tasks.
It seems to match what I see from the PRs I am reviewing: we are getting these slightly more often than before.
---
Instead, he could just use the references he needs in the new tree, delete/override the old tree's root node, and expect the Javascript GC to discard all the nodes that are now referenced.
Ideology
The Trump administration is basically following Kodak's strategy from the early 00s.
--
[1] https://www.iea.org/reports/world-energy-investment-2025/exe...
[2] https://www.eco-business.com/news/iea-renewables-will-be-wor...
While Mistral might not have the best LLM performances, their UX is IMO the best, or at least a tie with OpenAI's:
- I never had any UI bug, while these were common with Claude or OpenAI (e.g. a discussion disappearing, LLM crashing mid-answer, long context errors on Claude ...);
- They support most of the features I liked from OpenAI, such as libraries and projects;
- Their app is by far the fastest, thanks to their fast reply feature;
- They allow you to disable web-search.
For example, if the electricity price is -28€/MWh (like today in Germany), and your battery efficacy is 80%, you could get paid 28€/MWh charging, then only pay back 22€ discharging, generating a 6€/MWh profit.
Just charging your car when the demand is low is probably enough to drastically reduce the overall cost of the system. And this has basically no impact on the battery lifespan.
I went through the paper and I understood they made these improvements compared to "regular" MoE models:
1. Latent Multi-head Attention. If I understand correctly, they were able to do some caching on the attention computation. This one is still a little bit confusing to me;
2. New MoE architecture with one shared expert and a large number of small routed experts (256 total, but 8 in use in the same token inference). This was already used in DeepSeek v2;
3. Better load balancing of the training of experts. During training, they add bias or "bonus" value to experts that are less used, to make them more likely to be selected in the future training steps;
4. They added a few smaller transformer layers to predict not only the first next token, but a few additional tokens. Their training error/loss function then uses all these predicted tokens as input, not only the first one. This is supposed to improve the transformer capabilities in predicting sequences of tokens;
5. They are using FP8 instead of FP16 when it does not impact accuracy.
It's not clear to me which changes are the most important, but my guess would be that 4) is a critical improvement.
1), 2), 3) and 5) could explain why their model trains faster by some small factor (+/- 2x), but neither the 10x advertised boost nor how is performs greatly better than models with way more activated parameters (e.g. llama 3).
I went through the paper and I understood they made these improvements compared to "regular" MoE models:
1. Latent Multi-head Attention. If I understand correctly, they were able to do some caching on the attention computation. This one is still a little bit confusing to me;
2. new MoE architecture with one shared expert and a large number of small routed experts (256 total, but 8 in use in the same token inference). This was already used in DeepSeek v2;
3. Better load balancing of the training of experts. During training, they add bias or "bonus" value to experts that are less used to make them more likely to be selected in the future training steps;
4. They added a few smaller transformer layers to predict not only the first next token, but a few additional tokens. Their training error/loss function then uses all these predicted tokens, not only the first one. This is supposed to improve the transformer capabilities of predicting sequences of tokens. Note that they don't use this for inference, except for some latency optimisation by doing speculative execution on the 2nd token.
5. They are using FP8 instead of FP16 when it does not impact accuracy.
My guess would be that 4) is the most impactful improvement. 1), 2), 3) and 5) could explain why their model train faster, but not how is performs greatly better than models with way more activated parameters (e.g. llama 3).
My main complain is that the chat sometimes fails to correctly render some GPT-4o output (e.g. LaTeX expressions), but it's mostly fixed with a custom system prompt. It also significantly reduces the battery life of my Macbook M1, but that's expected.
A lot of these drivers debuted with karting, which while not being exactly cheap, is accessible to the most of the middle class.