4,463 karma · joined March 26, 2016
There may be an economic challenge - which seems to be the sort of problem mass manufacturing can solve very well.
There may be a compute model & latency problem, how do you organize model training when racks are much further apart than in traditional data centres (although speed of light is 50% faster in vacuum than glass fibre). But that's algorithms.
Relative to everything else in orbit, powering a rack of compute and some comms per satellite seems not really to be a physics problem.
This is interesting in that it seems to be making the grind more efficient. I think the true breakthrough will likely be proper scale quantum computing to make the first principles design feasible
For those unfamiliar with the reference, in Hitchikers guide to the galaxy series, Mice are projections of hyper intelligent extra terrestrial beings from a higher dimension who commission the creation of planet earth complete with fossils as a giant organic computer
Even in the top echelon of the Indian government (Indian Administrative Services), a similar culture exists and has interestingly been strengthend as the power has fragmented across political parties. I personally know several of my top B School/Engineering college mates who joined and any of them would do very well in top management in private industry. You can see the result in the rapidity with which India is pulling people out of poverty and modernizing. Certainly alower than China, but pretty amazing seei living outside for a decade now - many urban services are already far superior in quality to what you get in Europe, US or even Singapore. The challenge in India is the bulk of the bureaucracy below them isn't held to the same standard.In India it's still a prestige thing only in the top echelon - the lower echlons are more motivated by job security, pensions and avenues for corruption.
The difference in Singapore is the government pays well enough to attract and retain higher quality talent deeper.
This is nonsense that AI providers want to peddle. Inference is wildly gross margin profitable - likely 90%+ gross margins. It's very easy to work out the cost structures bottoms up. All providers can drop costs to a third and still keep positive gross margins.
The problems are 1. It possibly still doesn't pay out on training investment in a reasonable time frame without a massive expansion of the 90% gross margin.
2. There is no moat. As we see Mac Mini & High End GPUs stock outs and the pricing offered by DeepSeek and Qwen, the performance of Open Weight models are good enough that people can and are already shifting many inference workloads out of these 90% margin players
The first question to ask is does your use case require handling personal or sensitive data.
If you're using the LLM for OpenClaw or you want to handle sensitive or medical data, a local model generally is necessary.
if it's not so sensitive - Cloud providers with some sort of user agreement guarantee on not using your data for training would be the next bet. I personally generally use Gemini or Sonnet as my cloud backup. As I understand, OpenAI, Cloudflare (which bought replicate) and Qwen also seem to provide such guarantees and make SOTA models available. Others like DeepSeek seem to have an opt-out setting. Open router & co I avoid except for benchmarking models with public or dummy data as there is absolutely zero guarantee or ability to enforce terms on providers where your data might be sent.
Gemini and Anthropic (and OpenAI) tend to be expensive - it's very easy to run up 15 dollars a day or so bills which puts you solidly in 1 year pay out on Mac Mini territory - at this point I decided to buy. Gemini Flash Lite 3.1 is however surprisingly good value.
the next question is Mac or CUDA. If your expected use is serving LLM models for inferences, the latest large memory Macs give pretty good inference speed (better than DGX Spark) at a reasonable cost - I think there offer much better value than CUDA if the only use case is LLM inference & harnesses.
if you plan to also fine tune models, experiment with other types of ML on GPUs, do computer vision stuff etc. the development tooling on CUDA is far in advance of all other platforms.
Lastly if you choose CUDA, the question is GB10 family (DGX Spark - cluster able with 128Gb RAM et all) or dedicated GPUs workstations. What I found is practically any serious models weighs in requiring at least 96GB VRAM - Antirez's 2 bit quant of Deepseek 4 flash (my current daily driver) [2] , the Qwen 3.5 122B A10B 4-bit quant, the Qwen 3.6 27B Dense and 35B A3B 8 but quants etc. So you're well out of the consumer GPU territory into 1 or more RTX 6000 Pros or Data center grade devices. Yes you can try to hack away with multiple consumer cards or SSD streaming but it's very fiddly and you probably have better things to do with your life.
The GB10 system - which I ultimately went with - is certainly much cheaper and can be clustered through the Special NVLink cable to get 256, 384 or 512 GB setups but comes with severely constrained bandwidth. The Pro GPUs blast these out of water on performance but are expensive.
Lastly, renting a cloud GPU machine doesn't really make sense except to run already debugged fine tuning workloads. You'll probably spend at least 4 dollar an hour for sufficient capacity which if it's personal use, will mostly sit idle.
1. https://srinathh.medium.com/mid-size-local-models-are-now-co...
The history of middle East for the last 5000 years (since Sargon of Akkad) is replete with 'the king X "pacified" (the most commonly used euphemism) the people in the conquered territory'. It has never gone well for the victors under successors of king X, often within a generation or two.
In today's age when access to technology and information is such that any small sufficiently competent and motivated group can cause massive destruction, is it wise to keep creating motivated enemies and expect they will somehow never become competent or that the competent won't become motivated? It is doubly ironic given Israel's own defense industrial complex is filled with such small motivated and competent groups and the evidence of Ukraine/Russia conflict is staring in the face. This situation will blow up I fear within a generation unless Israeli society chooses different leaders.
I think there's possibly value to try fine tuning Qwen 3.5 on my OpenClaw turns log to see if performance improves. The one recent model I haven't tested yet is Nemotron 3 Super which I might bench soon.
[1] https://srinathh.medium.com/mid-size-local-models-are-now-co...
The present US President is smart enough to realise that now. In this case he let himself be misled by Israelis & their supporters that Iranians would rise up and replace their own Government. That indeed might have happened if the US had intervened when the Iranians were actually protesting some months before the present crisis. Now the US administration is looking for a way out without putting boots on the ground and Iran is looking to haggle on the price and for the US. this very business like cutting of ones losses is almost certainly the right move.
This is also why historically companies have preferred being vertically integrated to avoid having their supply chains exposed until American economists and brokerages started pushing the cult of specialization. Outside of the US, big conglomerates still operate this way.
1. Would have much lower sonic booms thanks to recent research (quite a bit of it by NASA on wing geometry) and more importantly computer simulation available now
2. The engines would be far more fuel efficient
3. The flights would be able to have better efficiency in the subsonic regime as well. Just see what winglets and the like have done to fuel economy .
I fly 14 to 18 hour routes maybe 4-5 times a year on business paying 5x the economy cost and it still sucks. Breaking the flight with a connection (IMO) sucks more. My management flies such routes every month. There is a lot of revenue headroom in that fare gap for something that flies maybe 3x-4x as fast which military aircraft already do.
What will hold back the idea is conservatism among the business managers in aircraft manufactures and incumbent airlines who will "draw lessons" from a 50 year old experiment
BTW form my benchmarking, open weigh models are good enough for many agentic tasks starting with Qwen 3.5/6 family and Deepseek v4 family, so it's likely we'll see displacement of api usage from the premium priced providers. Yes trainingis expensive, this isn't training
This. The gross margin on inference is at least 95% if not higher - several open weight models on my tiny consumer DGX Spark easily replace the 15 dollars a day I was paying in tokens for Claw usage with a dollar a day electricity. You add data centre overhead and depreciation, the theoretical net margin will trend lower but depreciation is always far more aggressive than actual product degradation. The old NVIDIA GPU on a 9 year old second hand gaming PC I bought still serves up a small Gemma 4 variant quite reasonably.