Well actually, it might.
3,522 karma · joined May 20, 2014
Well actually, it might.
Extremely large models don't suffer as much from quantization due to its weight topology also contains encoded information, so the loss of info from any one weight is somewhat mitigated.
The 1bit quant model is at an astonishing 397GB with 95B active per MOE. This literally puts Opus 4.5 performance level into a machine a normal person could buy, and still gets usable tokens/second.
The full lossless model BF16 is clocking at 4.9TB. The model card claims the model to be between Opus 4.8 and Fable 5. Again that's astonishing as getting a machine with 7TB RAM (with context + KV cache) is still within the realm of medium size companies.
Bad things: The open source version has its vision capability removed, and the context capped at 250k . I expect someone to bolt a Kimi 2.6 vision tower to it to restore the vision capability (at less performance of course). For context, I played around with extending the context to 600k for Qwen 3.5 397b, and the context remained stable up to around 480k. It'd be interesting to see if the same can be done to Q3.8 .
Also no out of the box DSpark/DFlash support. MTP is present so we should at least get some boost in TP speed.
One line after another
Hard to read it is
Why force the LLM to use files over vector database or key-value stores, just because it's a design principal for UNIX (which is designed for human users, not LLMs.)
Rust is like the instructor demanding that you build your house up to the national code, down to the choice of nail for your floor board. Your house will be perfect, but building it is extremely difficult and high-friction. In many cases this is unnecessary.
Thank you for putting my thoughts in words I can never string together so well.
I too would just stash the letter in my ZFS dropbox and never let it see the light of day again.
The letter reads like a husband who was just served a divorce notice from his wife. The man is angry, and wants everyone to know that he is not angry, and he is very much not bothered by the whole affair even though he has misgivings since the first day of marriage.
The letter would be much more convincing if it has any technical rebuttal against the decision.
Eg. Is corp ABC's stock price going to go down by 10% in 30 days?
Orchestrate a disinformation campaign on the 29th day to tank ABC's stock price.
I assume Anthropic will continue to tune the model, so I am not too bothered by this.
Even then, I had multiple cases where files were corrupted, and once the whole array refused to be online due to corrupted metadata. I had to make ZFS to replay the journal log with undocumented commands. Sometimes it takes a few days of hair-rising recovery but I always manage to get the array back intact.
The files that are corrupted are always extremely large files (>50 GB) with many small read/writes (eg. iSCSI image files.)
It's pretty impressive how resilient ZFS is, really, given I had what likely to be the worst possible hardware combination.
Reading fast means you can take in more info per unit of time. It can be a useful ability, if tedious at times.
English is a major reason- Indians are just better at speaking English than Chinese and thrives in corporate America. Whereas mainland Chinese couldn't climb the corporate ladder and have to seek better opportunities back in China.
Find me a tech executive in a Big Tech firm who is from mainland China. You can't find one. Both Lisa Su of AMD and Jensen Huang of Nvidia are Taiwanese immigrants who grew up in US in the 70s and thus speak fluent English.
2 companies have functionally similar products, but behaves completely different. One company makes technical decisions with security as the fundamental principal, while for the other company, security is not a consideration.
It's like storing all your nuke launch codes in the same vault, right in the middle of Washington DC national mall. Things are okay, until they are not okay.
Running 167 agents in the accelerator? My gawd that would never fly at my previous company. I'd get dragged out in front of a bunch of senior principals/distinguished and drawn and quartered.
And 300k manual interventions per year? If that happened on the monitoring side , many people (including me) would have gotten fired. Our deployment process might be hack-ish, but none of it involved a dedicated 'digital escort' team.
I too have gotten laid off recently from said company after similar situation. Just take a breath, relax, and realize that there's life outside. Go learn some new LLM/AI stuff. The stuff from the last few months are incredible.
We are all going to lose our jobs to LLM soon anyway.
- Local LLM, with a powerful debugger as its oracle, is now powerful enough to run rudimentary malware analysis without consulting with external sources.
- More complex malwares are still beyond what local LLMs can handle. The local LLM can see all the behaviors by the malware, but the LLM fails to put the analysis together to deduce the true intention of a binary.
- Local LLM is a very lost-cost way to do malware analysis (about 5 US cents of electricity.)
- The biggest killer-app feature is having the LLM writes its analysis back to Ghidra. The more you interact with the LLM, the more data it will write back to Ghidra. This could potentially saves hours per manual debugging by skipping function/resources/variables labeling.
OpenClaw vs NemoClaw (NVIDIA)
Developer Peter Steinberger (individual project) vs NVIDIA Corporation
Current Status Acquired by OpenAI (Feb 2026) vs Upcoming release (GTC 2026)
Target Market General-purpose consumer AI assistant vs Enterprise AI agent platform
Core Strength Rapid deployment, viral adoption vs Security, privacy, enterprise reliability
Ecosystem Community-driven (NanoClaw variants) vs NVIDIA NeMo & NIM integration
Governance Transitioning to foundation management vs NVIDIA-backed with open-source access
GPU Acceleration Not natively optimized vs Native NVIDIA GPU acceleration
NemoClaw is not even out yet, so who knows what it might look like. I guess if you sprinkle the word 'Nvidia' around enough, your product is automatically better than the rest.I don't even like OpenClaw, but this is just silly.
https://nemoclaw.bot/claw-ecosystem-overview.html
NemoClaw NVIDIA (Python/NeMo) Enterprise-grade Security Platform Built-in privacy tools, compliance, GPU acceleration Enterprise server architectureInsurance and overhead (eg. safety harass) exist for a reason other than to drain your wallet. Roofing is also a physically difficult job. You won't find many 50+ year old to couch on a rooftop all day, regardless of pays.
The support is going to suppliers, who are the true victim, but it's privatize the gain, socialize the cost. JFR screwed up, so they should be the first to step up to assist the suppliers.