570 karma · joined December 6, 2012
Though this was a 100-year prediction so we still got three and half to go!
Some observations:
* Knights are color bound
* You can mate with Knight & King (K+K is still insufficient material)
* 3 fold repetition still applies (and has a popup!)
Page 18 of the paper: > As shown in Table 1, our approach outperforms other methods for both Llama-3.1-8B-Instruct and Ministral-7B-Instruct, achieving significantly higher average scores. We evaluate our method using 2.5-bit and 3.5-bit quantization during text generation. These non-integer bit precisions result from our strategy of splitting channels into outlier and non-outlier sets, and applying two independent instances of TurboQuant to each, allocating higher bit precision to outliers. This outlier treatment strategy is consistent with prior work [63, 51] . For example, in our 2.5-bit setup, 32 outlier channels are quantized at 3 bits, while the remaining 96 channels use 2 bits, leading to an effective bit precision of (32 ×3 + 96×2)/128 = 2.5. For 3.5-bit quantization, a different ratio of outliers and regular channels leads to a higher effective bit precision. Despite using fewer bits than competing techniques, TurboQuant maintains performance comparable to unquantized models
So they find channels / indicies-of-the-vector that are important and give them more bits (3 bits) than the rest (2 bits).
>Isn't the turbo codebook the irregularly spaced centroid grid?
yes I believe so. They mention it's informed by the concentration of measure and the uncorrelated/independent vectors after the initial conditioning rotation. I feel like it was informed by PolarQuant, but that may just be how I intuit what's going on (because thinking about this in polar coordinates makes more sense in my head). IOW, I think the irregular spacing is maybe informed by TurboQuant.
However they do say, slightly to the contrary: "We find optimal scalar quantizers for random variables with Beta distributions by solving a continuous 1-dimensional k-means problem using the Max-Lloyd algorithm."
PolarQuant does live on in TurboQuant's codebooks for quantization which borrows from the hyperpolar coords
The core idea is to quantize KV cache, but do so in a way that destroys minimal information. In this case, it's similarly scores between vectors. The simplest way to do this is to change all the elements from 16bit of precision to, say, 4 bits (Scalar Quant.). These papers improve on it by realizing: almost all the energy (concentration of measure) is towards the equator of the hypersphere (normally distributed as 1/d; d=vector dimensionality). (The curse/blessing of hyper dimensionality strikes again.) So when we quantize the elements (think "latitudes", e.g. to the nearest degree) we destroy a lot of information because basically all the vectors were around the equator (so some latitudes have a lot of vectors and some have very few). The idea is to rotate the vectors away from the equator so they're more consistently distributed (to better preserve the entropy during quantization, which I guess was amitport's DRIVE idea). PolarQuant does a hyperpolar coordinate transform which superficially seems neat for preserving entropy because of this equator/polar framing (and ultimately unnecessary as shown by TurboQuant). They also realized there's a bias to the resulting vectors during similarity, so they wrote the QJL paper to fix the bias. And then the TurboQuant paper took PolarQuant + QJL, removed the hyperpolar coords, and added in some gross / highly-pragmatic extra bits for important channels (c.f. elements of the vectors) which is sort of a pathology of LLMs these days but it is what it is. Et voila, highly compressed KV Cache. If you're curious why you can randomly rotate the input, it's because all the vectors are rotated the same, so similarity works out. You could always un-rotate to get the original, but there's no need because the similarity on rotated/unrotated is the same if you compare apples to apples (with the QJL debiasing). Why was PolarQuant even published? Insu Han is solely on that paper and demanded/deserved credit/promotion, would be my guess. The blog post is chock-full of errors and confusions.
BYD and Geely have similar systems. Their ICE are around 47% thermal efficiency so like ~double what you'd expect in a pure ICE car + regen and other bonuses.
https://carnewschina.com/2025/08/02/im-motors-launches-stell...
Kind of. EREVs are what locomotives have been doing for a century (and to a lesser extent barges), which is called diesel-electric in that field. I agree the terminology is lacking, but EREVs are quite compelling (and their high market share in China supports consumer demand).
Hybrid: * ICE must run during regular operation (except for ~very short distances at ~very slow speeds) -- this increases operational costs (oil changes, economy, engine designed for torque and wide RPM range). * Complex drivetrain with wheels moved by electric motors and ICE, axles, etc. * Generally 10-40 miles of EV range
EREV: * Basically an EV with a short range, and whenever you want to charge the battery on the go (or use the waste heat from the ICE) it can use an efficient (Atkinson cycle) engine to do so. (Though american EREVs have used poorly suited engines for parts availability and enormous towing numbers) * Generally 50-200 miles of EV range * Think "EV for daily commute; ICE for road trips (and heating)"
IMO EREVs would've been a better development path than hybrids or pure EVs.[0] Immediately lower TCO in various interest rate environments via highly-flexible battery sizes, no cold or range anxiety issues, technically simple drive train and BTMS.
[0] I mean the Prius made a lot of technical strides given the battery technology/costs and familiarity the industry had with ICE at time. Tesla went full EV which is a very optimistic approach, and works well enough if you stick around the charging network, but the batteries are still expensive and heavy compared to a small ICE + tank.
I have no idea in practice. But for the thermodynamic limit of actually making a difference, any irreversible change requires heat to be generated, e.g. initializing to zero, truncating, or bitshifts with discarded information. In contrast, addition/subtraction/multiplication/bitshifts without over-/under- flow will not necessarily generate heat.
https://en.wikipedia.org/wiki/Landauer%27s_principle
PS. you can also use mass-energy equivalence to extend this to calculate the lower limit of mass for a given quantity of information. TL;DR: The internet weighs 50g https://www.youtube.com/watch?v=WaUzu-iksi8
OOC Do you use ChatGPT/Gemini via the chat interfaces, pasting context and design and then copying out?
Thanks!
I will say SVT-AV1 has had some significant ARM64 performance improvements lately (~300% YoY, with bitrate savings at a given preset[1][2], so call it a 400% increase), so for many use-cases software AV1 encoding (rather than hardware encoding) is likely the preferred modality.
The exceptions, IMO, are concurrant gaming with streaming (niche on MacOS?) and video server transcoding. However, even these exceptions are tenuous: Because Apple Silicon doesn't play x86's logical core / boost clock games, and considering the huge multi-threaded performance of M4, I think streaming with SW encoding of AV1 is quite feasible (for single streams) for both streaming and transcoding. x86 needs a dedicated AV1 encoder more-so due to the single-threaded perf hit from running a multi-threaded background workload. And the bit-rate efficiency will be much better from SW encoding.
That said, latency will suffer and I would still appreciate a HW AV1 encoder.
[0] https://en.wikipedia.org/wiki/Apple_M4 [1] https://www.phoronix.com/news/SVT-AV1-1.8-Released [2] https://www.phoronix.com/news/Intel-SVT-AV1-2.0
Pertinent to the conversation though, the BIOS is very much lacking, and supposedly software-based fan control is not implemented. That said, running the fans constantly at ~silent levels of rotation keeps the temps cool (you can even run the heatsink without a fan if you want).
[0] https://store.minisforum.com/products/minisforum-bd770i?vari... [1] https://www.cpu-monkey.com/en/cpu-amd_ryzen_9_7945hx
Also Autogen seems popular and well-ish liked https://microsoft.github.io/autogen/
LangChain definitely has the most market-/mind- share. For example, GCP has a blog post on supporting it: https://cloud.google.com/blog/products/ai-machine-learning/d...
> no more than two cards per PC
I've seen quad 4090 builds, e.g. here[0]. What do you mean no more than two cards? Yes, power is definitely an issue with multiple 4090s, though you can limit the max power using `nvidia-smi`, which IME doesn't hurt (mem-bottlenecked) inference.
[0] https://old.reddit.com/r/watercooling/comments/16ed8fu/quad_...
Perhaps the resolutions to the Repugnant Conclusion (Section 2, "Eight Ways of Dealing with the Repugnant Conclusion") can also be applied to the tyranny of the marginal user. Though to be honest, I find none of the resolutions wholly compelling.
[0] https://plato.stanford.edu/ARCHIVES/WIN2009/entries/repugnan...
Regardless, I'm excited for passkeys: passwords are already abstracted out with password managers, and now they're going a step further of not being revealable w/ passkeys. However losing the escape hatch of reading a password off my phone into an untrusted computer is quite the use-case loss.
[0] No affiliation, other than being a happy customer. https://www.wealthfront.com/cash
[1] https://www.wealthfront.com/cash-account-participant-banks
Humans are incredible explorers, self-sacrificing for the greater (perceived) good. We should consider continuing our legacy by understanding ourselves better through rigorous science. (And hopefully end the historical legacy of making thousands of other species extinct via our exploration.)
I believe the Journal of Controversial Idea (https://journalofcontroversialideas.org/) will explore some of these topics, shortly.
These take minutes to send (very stressful!) preventing a nice venmo-like experience. They also pollute the planet
Technological: * First Oblivious RAM implementation, "fog", so that transacting parties cannot be revealed * Their Rust codebase is really nice * Instant transfers with little computing power (CO2 emissions)
Ethical: * Moxie and Josh Goldbard hold no MOB, along with the employees. The Mobilecoin foundation has some awesome partners, e.g. the Long Now Foundation. * Mining is not ethical, it pollutes the planet and is just bad. The only alternative is a "pre-mine" given to an independent org, ie Mobilecoin Foundation
Legal: * The US's laws are not clear on what is allowable with privacy coins, so Mobilecoin has played it conservatively by saying US residents can't own the coins.
In summary, the critiques of Mobilecoin (in any of its incarnations, foundation, moxie, etc.) are assuming the agents involved have a financial interest in MOB being expensive -- I contend that is not the case. Please show your evidence.
PS. I am assuming good faith and honesty in statements, eg "Marlinspike notes, however, that neither he nor Signal own any MobileCoins." https://www.wired.com/story/signal-mobilecoin-payments-messa...
PPS. Some direct responses:
>Let's assume that they integrated with Bitcoin or Litecoin or some "mainstream" CC, would it still be a good idea?
No, not private. Also slow. Also pollutes planet. Monero is close on the privacy front, but takes 3 minutes to send (very stressful). It's possible a coin with the proper attributes could be made on stellar, but that raises questions towards ownership of Lumens (and pumping them) and their stellar reimplementation in Rust is likely more secure.
>Niche coin
nit: MOB has a top 15 market cap with 250m coins in distribution. Though I would hesitate to compare to other cryptocurrencies which are almost entirely scammy, polluting garbage.
If I were a HVAC company with WiFi thermostats, I would look into including miners in heating solutions.