LWM – Open LLM with 1M Tokens Context Window
github.com
github.com
World model on million-length video and language with RingAttention - https://news.ycombinator.com/item?id=39367141 - Feb 2024 (58 comments)
> We trained our models using TPUv4-1024, which is approximately equivalent to 450 A100s
> Inference for such long sequences requires a minimum of v4-128
So you'll need ~60 A100 for inference.
> We additionally scale our inference code to support million-length sequences by implementing RingAttention for decoding. Inference for such long sequences requires a minimum of v4-128 with a TPU mesh sharding of 32 tensor parallelism, and 4 sequence parallelism (ring dimension).
Each TPU-v4 has 32 GiB of HBM memory, so about 4 TiB (128 x 32GiB) of memory in fp16, without quantization.
Brute forcing a quadratic complexity problem seems like wastefulness at its worst.
https://arxiv.org/abs/2310.01889 (Submitted on 3 Oct 2023)
Any views on the license? Github says Apache 2 for weights...but hugging face says llama license