- 180B parameters
- Trained on 3.5 trillion tokens
- 7 million GPU hours
- Quality on par with PaLM 2, outperforming Llama 2 and GPT
-3.5 across benchmarks
- 4-bit and 8-bit show little degradation
- Trained on 3.5 trillion tokens
- 7 million GPU hours
- Quality on par with PaLM 2, outperforming Llama 2 and GPT
-3.5 across benchmarks
- 4-bit and 8-bit show little degradation