Stanford benchmarks and compares numerous Large Language Models
crfm.stanford.edu
crfm.stanford.edu
- Not all these models are instruct aligned
- Cohere's suite is free for personal use
- anthropic uses reinforcement learning from ai feedback for instruct alignment, removing humans from the process almost entirely(humans still make the rules the language model checks the responses against).
- open ai's models are pretty poorly calibrated compared to its peers. Calibration is how well the model's confidence solving a problem tracks with how it actually performs that problem. It seems to be a consequence of the way they are performing the instruct tuning (or maybe it's a consequence of instruct tuning in general). The gpt-4 report actually shows this. Base gpt-4 was excellently calibrated but rlhf knocked it right out. The excellent calibration was also independently verified (in the medical domain) from the paper about gpt-4's medical prowess. https://arxiv.org/abs/2303.13375
Yep, anything that old is ancient in the AI world
Then that should square off against a model trained on everything..