Why is the 32K context only twice as expensive as the 8K context?
Are they using sparse attention or something? I don't think flash attention on it's own can explain it.
EDIT: Oh right, if the cost is per token so if you actually fill the context then it is 8x more, which makes much more sense.