If it switches mid conversation, this is a massive increase in token consumption because it has to re-read your conversation into cache, right?
yes. At the bottom of the release post it says that they are releasing two new features, one of which is customizing fallback behavior instead of blocking for restricted models