As far as I know Anthropic haven't released the tokenizer for Claude - unlike OpenAI's tiktoken - but your tool lists the Claude 3 models as supported. How are you counting tokens for those?
As far as I know Anthropic haven't released the tokenizer for Claude - unlike OpenAI's tiktoken - but your tool lists the Claude 3 models as supported. How are you counting tokens for those?
At this moment, Tokencost uses the OpenAI tokenizer as a default tokenizer, but this would be a welcome PR!
I've been bugging Anthropic about this for a while, they said that releasing a new tokenizer is not on their current roadmap.
Frequently, contracts will have room for additional charges if circumstances change even a little, or products will have a market rate (fish, equity, etc.).
It might seem absurd but variable cost things are not uncommon.
I am merely hypothesizing, it may not be nondeterministic but I'm not going to assume it's not.
I've only ever seen: fixed price based on destination (typically for fares originating from an airport), negotiated, or metered. A better analog analogy would be metered pricing, but where the cost per mile is a secret.
Similarly, as LLMs become more and more commonplace, the pricing models will need to be more predictable. My LLM expenses are only around $100/month, but it's a bigger impediment to pushing projects to production when I can't tell the boss exactly how it'll be priced.
It looks like tiktoken is the default for most of the methods.
Disclaimer: I didn't fully trace which are being used in each case/model.
There are no cases for Claude models yet.
I wonder if anyone has run a bunch of messages through Anthropic's API and used the returned token count to approximate the tokenizer?