OpenAI and Microsoft Azure to deprecate GPT-4 32K
twitter.com
twitter.com
gpt-4, gpt-4-turbo, and gpt-4o are not the same models. They are mostly close enough when you have a human in the loop, and loose constraints. But if you are building systems off of the (already fragile) prompt based output, you will have to go through a very manual process of tuning your prompts to get the same/similar output out of the new model. It will break in weird ways that makes you feel like you are trying to nail Jello to a tree
There are software tools/services that help with this, and a ton more that merely promise to, but most of the tooling around LLMs these days gives the illusion of a reliable tool rather than results of one. It's the early days of the gold rush still, and every one wants to be seen as one of the first
[2]: https://insurtechdigital.com/articles/chatgpt-the-risks-and-...
--- please disregard [1] it was a terrible initial source I pulled of Google
[1]: https://medium.com/artivatic/use-of-chatgpt-4-in-health-insu...
Where what it seems we are getting with a lot of these companies shoving AI into something and calling it a product, is an MVP that is an MVP due to an unknown and untested nature.
We get a lot of MVP crap before, don't get me wrong. But at least it was understood crap. Sure it may have bugs in it and that is to be expected. But there was a limit in how wrong it could go. Since at the end of the day it was still limited to the code within the application and the server (if there is one).
Meanwhile when an over-reliance on an LLM goes wrong, depending on how it goes wrong could be catastrophic.
As we have seen time and time again just in the last couple months, when LLM's are shoved into something we seem to get a serious lack of testing under the guise of "beta".
If you rely on third party packages of any type, you have dependencies that can rapidly and unexpectedly break with an update. Semantic versioning is supposed to help with this, but it doesn’t always help.
Probably the best description of working with LLM agents I've read
Back in June 2023 and November 2023, we announced the following models will be deprecated on June 13th, 2024:
gpt-3.5-turbo-0301
gpt-3.5-turbo-0613
gpt-3.5-turbo-16k-0613
We noticed that your organization recently used at least one of these models. To help minimize any disruption, we are extending your access to these models for an additional 3 month grace period until September 13th, 2024. After this date, these models will be fully decommissioned.Curious if that is a business decision or a technical decision aka if optimizations for cheap and fast 128k gpt4o work only for small outputs
1. Model Capability
You are right that mechanically, input and output tokens in a standard decoder Transformer are "the same". A 32K context should mean you can have 1 input tokens and 32K output tokens (you actually get 1 bonus token), or 32K input tokens and 1 output token,
However, if you feed an LM "too much" of its own input (read: have too long an output length), it starts to go off the rails, empirically. The word "too much" is doing some work here: it's a balance of both (1) LLM labs having data that covers that many output tokens in an example and (2) LLMs labs having empirical tests to have confidence that the model won't reasonably go off the rails within some output limit. (Note, this isn't pretraining but the instruction tuning/RLHF after, so you don't just get examples for free)
In short, labs will often train a model targeting an output context length, and put out an offering based on that.
2. Infrastructure
While mathematically having the model read external input and its own output are the same, the infrastructure is wildly different. This is one of the first things you learn when deploying these models: you basically have a different stack for "encoding" and "decoding" (using those terms loosely. This is after all still a decoder only model). This means you need to set max lengths for both encoding and decoding separately.
So, after a long time of optimizing both the implementation and length hyperparameters (or just winging it), the lab will decide "we have a good implementation for up to 31K input and 1k output" and then go from there. If they wanted to change that, there's a bunch of infrastructure work involved. And because of the economies of batching, you want many inputs to have as close to the same lengths as possible, so you want to offer fewer configurations (some of this bucketing may be performed hidden from the user). Anyway, this is why it may become uneconomical to offer a model at a given length configuration (input or output) after some time.
The docs still mention bigger models with 128k tokens and smaller models with 8k tokens. It seems reasonable to optimize for big and small use cases differently? I don't see how we are being "robbed".
And no one here touched how high a multiple the cost is, so I assume its pretty high.
The appropriate level of due diligence for each LLM model transition is to run your various prompts into the new model and make sure they still produce the correct output; and, if they don't produce good output, to update the prompts so that they continue to produce good output.
Just yesterday, I was experimenting with 4o and assumed I could do a flat migration for some work. 4o actually provided worse results - results I explicitly asked to *not* have in my 4 output (and that I didn't have in my 4 output).
It's tedious to have to change models after you've already done a proper validation suite against one model.
That would be (at least my) complaint.
I've even version-stamped the models I use on purpose to avoid surprises.
To do this, I've made, and continue to make tradeoffs. Among many of the tradeoffs I'm currently making that I intend to resolve ASAP is having OpenAI as a single source of failure. I intend to have some of the other hosted solutions as other options for LLM processing. One of the many, many options that will be considered at that time is self-hosting it, as well, as another option.
I've already spent more time than I should perfecting various smaller pieces, increasing reliability, etc. Each time I choose perfection, I lose more time, more runway, more potential market share; and, something I've recently had to learn:
Each time I lock myself in a previous step to get that step perfect, I miss the lessons I'm about to have to learn in the next stage of the process, including new issues I'll run into that increase the next step's complexity above my initial estimates.
Everything is a tradeoff. Choosing to use a commercially available solution with known and relatively set costs while accepting it may slowly change underfoot (while also knowing I have alternatives I can swap to if an emergency comes up that should only take a little bit to transfer to) is one I've made.