It looks like GPT-4-32k is rolling out
community.openai.com
community.openai.com
32,000 (tokens) * 4 = 128,000 (characters)
> While a general guideline is one page is 500 words (single spaced) or 250 words (double spaced), this is a ballpark figure - https://wordcounter.net/words-per-page
Assuming (on average) one word = 5 letters, context ends up being ~50 pages (128000 / (500 * 5)).
Just to put the number of "32k tokens" into somewhat estimated context.
Interestingly, using 2-space indentation, as soon as you are about 3-4 indentation levels deep you spend as many tokens on indentation as on the actual code. For example, "log::LevelFilter::Info" is 6 tokens, same as 6 consecutive spaces. There are probably a lot of easy gains here reformatting your code to use longer lines or maybe no indentation at all.
Using tabs for indentation needs fewer tokens than using multiple spaces for indentation.
32k is ~4.5k LoC
https://tiktokenizer.vercel.app/
Not sure, why they don't support GPT-4 on their own website.
As far as I can tell, the initial reading in of the document (let's say a 20k token one), will be a repeated cost for each subsequent query over the document. If I have a 20k token document, and ask 10 follow-up prompts consisting of 100 tokens, that would take me to a total spend of 20k * 10 + (10 * 11)/2 * 100 = 205,500 prompt tokens, or over $6. This does not include completion tokens or the response history which would edge us closer to $8 for the chat session.
TL;DR the puck should get there very soon, not just Really Soon Now.
https://www.pinecone.io/learn/langchain-conversational-memor...
But aside from this, if I have a large document that I want to "chat" with, it looks like I am either chunking it and then selectively retrieving relevant subsections at question time, or I am naively dumping the whole document (that now fits in 32k) and then doing the chat, at a high cost. So 32k (and increased context size in general) does not look to be a huge gamechanger in patterns of use until cost comes down by an order of magnitude or two.
I think we may end up with first and second stage completions, where the first stage prepares the context for the second stage. The first stage can be a (tailored) gpt3.5 and the second stage can do the brainy work. That way you can actively control costs by making the first stage forward a context of a given maximum size.
Is it $0.60 even for the input tokens?
38 seconds for that example.
i guess i’m running into a different limit (not context length), or maybe i’m misremembering
Also, I pay for ChatGPT but I have none of the new features except for GPT4. Very frustrating.
The really odd thing is that I was given GPT-4 with browsing alpha enabled - for a single session last week.
As soon as I reloaded the page, it was gone. Since then the picture has reverted back to the above.
Twitter has become a bit painful to read these days, with all the AI influencers posting about what GPT-4 and plugins, code interpreter etc. can do.
The latter is definitely the cheapest option; updates are trivial.
GPT-4 seems to show that linear algebra definitely can do the job, but training is so expensive and the model gets so huge and inflexible.
It seems like having fixed format vectors of knowledge that the model can use-- denser and more precise than just incorporating tool results as tokens like OpenAI's plugin approach-- is a path forward towards extensibility and online learning.
There may be a middle ground between these two approaches though. If every query used the same prompt prefix (because you only update the codebase + docs occasionally) then you could put it into the model once and cache the keys and values from the attention heads. I wonder if OpenAI does this with whatever prefix they use for ChatGPT?
I'm pretty sure they're using a 4k GPT-4 model for ChatGPT Plus, even though they only announced 8k and 32k... It can't handle more than 4k of tokens (actually a little below that, starts ignoring your last few sentences if you get close). If you check developer tools, the request to an API /models endpoint says the limit for GPT-4 is 4096. It's very unfortunate.
Another thing that annoys me is how most updates don't get a changelog entry. For whatever reason, they keep little secrets like that.
Every time I see a company act like this, more responsive and truly open competition eventually eats their lunch.
52/1.92 = 27 416/1.92 = 217
So using GPT-4 with 32k tokens, 27 times per hour, or 217 times per day, in terms of cost, is approximately the equivalent of another dev
Not that it matters for the calculation, but i wonder how long such a request (ingesting 32k tokens and responding with a similar amount) would take.
At the speed of regular ChatGPT take would take a good while.
At the upper bound, this would be $2 * 3 * 60 * 4 = $1440 a day.
Thankfully, I am using retriever-augmentation and context stuffing into the base 4k model, so costs are manageable.
The 32k context model cannot be deployed into a production app at this pricing as a more capable drop-in replacement for shorter-context models.
Care you elaborate? This sounds very interesting & useful. Just anything about the setup and implementation would be super helpful.
The initial blog post was only just over a month ago, and it was announcing alpha access for a few users and developers:
> Today, we will begin extending plugin alpha access to users and developers from our waitlist. While we will initially prioritize a small number of developers and ChatGPT Plus users, we plan to roll out larger-scale access over time.
https://openai.com/blog/chatgpt-plugins
We are literally 1 month into the alpha of plugins.
The developer livestream was on March 14th: https://www.youtube.com/live/outcGtbnMuQ?feature=share.
The time since GPT-4 already feels something like 6 months. So far I'm perpetually feeling behind.
I struggle to keep up and all I need to do is understand developments well enough to simplify them in to palatable morsels for my tech skeptic colleagues in politics and non profits.
Challenging because they have a form of technology PTSD. when they hear "new technology" nft's of monkeys with 6 digit prices and peter thiel's yacht flash before their eyes and they see red.
And I can't really blame them, the rhetoric around crypto was enough to sour most non techies (in my little corner of lefty politics anyway) against the idea that any tech advancement is noteworthy. One of the first more serious individuals in politics to hear me out did so because "i sounded like one of the early linux proselytizers" lol.
Completely agree how time has slowed. I rotate between absolute giddy anticipation at our future thanks to the tech and nihilistic doomerism. Even as a hobbyist though I knew to take this seriously since I saw robert miles talk about gpt 2 in 2017(?) and note there's zero sign of these things plateauing in ability simply by ramping up parameter count.
I've gone on long enough but that live stream felt like the intro to a sci fi movie at points. Can't wait to have multi modal and plugins rolled out.
I expect that in the next 5yrs developer workflows will completely change based on all the LLM stuff.
I think it's always difficult to tell if new tech is just hype or will have real impact, but it really feels to me like LLMs will have real impact. Maybe not as much as they are being hyped, but definitely legit impact. There's a possibility of even greater impact than the hype as well.
It’s very slow, almost 10X slower than ChatGPT
It’s integration is bad. For most plugins it doesn’t do anything smart with its API call. For example if I ask “Nearest cheap International flight”, it literally goes to Kayak and searches Nearest Cheap International Flight, if Kayak can’t handle that query, GPT can’t either.
The only plug-in with good integration is Wolfram and it makes so many syntax errors calling Wolfram that it’s thrash. Often it just syntax errors out for half my queries
I wouldn’t have minded if they spent a few more months internally testing plug-ins before rolling it out to me, seeing it’s current state. The annoying thing is the chat website automatically starts at plugins mode which is borderline unusable. So every time I have to click on the drop-down and then choose ChatGPT or GPT4.
For code, I use phind.com.
I also find it strange they don't contrast gpt4 and gpt3.5
And don't forget that all the LLaMA-based models only have 2K context size. It's good enough for random chat, but you quickly bump into it for any sort of complicated task solving or writing code. Increasing this to 4K - like GPT-3.5 has - would require significantly more RAM for the same model size.
For things where correctness matters, the majority of cost will still come from humans who are in charge of ensuring correctness.
Once I filled that in I got access within a few days.
https://openai.com/waitlist/plugins
If you are the person to say: "I am a developer and want to build a plugin"
Then it is likely you missed the option to request which plugins you want access to.
I stopped paying the pro because without plugins it didnt do that much tbh
(Also UI probably tanks too. I dread what the OpenAI Playground will do when you start actually using 32k model for real, like throwing a 15k token long prompt at it. ChatGPT UI has no chance.)
Until they cut down the cost then they should worry yeah
Those startups killed themselves. A 32K context was advertised as a feature to be rolled out the same day GPT-4 came out.
Also - what startups are getting even remotely close to 32K context at GPT-4’s parameter count? All I’ve seen is attempts to use KNN over a database to artificially improve long term recall.
Something like the following
Here is the Template class:
…
Here is an example component:
…
Here is an example Input element:
…
I need to create another input element that allows me to select a number from a drop down between 1/32 and 24/32 in 1/32 increments
It works really well, you can tell it to implement new features or mutate parts of the code and it having the entire (or a lot of) the code in its context really improves the output.
The biggest caveat: shit is expensive! A full 32k token request will run you like $2, if you do dialog back and forth you can rack up quite the bill quickly. If it was 10x cheaper, I would use nothing else, having a large context window is that much of a game changer. As it stands, I _very_ carefully construct the prompt and move the conversation out of the 32k into the 8k model as fast as I can to save cost.
This is why it's so easy to burn up lots of tokens very fast.
Wonderous as this new tech is, it seems a bit much to be paying $2 a question in a conversation about a 32k token text.
Using it on a simple conversation is not its intended purpose, that's like using a supercomputer to play pong.
It's quite interesting if you can tweak your search just right. You can even use less tokens than 8K even!
That's not really how conversation/chat works is it?
It is like complaining that HTTP is limiting because it is stateless. Build state on top of it.
I haven't actually even hit the 8k limit yet, and even experimenting with 32k is pretty expensive, so I'm not sure what I'd do with it.
It is strange they can roll out 32k for some, while not even having 8k for everyone yet.
Plus seems expensive to me and it is still rate limited quite a lot.
I guess it's going to take further optimisation to make it worthwhile for OpenAI.
I’m wondering if this is a quick way to get to $5 each round trip (send and receive)
It gave me the same vague feeling of annoyance and disgust at the "look how smart I am" linguistic obstreperousness I get when reading the real deal.
The output from that prompt seems spectacular, so I'm wondering if there are any other differences.
I just tried the same prompt with GPT-4 and the style was much more GPT-like, what I'm used to, not near the same quality as in the OP, although maybe it's just luck?
I'm actually not sure if longer responses can be expected with the 32k vs 8k models. Anyone from OpenAI care to comment on that?
Once the limit is hit it will stop, sometimes mid sentence.
I have a slack bot at work with a system document which is a little over 1k tokens, meaning there is around 3k tokens left for questions and replies.
Trick I am currently doing is to prune older messages to keep it under the limit.
Oh that’s why GPT does that in a long thread.
It makes sense in retrospect it’s the whole conversation not the individual messages.
ChatGPT UX has much to be desired, there should be error messages that communicate this stuff better.
There's also the token context: how many words into the "past" it considers when formulating the next word.
They're different things. You can generate very, very long responses with a model with a short context window; it just will have amnesia about what it said earlier-- though OpenAI often seems to restrict you / prevent you from having context scroll off in this way.
import tiktoken
len(tiktoken.encoding_for_model("gpt-4").encode(contents))Direct link to the relevant sign up form fwiw:
https://customervoice.microsoft.com/Pages/ResponsePage.aspx?...
We got access a few weeks back via Azure.
LoRA is for adapting the model to a certain model. It usually means you need to give it shorter prompts, but for book summarization, it wouldn't help.