GPT-4 API General Availability
openai.com
openai.com
If you use GPT, you're giving OpenAI money to lobby the government so they'll have no competitors, ultimately screwing yourself, your wallet, and the rest of us too.
OpenAI has no moat, unless you give them money to write legislation.
I can currently run some scary smart and fast LLMs on a 5 year old laptop with no GPU. The future is, at least, interesting.
It's been a month or two since I've tried but the results were depressingly slow and useless for more or less every task I tried.
Every time a model is claimed to be "90% of GPT-3" I get excited and every time it's very disappointing.
(On that note, after using GPT-4, GPT-3 now seems disappointing almost every time I interact with it.)
Text generation style instead of chat style is another avenue that makes the feedback time not so annoying for a developer.
at 100ms/token, it's faster than most people type, I think. That's what you might get on an old laptop with a 7B model.
There's a useful leaderboard here to help you pick a model: https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderb...
It really depends on your task, lots and lots of natural language type tasks give great results, the models seem to have extensive knowledge of many fields. So for some kinds of Q&A bot (technical or not), for copy blurbs, for fiction, game NPCs, etc, the models (especially 13B and up) can be breathtaking, even moreso considering they run on bottom-dollar consumer hardware (I paid $250 for the laptop I'm developing on).
There are of course some things that neither the local LLMs nor GPT4 can do, like create useful OpenSCAD models :)
Things keep getting better, newer quantization methods give you more smarts in the same amount of RAM at basically the same speed -- the models are getting better, there are more permissively licensed ones now.
Like, not vaguely hand wavey stuff, specifically, what model and what inference code?
I get nothing like that performance for the 7B models, forget the larger models, using llama.cpp on a pc without an nvidia GPU.
You can take a look at a list of evals here: https://llm-tracker.info/books/evals/page/list-of-evals - for general usage, I think home-rolled evals like llm-jeopardy [2] and local-llm-comparison [3] by hobbyists are more useful than most of the benchmark rankings.
That being said, personally I mostly use GPT-4 for code assistance to that's what I'm most interested in, and the latest code assistants are scoring quite well: https://github.com/abacaj/code-eval - a recent replit-3b fine tune the human-eval results for open models (as a point of reference, GPT-3.5 gets 60.4 on pass@1 and 68.9 on pass@10 [4]) - I've only just started playing around with it since replit model tooling is not as good as llamas (doc here: https://llm-tracker.info/books/howto-guides/page/replit-mode...).
I'm interested in potentially applying reflexion or some of the other techniques that have been tried to even further increase coding abilities. (InterCode in particular has caught my eye https://intercode-benchmark.github.io/)
[1] https://github.com/turboderp/exllama#results-so-far
[2] https://github.com/aigoopy/llm-jeopardy
[3] https://github.com/Troyanovsky/Local-LLM-comparison/tree/mai...
[4] https://github.com/nlpxucan/WizardLM/tree/main/WizardCoder
Is exllama an alternative to llama.cpp?
There's definitely some prompt magic openai does behind the scenes that helps beat the raw style local llms usually go for. With proper prompting you can get chatgpt like answers.
But to address both: is it very relevant what LLM you use right now? Local or hosted, openAI or other?
It seems like the interface has converged around chat-based prompts.
New ideas for tuning or improving the efficiency of foundational models are published almost every week.
If one wants to build a product on top of of generative AI, why not simply start with what’s free or works with one’s dev environment?
Presumably, the interaction with or API to text-based gen AI will be very similar no matter what engine is best for your use case at any given time.
This would imply these backends will be swappable, the way web services are that copy AWS S3 APIs.
So, to return to the point, can’t people just build their product with openAI or other and plan to move away based on the cost and fit for their circumstances?
Couldn’t someone say prototype the entire product on some lower-quality LLM and occasionally pass requests to GPT4 to validate behavior?
It seems far-fetched to believe this tech can be constrained by legislation.
OpenAI can lobby all they want, it won’t necessarily buy them anything. Look what happened with FTX.
Since LLMs can be run locally and the engines be black boxes to the user, how could a legislative act really prevent them from being everywhere—-especially given the public utility.
It can be done -- it is the basis for assisted generation and related work. It does require full access to the model, to be time and money-efficient. See https://huggingface.co/blog/assisted-generation
Disclaimer: I'm the author of the blog post linked above.
This, infact, might be a better way to do inference anyway: https://twitter.com/Francis_YAO_/status/1675967988925710338
> So, to return to the point, can’t people just build their product with openAI or other and plan to move away based on the cost and fit for their circumstances?
Depends. There are signs that folks are buying into GPT-specific APIs (like function calls) which may not be as easy to migrate away from.
I have an old laptop with 16GB RAM and no GPU. Can I run these models?
https://huggingface.co/TheBloke
There's a LocalLLaMA subreddit, irc channels, and a whole big community around the web working on it on GitHub nd elsewhere.
edit: I forgot to directly answer you: yes you can run these models. 16GB of plenty. Different quantizations give you different amounts of smarts and speed. There are tables that tell you how much RAM is needed per which quantization you choose, as well as how fast it can produce results (ms per token). e.g. https://github.com/ggerganov/llama.cpp#quantization where RAM required a little more than the file size, but there are tables that list it explicitly which I don't have immediately at hand.
While you can run all kinds of GPTs locally, GPT-4 still smokes everything right now – and even it is not actually good enough to not be a lynchpin for a lot of cases yet.
Not if you want it to write adult (graphically pornographic or violent) content.
Here's a handy model comparison chart (this is a coding benchmark, so coding-only models tend to rank higher): https://i.imgur.com/AqSjjj2.jpeg
It beats Claude and Bard.
You could probably get a 4bit 15B model going in 16GB of RAM and be approaching GPT4 in capability.
...on an old laptop, lol
Let's eat OpenAI's lunch! They deserve it for trying to steal this tech by "privatizing" a charity, hiding scientific data that was supposed to be shared with us by said charity whose purpose was to help us all, and dishonestly trying to persuade the government not to let us compete with them.
It used to be this took a few people to come up with writing actual responses to forum posts all day, or marketing operations plans, or pro- or anti-thing propaganda plans.
But now, you could astroturf a movement with a GPU, a ChatGPT clone, some bots and vpns hosted from a single computer, a cron job, and one human running it.
If you thought disinformation was bad 2 years ago, get ready for fully automated disinformation that can be targeted down to an online community or specific user in an online community...
[1] GPT4 is 8 x 220B params = 1.7T params: https://news.ycombinator.com/item?id=36413296
I would guesstimate that the great majority of prompts are trash. People playing with a toy and amusing themselves. The platform sends those to the trash models.
For the other tiny percentage that produces a prompt the size of a paragraph, using the techniques published by OpenAI themselves, they likely get the higher tier models. This is also why I believe many are recently complaining about the quality of the outputs. When your chat history is filled with "have waifu pretend to be my girlfriend" then whatever memory the model is maintaining will be poisoned by the quality of your past prompts.
Garbage in, garbage out. I am certain that the #1 priority for OpenAI/Microsoft is lowering the cost of each prompt while satisfying the majority.
The majority is not in HN.
That's certainly true, but it's hard to deny the quality of gpt 4. If the issue is the training data, let's just use their training data, it's not like they had to close up shop because of using restricted data.
I think the issue is more on the financial side, it must have been extremely expensive to train gpt 4. Open source models don't have that kind of money right now.
I'll finance open source models once they are actually good, or show realistic promises of reaching that level of quality on consumer hardware. Until then, open source will open source.
I've never bought any kind of subscription or paid api costs to openai, but if gpt 4 finally reached the point where I feel like it's a lot better than just good enough, I'll happily pay for it (while still being on the lookout for open source models that fit my hardware).
a 1billion parameter model beats 175billion parameter GPT3.5
OpenAI wants us all to drink the kool-aid.
Depending on your pre-prompt, your fine-tune (i.e. which model you downloaded), and your specific task, the results can be startlingly good, it's crazy that you can do this on a $250 laptop. I stay up nights working on it lately, it's so interesting.
More importantly, things change by the day. New models, new methods, new software, new interfaces... the possibilities are endless... unless we let OpenAI corrupt our government(s).
Is there a local solution that is at least as intelligent as GPT 3.5 in that regard that I can run in a container?
You can rent time on a hosted GPU, sharing a hosted model with others.
Most of peoples' problem is watching the AI type, it's not instant, but then not all (or even most) applications need to be instant. You can also avoid that by having it return everything at once instead of streaming style.
Local absolutely can scale. All kinds of fun things can be done on a machine with 16GB of RAM, or 8GB if you work harder.
Funny, for me it is the complete opposite. I created an interface in Matrix that does just that: return everything at once. But the lag annoys me more than the slow typing in the regular chat interface. The slow typing helps me keep me focused on the conversation. Without it, my mind starts wandering while it waits.
https://github.com/imartinez/privateGPT is great if you want do it with code.
Their moat is that they had access to data sources which have since been clamped down on, eg reddit and twitter apis.
I sympathize with the idea of wanting to run a local LLM, but IMO, this would require building a desktop with a GPU and plenty of horsepower + silent cooling and put it somewhere in a closet in my apartment. Running LLMs on my laptop is (to me) clearly a waste of my time and its battery/cooling.
> I can currently run some scary smart and fast LLMs on a 5 year old laptop with no GPU.
And, something useful or just playing? I played with local models, and will keep playing, training, experimenting. It's interesting, but not a solution, not yet.
This will end my usage of openai for most things. I doubt my $5-$10 API payments per month will matter. This just lights more of a fire under me to get the 65B llama models working locally.
What are funs things we can with it until it sunsets on January 4, 2024?
When using davinci the onus is on you to construct prompts (memories) which is fun and powerful.
====
97% of API usage might be because of ChatGPT's general appeal to the world. But I think they will be losing a part of the hacker/builder ethos if they drop things like davinci-003, which might suck for them in the long run. Consumers over developers.
I used the "Complete" UI (from the Playground) for a bit before the "Chat" interface was available; I don't really think there's anything you couldn't do in the "Complete" UI that you couldn't also do in the "Chat" UI.
One supposes openai has a 6 month notice period vs a 12 month period for azure. This might generally effect one’s appetite in choosing which endpoint to use for any model.
But with davinci at the same price point as GPT-4 I'm hoping the latter is enough of a step up in its variety of vocabulary and nudgeable sophistication of language to be a drop in replacement.
Though in general I think there's an under appreciation for just how much is being lost in the trend towards instruct models, and hope there will be smart actors in the market who use a pre-optimization step for instruct prompts that formats it for untuned models. I'd imagine that parameter size to parameter size that approach will look much more advanced to end users just by not lobotomizing the underlying model.
More info:
https://platform.openai.com/docs/model-index-for-researchers
I guess they'll give you early access to it.
*"They have just locked down access to a model which they basically realized was way more valuable than even they thought - and they are in the process of locking in all controls around exploiting the model for great justice?"*
Trying the same prompts that gave nerfed "I am just an AI I can't speculate about the future" bs on completion API gave somewhat better results, but most of the time they were flagged as breaking the guidelines which is a TOS breach if done enough times.
This can be solved other than open models. The same thing happened with stable diffusion. Good thing it's open so you can still use the pre-nerfed 1.6 models.
I know it might be edgy or unpopular but I don't think one entity should decide how we can use this powerful tool. No matter its implications and consequences.
FOS for the win.
But mostly it has to do with the fact that LLM do what they've seen. And if they've been fine-tuned to not respond to some classes of things they'll misapply that to lots of other things. That's why most people go for the "uncensored" fine tuning datasets for the llamas even for completely sfw use cases.
Wait, they're not letting you use your own fine-tuned models anymore? So anybody who paid for a fine-tuned model is just forced to repay the training tokens to fine-tune on top of the new censored models? Maybe I'm misunderstanding it.
It's a deal-breaker for many.
I assume if you reach out they throw some credits at you
I have much better performance by "prompt tuning" - when question arises, I search 30 similiar examples in training set, and send it to non-tuned GPT and ask the question and get much better performance than fine-tuned older models.
It's not cost-effective, but it may be part of a valid business plan.
Do you have any recommendations for good open models that businesses could use today?
From what I've seen in the space, I suspect businesses are building fine tuned models against closed models because those are the only viable models to build a business model on top of. The quality of open models isn't competitive.
Just sell access at a higher price than you get it
Either directly, on on average based on your user stories
On fine-tuning:
> We will be providing support to users who previously fine-tuned models to make this transition as smooth as possible.
On embeddings:
> We will cover the financial cost of users re-embedding content with these new models.
gpt-3.5-turbo is the model behind ChatGPT. It's chat-fine-tuned which makes it very hard to use for use-cases where you really just want it to obey/complete without any "chatty" verbiage.
The "davinci-003" model was the last instruction tuned model, but is 10x more expensive than gpt-3.5-turbo, so it makes economical sense to hack gpt-3.5-turbo to your use case even if it is hugely wasteful from a tokens point of view.
Think of the properties you want in the JSON object, then send those to ChatGPT as required parameters for a function (even if that function doesn't exist).
# Definition of our local function(s).
# This is effectively telling ChatGPT what we're going to use its JSON output for.
# Send this alongside the "model" and "messages" properties in the API request.
functions = [
{
"name": "write_post",
"description": "Shows the title and summary of some text.",
"parameters": {
"type": "object",
"properties": {
"title": {
"type": "string",
"description": "Title of the text output."
},
"summary": {
"type": "string",
"description": "Summary of the text output."
}
}
}
}
]
I've found it's not perfect but still pretty reliable – good enough for me combined with error handling.If you're interested, I wrote a blog post with more detail: https://puppycoding.com/2023/07/07/json-object-from-chatgpt-...
Have you tried getting your formatted JSON out via the new Functions API? I does cure a lot of the deficiencies in 3.5-turbo.
The difference between them is that the chat models are much more... chatty - they're trained to act like they're in a conversation with you. The chat models generally say things "Sure, I can do that for you!", and "No problem! Here is". The conversation style is generally more inconsistent in it's style. It can be difficult to make it only return the result you want, and occasionally it'll keep talking anyway. It'll also talk in first person more, and a few things like that.
So if you're using it as an API for things like summarization, extracting the subject of a sentence, code editing, etc, then the chat model can be super annoying to work with.
OK maybe I'm stupid but I am a paying OpenAI API customer and I don't have it yet. I see:
gpt-3.5-turbo-16k
gpt-3.5-turbo
gpt-3.5-turbo-16k-0613
gpt-3.5-turbo-0613
gpt-3.5-turbo-0301
I don't see any gpt-4Edit: Probably my problem is that I upgraded to paid API account within the last month, so I'm not technically a "paying API customer" yet according to the accounting definitions.
Same for me. I signed up only a few days ago and was excited to switch to "gpt-4" but I haven't paid the first bill (save the $5 capture) so I probably have to continue to wait for this.
I made a very simple command-line tool that calls the API. You run something like:
> ask "What's the opposite of false?"
https://github.com/codazoda/askaihttps://www.pastery.net/ccvjrh/
It also does streaming, so it live-prints the response as it comes.
1. chat subscription only
2. i have paid for api calls but don't have a subscription
and only #2 currently has gpt4 available in the playground
The mass adoption of the ChatGPT APIs compared to the old Completion APIs proves my initial blog post on the ChatGPT API correct: developers will immediately switch for a massive price reduction if quality is the same (or better!): https://news.ycombinator.com/item?id=35110998
Medical writing is the opposite, with unstated premises, semi-random associations, and rarely a meaningful sentence.
So still waiting to be on the same 32 pages...
I mostly use it for generating tests, making documentation, refactoring, code snippets, etc. I use it daily for work along with copilot/x.
In my experience GPT3.5turbo is... rather dumb in comparison. It makes a comment explaining what a method is going to do and what arguments it will have - then misses arguments altogether. It feels like it has poor memory (and we're talking relatively short code snippets, nothing remotely near it's context length).
And I don't mean small mistakes - I mean it will say it will do something with several steps, then just miss entire steps.
GPT3.5turbo is reliably unreliable for me, requiring large changes and constant "rerolls".
GPT3.5turbo also has difficulty following the "style/template" from both the prompt and it's own response. It'll be consistent then just - change. An example being how it uses bullet points in documentation.
Codex is generally better - but noticeably worse then GPT4 - it's decent as a "smart autocomplete" though. Not crazy useful for documentation.
Meanwhile GPT4 generally nails the results, occasionally needing a few tweaks, generally only with long/complex code/prompts.
tl;dr - In my experience for code GPT3.5turbo isn't even worth the time it takes to get a good result/fix the result. Codex can do some decent things. I just use GPT4 for anything more then autocomplete - it's so much more consistent.
Where 3.5 excels is with programmatic access. You can ask it for 2x as much text between setup so the end result is well formed and still get a reply that's cheaper and faster than 4 (for example, ask 3.5 for a response, then ask it to format that response)
I built notionsmith.ai on 3.5: for some time I experimented with GPT 4 but the result was significantly worse to use because of how slow it became, going from ~15 seconds per generated output to a minute plus.
And you could work around that with things like streaming output for some use cases, but that doesn't work for chain of thought. GPT 4 can do some tasks without chain of thought that 3.5 required it for, but there are still many times where it improves the result from 4 dramatically.
For example, I leverage chain of thought in replies to the user when they're in a chat and that results in a much better user experience: It's very difficult to run into the default 'As a large language model' disclaimer regardless of how deeply you probe a generated experience when using it. GPT 4 requires the same chain of thought process to avoid that, but ends up needing several seconds per response, as opposed to 3.5 which is near-instant.
-
I suspect a lot of people are building things on 4 but would get better quality of output if they used more aspects of chain of thought and either settled for a slower output or moved to 3.5 (or a mix of 3.5 and 4)
poe law
From using voice in the ChatGPT iOS app, I surmise that Whisper is very good at working out what you've actually said.
But it's really annoying to have to say my whole bit before getting any feedback about what it's gonna think I said. Even if it's getting it right at an impressive rate.
Given this is how OpenAI themselves use it (say your whole thing before getting feedback), I don't know that the API is set up to be able to mitigate that at all, but it would be really nice to have something closer to the responsiveness of on-device dictation with the quality of Whisper.
Superficially, I think this will work very well, but slightly worse than whisper (with the advantage ofc being that its better at real-time transcription).
[0]https://machinelearning.apple.com/research/attention-free-tr...
Pretty crazy stuff — perfectly understandable translations.
Using built-in text input showed quite good results since ChatGPT is still understanding the ask quite well
Wait for what!? Christmas? When we can open our presents and have a GPT 4 inside?
It's like they took a leaf from Google's "how to guarantee the failure of a new product" marketing. That is: restrict access, ensuring that to word-of-mouth marketing can't possibly work because none of your friends are allowed to try the product.
The announcement here is "general availability" of the GPT-4 model...
...but not the 32K context model. Not the multi-modal version with image input. No fine-tuning. Only one model (chat).
As of today, I can only access GPT 3.5 via Azure Open AI service and the Open AI API account that I have.
What's the point of all these arbitrary restrictions on who can access what model!?
I can use GPT 4 via Chat, but not an API. I can use an enhanced version of Dall-E via Bing Image Creator, but not the OpenAI API. Some vendors that have been blessed by the Great and Benevolent Sam Altman have access to GPT-4 32K, the rest of us don't.
Sell the product, not the access to it.
Don't be like the Soviet Union, where you had to "know someone" to get access.
So they open things carefully, pull back when necessary like when they limited use of the public GPT-4 version of ChatGPT. That doesn’t seem too unreasonable. And yes sure, some amount of it might be attempts to manufacture scarcity to increase the hype. It’s an old tactic and hardly comparable to Soviet Russia.
If there's not enough GPUs at a certain price point, raise prices. Then lower prices later when GPUs become available.
They did it with GPT 3.5, so why not GPT 4?
If we guesstimate that every 100 customers needs 1 NVIDIA GPU (completely random guess), then that means OpenAI needs to buy more GPUs for every 100 new customers using GPT-4. The problem is there's a GPU shortage so it's hard to add more GPUs by just throwing money at the problem.
https://www.fierceelectronics.com/electronics/ask-nvidia-ceo...
Infrastructure.
> It's like they took a leaf from Google's "how to guarantee the failure of a new product" marketing.
Yeah, an infamous guaranteed failure: GPT-4. (canned laughter)
"Lick my boots, in person, and you can be one of the privileged few" is very much the behaviour of a Communist dictatorship, not a capitalist corporation.
I can spin up an Azure VM right now in almost any country I choose... except China. That's the only one where I have to beg the government for permission.
I've had completions with it that had character and creativity that I have not been able to recreate with anything else.
Brilliant and hilarious things that are a permanent part of my family's cherished canon.
https://raw.githubusercontent.com/thomasdavis/omega/master/s...
Had it hooked up to speech so you could just talk at it and it would talk back at you.
Gave incredible answers that ChatGPT just doesn't do at all.
I've been working on LLMs for creative tasks and believe a mix of chain of thought and injecting stochasticity (like instructing the LLM to use certain random letters pulled from an RNG in a certain way at certain points) can go a long way in terms of getting closer to human-like creativity
I am especially interested in gpt-3.5-turbo-instruct, as I think that the hype surrounding ChatGPT and "conversational LLMs" has sucked a lot of air out of what is possible with general instruct models. Being able to fine tune it will be phenomenal as well.
I do not really understand the efforts that went on behind the scenes to train GPT models on factual data. Did humans have to hand approve/decline responses to increase its score?
"America is 49 states" - decline
"America is 50 states" - approve
Is this how it worked at a simple overview? Do we know if they are working on adding the rest of 2021, then 2022, and eventually 2023? I know it can crawl the web with the Bing addon but, it's not the same.
I asked it about Maya Kowalski the other day. Sure it can condense a blog post or two, but it's not the same as having the intricacies as if it actually was trained/knew about the topic.
OpenAI will never fulfill the entire market, and their moat is in danger with every other company that has LLM cash flow.
They want to become the AWS of AI, but it's becoming clear they'll lose generative multimedia. They may see the LLM space become a race to the bottom as well.
There's still no news about the multi-modal gpt-4. I guess the image input is just too expensive to run or it's actually not as great as they hyped it.
https://help.openai.com/en/articles/7102672-how-can-i-access...
The decision of burying these extra information in a support article, not cool!
We already know they have a SOTA model that can turn images into latent space vectors without being some insane resource hog - in fact, they give it away to competitors like Stability. [0]
My guess is a limited set of people are using the GPT-4 with CLIP hybrid, but those use-cases are mostly trying to decipher pictures of text (which it would be very bad at), so they're working on that (or other use-case problems).
There was no question something happened in January with ChatGPT, weirdly would refuse to answer questions that were harmless but difficult(Give me a daily schedule of a stoic hedonist)
Every once in a while, I see redditors complain of it being nerfed.
Sometimes I go back to gpt3.5 and am mind boggled how much worse it is.
Makes me wonder if they keep increasing the version number while dumbing down the previous model.
With an API, being unreliable would be a deal-breaker. Looking forward to people fine-tuning LLMs with GPT4 API. I'd love it for medical purposes, I'm so worried of a future where the US medical cartels ban ChatGPT for medical purposes. At least with local models, we don't have to worry about regression.
https://chat.openai.com/share/04c1dbc0-4890-447f-b5a5-7b1bc5...
My guess is that they began to restrict ChatGPT because they can't sell that. They probably want to sell you CodeGPT or other products in the future so why would they give that away for free? ChatGPT is just a teaser.
Is there any actual evidences other than some user subjective experiences?
You can apparently have it be nice or smart, but not both.
Has anyone fared better? I might be doing something wrong but I can't see what that could possibly be.
We use it to generate automatic insights from survey data at a weekly cadence for Zigpoll (https://www.zigpoll.com). This makes getting an instant response unnecessary but still provides a lot of value to our customers.
> We recognize this is a significant change for developers using those older models. Winding down these models is not a decision we are making lightly. We will cover the financial cost of users re-embedding content with these new models. We will be in touch with impacted users over the coming days.
The API will make you wait up to 10 minutes, and then time out. What's worse, it will time out between their edge servers (cloudflare) and their internal servers, and the way OpenAI implemented their billing you will get a 4xx/5xx response code, but you will still get billed for the request and whatever the servers generated and you didn't get. That's borderline fraudulent.
Meanwhile, their status page will happily show all green, so don't believe that. It seems to be manually updated and does not reflect the truth.
Could it be that it works better in another region? Could it be just my region that is affected? Perhaps — but I won't know, because support is non-existent and hidden behind a moat. You need to jump through hoops and talk to bots, and then you eventually get a bot reply. That you can't respond to.
My support requests about being charged for data I didn't have a chance to get have been unanswered for more than 5 weeks now.
There is no way to contact OpenAI, no way to report problems, the API sometimes kind-of works, but mostly doesn't, and if you comment in the developer forums, you'll mostly get replies from apologists that explain that OpenAI is "growing quickly". I'd say you either provide a production paid API or you don't. At the moment, this looks very much like amateur hour, and charging for requests that were never fulfilled seems like a fraud to me.
So, consider carefully whether you want to build against all that.
Very sorry to hear about these issues, particularly the timeouts. Latency is top of mind for us and something we are continuing to push on. Does streaming work for your use case?
https://github.com/openai/openai-cookbook/blob/main/examples...
We definitely want to investigate these and the billing issues further. Would you consider emailing me your org ID and any request IDs (if you have them) at atty@openai.com?
Thank you for using the API, and really appreciate the honest feedback.
Borderline!? They're regularly charging customers for products they know weren't delivered. That sounds like straight-up fraud to me, no borderline about it.
https://www.reddit.com/r/ChatGPT/comments/14ruui2/comment/jq...
I've heard LLMs described as "setting money on fire" from people that work in the actually-running-these-things-in-prod industry. Ballpark numbers of $10-20/query in hardware costs. Right now Microsoft (through its OpenAI investment) and Google are subsidizing these costs, and I've heard it's costing Microsoft literally billions a year. But both companies are clearly betting on hardware or software breakthroughs to bring the cost down. If it doesn't come down there's a good chance that it'll remain more economical to pay someone in the Philippines or India to write all the stuff you would have ChatGPT write.
I’m pretty sure they tuned the Cloudflare WAF rules on GPT 3 and forgot to increase the request size limits when they added the bigger models with longer contest windows.
I too had an issue and put in a request. Took about 2.5 months to get a response, so 5 weeks you are almost half way there.
However, I'm virtually certain somethings wrong on your end, I've never seen a wait even close to that unless it was completely down. Also the thing about "small prompts"...it sounds to me like you're overflowing context, they're returning an error, and somethings retrying.
For example exponential back off is the first step, then adding retrying on timeouts (we use streaming and if there are 30 seconds in between getting data back we retry the whole request - rare but happens), then fixing anything else that pops up
It is possible to have a stable production app on top of it
Just fix your code and stop expecting OpenAI to hold your hand
I tried again just now and I got "Oops! We ran into an issue while authenticating you." but it works on chromium.
It's fraudulent, full stop. Maybe they're able to weasel out of it with credit card companies because you're buying "credits."
I suspect it was done this way out of pure incompetence; the OpenAI team handling the customer-facing infrastructure have a pretty poor history. Far as I know you still can't do something simple like change your email address.
as far as I know OpenAI only has one region, that is out in Texas.
even more hilariously, as far as I can tell, Azure OpenAI -also- only has one region.. cant imagine why
If you want better latency and sane billing you need to go through Azure OpenAI Services.
OpenAI also offers decreased latency under the Enterprise Agreement.
In our own tracking, the P99 isn't exactly great, but this is groundbreaking tech we're dealing with here, and our dissatisfaction with the high end of latency is well worth the value we get in our product: https://twitter.com/_cartermp/status/1674092825053655040/
I'm not as excited about this as I am about most new tech. I'm sure there will be cool uses eventually but right now it seems like the primary use is to cheat at exams and write bad articles, and to ask the kind of questions Google could have answered 5 years ago.
I do think it's cool that it can debug and review code though.
Recalibrate where you think AI provides "real" value
Unfortunately it's still too expensive and the completion speed is not as high as GPT-3.5 but I hope both problems will improve over time.
I had the silly idea of trying to change the signin method of my account. Which isn't possible. So I figured to just delete the account and create a new one with the correct signin method.
Turns out they don't delete anything. Both the email address and phone number are held hostage. As you try to create a new account, it will point out that those are in use. I can easily change my email address but not my phone number, I only have one.
I've contacted support 4 times, but it's just bot replies. There's entire Reddit threads full of us perma-banned potential customers. Money in hand, but permanently locked out.
What a ridiculous company.
But yeah whining on HN would be more productive
Things that GPT-4 would easily, and correctly, reason through in April/May it just doesn't do any longer.
I checked many of the links you posted and, although I am a fluid programmer in several languages, I lack the specific python background that many of these links seem to state as a requirement.
Would any kind soul put me in the direction of an easy solution to run LLMs, potentially in AWS, that answers questions about your own docs? (I use Confluence but I can happily export pages.)
Thanks a lot in advance!!!
All of them are open source minus the gpt part so you can get a feel for how it works.
It's also a shame that the API is so cut down, and removes all the good options that text completions had.
I only hope their competition are better.
nooooo they are deprecating the remnants of the base models
OpenAI even has a whole repository specifically for this - GPT-eval. No one uses it.
I'm not saying the theories are wrong. Maybe there is something behind the hunches that so many people seem to have about degradation. But there isn't _any_ proof. None. Whatsoever. And people are taking _internet comments_ as that proof instead? I mean, sure, it's easy to be cynical about companies in this day and age; which is why I would ultimately believe someone if they provided actual evidence. But, again - not a single ounce of proof has been provided in any one of these threads.
Furthermore, the lack of rigor being applied even with the various anecdotes is appalling.
Which version are you talking about? GPT-4 or GPT-3? Are you using the API or the web interface? Are you aware that output is non-deterministic? Are you aware that your own psychological biases will skew your opinions on the matter? One or more of these questions tend to go unanswered.
Just please, show me some robust proof. If you can't because you didn't think to; you _surely_ must realize that many people are building entire businesses on top of this tech and at least _one_ of them is running these types of evaluations. Furthermore, the model is state-of-the-art for research now as well and if you can _prove_ that there is degradation in the model that they are lying about (in a research paper), you will get citations. And yet, there is nothing. Zilch. Nada.
For anyone who wants to quickly try this out in VSCode for your custom prompts - https://marketplace.visualstudio.com/items?itemName=ppipada....
So need to pay to fine tune again?
500 {'error': {'message': 'Request failed due to server shutdown', 'type': 'server_error', 'param': None, 'code': None}} {'Date': 'Thu, 06 Jul 2023 20:48:07 GMT', 'Content-Type': 'application/json', 'Content-Length': '141', 'Connection': 'keep-alive', 'access-control-allow-origin': '*', 'openai-model': 'gpt-4-0613', 'openai-organization'Sadly, that niche is still better served by openAI.
Now, I am not superstitious, but...
GPT-4 is on a completely different level of consistency and actually listening to your system prompt than chagpt-3.5. It trails off much more rarely.
If only it wasn't so slow/expensive... (it really starts to hurt with large token counts).
What is the difference? Replacing evil with another evil.
This is just behemoths exchanging hands.