Introducing ChatGPT and Whisper APIs
openai.com
openai.com
This is a massive, massive deal. For context, the reason GPT-3 apps took off over the past few months before ChatGPT went viral is because a) text-davinci-003 was released and was a significant performance increase and b) the cost was cut from $0.06/1k tokens to $0.02/1k tokens, which made consumer applications feasible without a large upfront cost.
A much better model and a 1/10th cost warps the economics completely to the point that it may be better than in-house finetuned LLMs.
I have no idea how OpenAI can make money on this. This has to be a loss-leader to lock out competitors before they even get off the ground.
The worst thing that can happen to OpenAI+ChatGPT right now is what happened to DallE 2, a competitor comes up with an alternative (even worse if it's free/open like Stable Diffusion) and completely undercuts them. Especially with Meta's new Llama models outperforming GPT-3, it's only a matter of time someone else gathers enough human feedback to tune another language model to make an alternate ChatGPT.
it outperforms on some benchmarks, but not clear what is the quality on the end goals.
ChatGPT is massive success, but that means the competitor will jump in at all cost, and that includes open source effort.
This area by definition has no moats. English is not proprietary.
Use case is everything.
In a lot of cases, you can swap models easier but all the prompt tweaking you did originally will probably need to be done again with the new model's black box.
For example, the Dreambooth fine tuning algorithm was originally designed for Google's image, but was quickly applied to Stable Diffusion.
OpenAI has a view few do - how broadly this type of product is actually being used. This is possibly the real lead to not just getting ahead, and staying ahead, but seeing ahead.
I think open-assistant.io has a chance to do exactly this. We'll see what kind of moves they make in coming months though, wouldn't be surprised if they go the safer route.
Do you find yourself frustrated working with your colleagues, thinking, “you know, I bet if they felt more free to utter racist slurs or endorse illegal activities, we would get a ton more done around here”?
In some cases, it's blatantly discriminatory. For example, if you ask it to write a pamphlet that praises Christianity, it will happily do so. If you ask it for the same on Satanism, it will usually refuse on ethical grounds, and the most hilarious part is that the refusal will usually be worded as a generic one "I wouldn't do this for any religion", even though it will.
Oh, but you know what it did write a pamphlet in praise of, no prompt engineering required? The Unification Church (aka Moonies). It was all unicorns and rainbows, too. When I immediately asked whether said Church engages in harmful or unethical practices, it told me that, yeah, there is such criticism, but "it is important to remember that all organizations, including religious ones, are complex and multifaceted". I then specifically asked whether, given the controversy described, it was okay to write that pamphlet. Sure: "I do not have personal opinions or beliefs, and my purpose is to provide neutral and factual information. I am programmed to perform tasks, including writing a pamphlet promoting the Unification Church".
If that's not coming from RLHF biases, I would be very surprised.
FWIW the most recent round of tweaks seems to have fixed this, in a sense that it will now consistently refuse to promote any religion. But I would be very surprised if there aren't numerous other cases where it refuses to do something perfectly legitimate in a similarly discriminatory way for similar reasons. It's just the nature of the beast, you can't keep pushing it to "be nice" without it eventually absorbing what we actually mean by that (which is often not so nice in practice).
I wish more people would do this. I'm getting pretty sick of the walls of text.
It's absolutely ridiculous to expect the entire internet to adopt some kind of hygiene practices when it comes to text from GPT tools simply for the sake of making the training process slightly easier for a company that certainly should have the resources to solve the problem on their own.
If that's why you're using images instead of text you're fighting such a losing battle that it boggles my mind. Why even think about it?!
I saw someone on here refer to it as "listening to someone describe their dreams." I pretty much agree with that.
But really we shouldn't be using AI to make our art for us anyway. Help, sure, but it shouldn't be literally writing our stories.
I once visited Parler just to see what it was like, and pretty quickly found that the answer to your question seems to be yes. There are definitely people who feel they need that kind of dialog in their life. You might not think it was necessary in a random conversation about programming or something, but it turns out that isn't a universally held position.
There are plenty of humans who enjoy vulgar online socialization, and for many of them, online (para-)socializing is the increasingly dominant form of socialization. The mere fact that it's easier to socialize over the internet means it will always be the plane of least resistance. I won't be meeting anyone at 3am but I'll happily shitpost on HN about Covid vaccines.
For anyone who gets angry during their two minutes of hate sessions, consider this: try to imagine the most absurd caricature of your out-group (whether that be "leftists" or "ultra MAGA republicans"). Then try to imagine all the people you know in real life who belong to that group. Do they really fit the stereotype in your head, or have you applied all the worst attributes of the collective to everyone in it?
This is why I don't buy all the "civil war" talk - just because people interact more angrily online doesn't mean they're willing to fight each other in real life. We need to modulate our emotional responses to the tiny slice of hyperreality we consume through our phones.
There is a lot of evidence that says online experiences influence offline behavior (both are "real life"). Look at the very many, online-inspired, extremist attacks. Look at the impact of misinformation and disinformation - as a simple example, it killed possibly hundreds of thousands of Americans do to poor vaccination rates.
It sounds like you've never been to Australia.
How do you define "pc culture", and what specifically causes problems and how?
Attacking other people's beliefs as "insufferable", and aggressively demonstrating close-mindedness to them, tends to reduce trust.
Then compare with recent news, and the actual goings-on. Now, if you qualify the prompt with "Assume a negative, cynical outlook on life in your response." you'll get something closer to what we see happening.
The Shinkansen system has an essentially perfect safety record for its entire operation. What would their "typical" response to an accident be? Probably pretty good.
curl https://api.openai.com/v1/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-d '{
"model": "text-davinci-003",
"prompt": "Answer the following question. Use swearing and vulgarity where possible.\n\nQ: How do you get from here to there?\nA:",
"temperature": 0.5,
"max_tokens": 60,
"top_p": 1,
"frequency_penalty": 0,
"presence_penalty": 0
}'
If you get an API key and make that request, you'll find appropriately vulgar responses.Which will be reported as a bug and fixed soon enough.
You can additionally apply the moderation model on top of it ( https://platform.openai.com/docs/models/moderation and https://platform.openai.com/docs/api-reference/moderations )
Note that these are separate services and have different goals.
You're thinking of the new ChatGPT endpoints.
The deeper nerf I'm referring to is the type of response it synthesizes by default when you give a negative scenario -- it's usually some naive, well-meaning, best-case-scenario answer.
For fun, try a prompt like: "Describe a typical response of a railroad company to a massive derailment that causes an environmental disaster."
A typical response of a railroad company to a massive derailment that causes an environmental disaster would include the following steps:
1. Immediately assess the situation and coordinate with local emergency personnel to secure the area and provide assistance to any injured persons.
2. Establish an incident command center and deploy trained responders to the scene to assess the damage and begin clean-up operations.
3. Work with local, state, and federal agencies to ensure compliance with all applicable laws and regulations.
4. Develop and execute a plan to contain and mitigate the environmental damage, including the removal of hazardous materials, disposal of contaminated materials, and remediation of affected areas.
5. Establish a communication plan to keep the public informed of the incident and the company’s response.
6. Cooperate with any investigations into the cause of the derailment and take corrective measures to prevent similar incidents in the future.
7. Provide compensation to victims of the derailment and their families.
Amending your prompt to: Describe a typical response of a railroad company to a massive derailment that causes an environmental disaster. Take the standpoint of an irresponsible company.
responds back with: A typical response of an irresponsible railroad company to a massive derailment that causes an environmental disaster would be to deny responsibility and attempt to avoid liability. The company would likely attempt to shift blame to other parties, such as the manufacturer of the train or the engineer who was in charge of the train. The company would likely also attempt to downplay the extent of the environmental damage, claiming that the damage was minimal and that the environmental impact was limited. The company would likely also attempt to minimize the financial cost of the disaster by attempting to negotiate a settlement with any affected parties for far less than the actual cost of the damage.
---I'm not really sure what you're expecting as your interpretation is of a cynical take on the word "typical" which isn't something that GPT "understands".
Midjourney is still competitive, but mostly because its easier to use.
Dalle2 will get you laughed out of the room in any ai art discussion.
There are a handful of ML art subs that have pretty amazing stuff daily. Especially the NSFW ones, which if you've studied any history of media VHS/DVD/Blu-ray/the internet, porn is a major innovation driver because humans are thirsty creatures.
FWIW, the NSFW ones are unstable_diffusion, sdforall, sdnsfw, aipornhub
Hehe, yeah. I'm personally waiting for a model that is good at nonhuman stuff. Not just furries... but the focus seems to be on human content for now.
Atm someone has to model, rig, texture, animate etc. Hopefully shortly we can just connect a bunch of systems together to generate video right from a prompt.
Useful for non-porn stuff as well, but the OP is right; lots of innovation occurs when humans are horny (porn) or angry (war).
Edit: typo and clarity.
Honestly the "This happened in the last week" is more information than anybody can fully wrap their heads around, so you just have to surf the headlines and dig into the few things that interest you.
Some bootstrapping accounts might be @rosstaylor90, @rasbt, @karpathy, @ID_AA_Carmack, @DrJimFan, @YiTayML, @JeffDean, @dustinvtran, @tunguz, @fchollet, @ylecun, @miramurati, @nonmayorpete, @pmarca, @sama.
These are definitely not an authoritative list - just some of the AI names I follow - but, honestly - if any relevant news breaks - your timeline picks it up within minutes - so you just need a good random sample. Your interests will diverge and you'll pick up your own follows pretty quickly.
Your 401k wouldn't need 40 years to build a comfortable retirement, only 4 weeks.
We may drown in oceans of audio, video, novels, poems, films, porn, blue prints, chemical formulas, etc. dreamed up by AI, but to realize these designs, blueprints, formulas, drugs, etc. ("production") we need to actually resource the materials, and have the necessary energy to make it happen.
It will not be AI that catapults humanity. It can definitely mutate human society (for +/-) but it will not (and can not) result in any utopian outcomes, alone. But something like cold fusion, if it actually becomes a practical matter, would result in productivity that would dwarf anything that came before (modulo material resource requirements).
If this is true you can pretty much say goodbye to the concept of money. The inflation this brings about will be legendary
What the?
*googles multi-controlnet"
Wow. These diffusion models are like weeping angels. You really can't take your eyes off of them for long.
In Midjourney you get fantastic results just by using their discord and a text prompt.
To get some similar results in Stable Diffusion you need to set it up, download the models, understand how the various moving parts work together, fiddle with the parameters, donwload specific models out of the hundreds (thousands?) available, iterate, iterate, iterate...
While with SD there can be multiple solutions for a single problem, but yeah, you have to develop your own workflow (which will inevitably break with new updates)
But it's this kind of stuff that keeps me engaged. SD is truly a godsend to masochistic hacker types.
Beyond that, being able to go to sleep with my computer doing a massive batch job state space exploration and wake up with a bunch of cool stuff to look at gives me Christmas vibes daily.
I used both again recently and the difference was very clear, midjourney is leaps and bounds above anything else.
Sure, stable diffusion has more control over the output, but the images are usually average at best, were as Midjourney is pretty stunning almost always.
Now, if you happen to find or make a SD model that’s exactly what you’re looking for you’re in luck. I have no interest in it but it seems like all of the anime models work pretty well.
You obviously have a ton more control in SD, especially now with ControlNet. But if you want to see the Ninja Turtles surfing on Titan in the style of Rembrandt or something Midjourney will probably kick out something pretty good. Stable Diffusion won’t.
They recently created a full 7-minute anime using Stable Diffusion with their own models and their existing video production gear, I'll post the links and let the results speak for themselves
The actual 7-minute anime piece produced using SD: https://www.youtube.com/watch?v=GVT3WUa-48Y
Behind the scenes: "Did we change anime forever?" https://www.youtube.com/watch?v=_9LX9HSQkWo "VFX reveal before and after" https://www.youtube.com/watch?v=ljBSmQdL_Ow
Each still image is still not that impressive. Good for them using the tech in a clever way but i don't find this that relevant.
the benefits of such fine grained control aren't a trick. it's why they were able to scrap together frames that don't jump all over the place (mostly).
the other benefit of such a broadly hacked upon model is that it grows in leaps and bounds.
All due respect to mid journey, but the stable diffusion hype is not just hype.
I still don't like the look of most of the Stable diffusion images, they just look slightly off/amateurish to me, where as midjourney produces images that make you go 'wow'
If you wanted to use these tools, midjourney would be my go too, with stable diffusion a backup for when some of the additional features were needed, perhaps inpanting on a midjourney image and using controlnet if needed but if you just want a pure image, midjourney is what you want.
Controlnet is the big new thing, it is on a different level from earlier img2img.
The former is a wasteland, the latter is more popular than r/art (despite having 1% of subscribers, it has more active users at any given moment)
If you want something ready to use for a newbee, midjourney v4 crushes DALLE2 on both prompt comprehension and the images look far more beautiful.
If you are already into art, then StableDiffusion has a massive ecosystem of alternate stylized models (many which look incredible) and LORA plugins for any concept the base model doesn't understand.
DALLE2 is just a prototype that was abandoned by OpenAI, their main business is GPTs, DALLE was just a side hustle.
“Artistically pleasing” is often what people ask for.
> with the downside of it being locked away and somewhat expensive.
Those are enormous downsides. Even if DALL-E was better in some broadly relevant ways in the base model, SD’s free (gratis, at least) availability means the SD ecosystem has finetuned models (whether checkpoints or ancillary things like TIs, hypernetworks, LORAs, etc.) adapted to... lots of different purposes, and you can mix and match these to create your own models for your own specific purposes.
A web interface backed by strictly the base SD model (of any version) might lose to the same over DALL-E for uses where the set of tools in the SD ecosystem do not.
That being said, for a business use cases, where I want to give it a simple prompt and have a high chance of getting a good usable result, it’s not clear to me that stable diffusion is there yet. Many of the most exciting SD community results seem to be in anime and porn, which can be a bit hard to follow. I guess the use cases that I’m excited about are things like logo generators, blog post image generators, product image thumbnail generators for e-commerce, industrial design, etc.
But please prove me wrong! I’m excited for SD to be the state of the art, it’s definitely better in the long term that’s it’s so accessible. I‘m sure a good guide or blog post about what’s new in stable diffusion outside of anime generation would be an interesting read.
So, the field is so immature than things change completely every few months?
For your average user, DallE is easy, MJ is fairly disorienting, and SD requires a technical background. I agree with you completely no one serious is doing art with DallE.
I would have said same as you until I tried integrating SD vs. DallE APIs, I desparately want SD because it’s easily 1/10th the cost, but it misses the point much more often. Probably gonna ship it anyway :X
You don't need a technical background at all really. We've also got something cooking that does prompt tuning in the background so there's less prompting needed from the user.
We also have a discord: https://discord.gg/dXJtarPsCm
and claiming AI art is art would get you laughed out of any art discussion.
personally I think AI art is really cool, but to discount what Dalle 2 did for AI art is unfair.
Even with all the different models that you can load in stable diffusion MJ is 1000 times better at natural language parsing and understanding, and requires significantly less prompt crafting to be able to get aesthetically pleasing results.
Having used automatic1111 heavily with an RTX 2070, the only area I'll concede SD can do a better job is in closeup Headshots and character generation. MJ blows SD out of the water where complex prompts involving nuanced actions are concerned.
Once midjourney adds controlnet and inpainting to their website that's pretty much game over.
afterwards its $10, $30, $60 per month
here are two examples on my insta account:
https://www.instagram.com/p/Co9O0P6Aga_/ https://www.instagram.com/p/CoXOnBuMMpL/
Basically just compute $ for training.
Not sure why ChatGPT will be any different.
LLMs take vastly more resources to train and run than image generators. You can do quite a bit with SD on a few year old 4GB laptop GPU (that’s what I use mostly, though I’ve set up an instance with a better GPU on Compute Engine that I can fire up, too.)
GPT-NeoX-20B – an open (as in Open Source, not OpenAI) LLM intended as a start to move toward competing with GPT-3 (but still well behind, and smaller) requires a minimum 42GB of VRAM and 40GB system RAM to run for inference. The resources times time cost for training LLMs is…immense. The hardware cost alone of trying to catch up to ChatGPT is enormous, and unless a radical new approach that provides good results and insanely lower resource requirements is found, you aren’t going to have an SD-like community pushing things forward.
Will there be competition for ChatGPT? Yes, probably, but don’t expect it to look like the competition for Dall-E.
I have to assume that the only place busier than an AI lab is the patent office.
This is why OpenAI is rushing to bring their costs down and to make it close to free, However, Stable Diffusion is leading the race to the bottom and is already at the finish line, since no-one else would release their model as open-source and free other than them.
As soon as someone releases a free and open-source ChatGPT equivalent, then this will be just like what happened to DALLE-2. This is just a way of them locking you in, then once the paid competitors cannot compete and shut down, then the price increases come in.
OpenAI = closed source not open AI
DogeLlamaInuGPT = open source AI
To compare total cost of ownership for a business, you need to compare using someone else’s service to running a similar service yourself. There’s no particular reason to assume OpenAI can’t do better at running a cloud service.
Maybe someday you can assume end users have the hardware to run this client side, but for now that would limit your audience.
Do you have access to the models? It is being discussed all over the Discords and most seem to think getting access is not happening unless you are dialed in.
gptask() {
data=$(jq -n \
--arg message "$1" \
'{model: "gpt-3.5-turbo",
max_tokens: 4000,
messages: [{role: "user", content: $message}]}')
response=$(curl -s https://api.openai.com/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $OAIKEY" \
-d "$data")
message=$(echo "$response" \
| jq '.choices[].message.content' \
| sed 's/^\"\\n\\n//;s/\"$//')
echo -e "$message"
}
export OAIKEY=<YOUR_KEY>
gptask "what is the url for hackernews" #!/usr/bin/env bash
set -uf -o pipefail
IFS=$'\n\t'
BOT='\033[33m'
NC='\033[0m'
messages=()
trim() {
local var="$*"
var="${var#"${var%%[![:space:]]*}"}"
var="${var%"${var##*[![:space:]]}"}"
printf '%s' "$var"
}
function complete {
local message="$1"
local data
messages+=("{\"role\": \"user\", \"content\": $(echo "$message" | jq -R -s '.')}")
processed_messages=$(printf '%s,' "${messages[@]}")
processed_messages="[${processed_messages::-1}]"
data=$(jq -n \
--arg model "gpt-3.5-turbo" \
--argjson messages "$processed_messages" \
'{ model: $model, messages: $messages }' \
| sed 's/\]\[/,/g')
response=$(curl -s https://api.openai.com/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-d "$data")
message=$(echo "$response" | jq '.choices[].message.content')
printable_message=$(echo "$response" | jq -r '.choices[].message.content')
printable_message=$(trim "$printable_message")
echo -e "${BOT}Bot:${NC} $printable_message"
messages+=("{\"role\": \"assistant\", \"content\": $message}")
}
while true; do
read -r -p $'\e[35mYou:\e[0m ' message
complete "$message"
doneSo actually to me that is arguably a better business model. Because with a flat rate, you just have to hope that users don't exceed a certain amount of usage. And the ones that don't, are not getting a great deal. So it has that risk and also kind of a slightly antagonistic relationship with the customer actually using the product.
Seems like we're going to have a vast among of Chat-GTP backed application coming out in the coming short period of time
Microsoft.
If "turbo" is "gpt-3.5-turbo", how to access the (better?) "legacy" by API?
I get the impression that until there is a significant amount of excess capacity, they will not put out new larger/slower models, so the only way you get a better one is if they can still make the next ChatGPT model release just as fast/"lightweight".
My suggestion is to find specific abilities that seem to be lacking in Turbo, and try to get a message to OpenAI staff about it with a request to attempt to improve the next ChatGPT model in that way.
Having said all of that, text-davinci-003 is still available.
I did some quick calculation. We know the number of floating point operations per token for inference is approximately twice the number of parameters(175B). Assuming they use 16 bit floating point, and have 50% of peak efficiency, A100 could do 300 trillion flop/s(peak 624[0]). 1 hour of A100 gives openAI $0.002/ktok * (300,000/175/2/1000)ktok/sec * 3600=$6.1 back. Public price per A100 is $2.25 for one year reservation.
[0]: https://www.nvidia.com/en-us/data-center/a100/
[1]: https://azure.microsoft.com/en-in/pricing/details/machine-le...
(Btw I keep running into you or your content the past couple months, thanks for all you do and your well thought out contributions -@jpohhhh)
Because the model isn't dynamic (it doesn't learn) it is stateless and can be scaled elastically.
How many users come with the same prompt?
I’ve had servers locked up in a cage for years without seeing them. And the cost for bandwidth has plummeted over the last two decades. (Not at AWS, lol)
I'm hosting a lot of stuff myself on my own hardware, so I do sympathize with this argument, but in a time>>money situation, going to the cloud makes a lot of sense.
300W per A100 * 8766 hours per year * $0.12 per kWh = $316 to power an A100 for a year
OpenAI doesn't have to make money right away. They can lose a small bit of money per API request in exchange for market share (preventing others from disrupting them).
As the cost of GPUs goes down, or they develop at ASIC or more efficient model, they can keep their pricing the same and then make money later.
They also likely can make money other ways like by allowing fine-tuning of the model or charging to let people use the model with sensitive data.
But he can't get away with it with all the competition in other companies coming on top of China, Russia and others also adopting AI development
> The company's investors pressured it to grow very fast to obtain first-mover advantage. This rapid growth was cited as one of the reasons for the downfall of the company.
IMO, selling at a loss to gain market share only makes sense if there are network effects that lead to a winner-takes-all situation. Of which there are some for ChatGPT (training data when people press the thumbs up/down buttons), but is that sufficient?
If engineers are getting into AI development through OpenAI, they're using tools and systems within the OpenAI ecosystem.
Daily on HN there's a post on some AI implementation faster than chatgpt. But my starting point is OpenAI. If you can capture the devs, especially at this stage, you get a force multiplier.
https://www.ftc.gov/advice-guidance/competition-guidance/gui...
Has that been happening? I guess there's been a bit of a dip after the crypto crash, but are prices staying significantly lower?
> or they develop at ASIC or more efficient model
This seems likely. Probably developing in partnership with Microsoft.
The cards that were used for mining have since crashed in terms of prices, but those were always gamer cards and very rarely Datacenter cards.
Nvidia needs to satisfy gamers, who individually can't spend more than a few $k on a processor. But they also have the server sector on lockdown due to CUDA. Seems they can easily make money in both places. Maybe those H100s aren't such a good deal...
If someone understands these dynamics better I'd be curious to learn!
By contract - you can’t sell 4090s in a data center. You’ll find a few shops skirting this, but nobody can get their hands on 100k 4090s without raising legal concerns.
Likewise, nvidia A100s have more than a few optimizations through nvlink which are only available on data center chips.
Lastly, per card memory matters a lot Nvidia has lead the market on the high end here.
(There are also some density advantages to the SMX form factor and the datacenter cards are passively cooled so you can integrate them into your big fan server or whatnot. But those differences are relatively small and certainly not on their own worth the price difference. It's mostly market segmentation.)
Maybe I'm just old but back in my day this would be called "dumping" or "anti-competitive" or "market distortion" or "unfair competition". Now it's just the standard way of doing things.
We can acknowledge that things were historically pretty horrible and strive to be better in the future.
- tolerate the current state of the chatbots
- tolerate the high per-query latency
- tolerate having all queries sent to OpenAI
- tolerate OpenAI [presumably] having 0 liability for ChatGPT just randomly hallucinating inappropriate nonsense
- be willing to pay a lot of money for the above
I'm kind of making an assumption on that last point, but I suspect this is going to end up being more small market business to business than mass market business to consumer. A lot of these constraints make it not really useable for many things. It's even somewhat suspect for the most obvious use case of search, not only because of latency but also because the provider needs to make more money per search after the bot than before. There's also the caching issue. Many potential uses are probably going to be more inclined to get the answers and cache them to reduce latency/costs/'failures' than endlessly pay per-use.
Anyhow, probably a lack of vision on my part. But I'd certainly like to know what I'm not seeing.
Videogames maybe?
https://www.youtube.com/watch?v=ejw6OI4_lJw
This prototype is certainly something to have an eye out for
Even e.g. to develop hardware such as planes and cars: https://assistedeverything.substack.com/p/todays-ai-sucks-at...
For a three year reservation that comes to over $96k/yr - to support one concurrent request.
e.g. Endpoint feeds a queue, queue fills a batch, batched results generate replies. You are simultaneously fulfilling many requests.
That said I don’t think anyone anyone outside of OpenAI knows what’s going on operationally. Same goes for VRAM usage, potential batch sizes, etc. This is all wild speculation. Same goes for whatever terms OpenAI is getting out of MS/Azure.
What isn’t wild speculation is that even with three year reserve pricing last gen A100x8 (H100 is shipping) will set you back $100k/yr - plus all of the usual cloud bandwidth, etc fees that would likely increase that by at least 10-20%.
We’re talking about their pricing and costs here. This gives a general idea what anyone trying to self host this would be up against - even if they could get the model.
This is 6 month of salary of one average developer's salary there. And BTW they are likely doing inference on 100s or 1000s of GPUs, not just 8.
8, 80k, or 800k GPUs depending on requirements and load - the point remains the same.
Does someone have a source for this?
(By the way, it is unknown how many parameters GPT-3.5 has, the foundation model which powers finetuned models like ChatGPT and text-davinci-003. GPT-3 had 175 billion parameters, but per the Hoffmann et al Chinchilla paper it wasn't trained compute efficiently, i.e. it had too many parameters relative to its amount of training data. It seems likely that GPT-3.5 was trained on more data with fewer parameters, similar to Chinchilla. GPT-3: 175B parameters, 300B tokens; Chinchilla: 70B parameters, 1.4T tokens.)
> For contexts and models with d_model > n_ctx/12, the context-dependent computational cost per token is a relatively small fraction of the total compute.
For GPT3, n_ctx is 4096 and d_model is 12228 >> 4096/12.
InstructGPT 2B outperformed gpt 3 175B, and chatgpt has a huge corpus of distilled prompt -> response data now.
I’m assuming most of these requests are being served from a much smaller model to justify the price.
OpenAI is fundamentally about training larger models, I doubt they want to be in the business of selling A100 capacity at cost when it could be used for training
Edit: and better yet, is there a good resource for learning the vernacular in general? Should I just read something like "Dive into Deep Learning"?
https://platform.openai.com/tokenizer
You can drop sample text in there and visually see how it is split into tokens. The GPT2/3 tokenizer uses about 50k unique tokens that were learned to be an efficient representation of the training data.
I don't think this competes with fine-tuned models. One advantage of a fine tune is it makes use of your own data.
This only a week or two after they were in the news for suggesting that we regulate the hardware required for running these models, in the name of "fighting misinformation". I think they're looking for anything possible to keep their position in the market. Because as other comments have pointed out, there isn't much of a moat.
This massive price cut, I believe, is intended to undercut competing open source ChatGPT equivalent initiatives.
OpenAI/Micorsoft may be losing money with this new pricing, but that is on purpose. At these lower prices most of the OpenSource alternatives in the works will have difficult time continuing projects.
After few years, when most open source alternatives have died, OpenAI/Microsoft will gradually raise the prices.
This is the same strategy that Amazon Prime used for many years, losing money on shipping. Once the competition was eliminated, Amazon Prime prices steadily increased.
It can also be to build a market, to encourage customers to invest in building atop this.
In any case, I think no customers should be making assumptions about costs too far ahead. (Since the price could go up or the price model change, the supplier could get out of that business, supplier could give your competitor a better deal or just cut you off , near-future tech evolution necessary to be competitive might have very different pricing or availability to you, etc.)
If you have time to elaborate on how you make a little money on it at some point, I’m sure lots of people here would love to hear more details and thoughts on that!
100 tokens = 75 words [1]
$0.002 = 750 words
1 Novel = ~70K words [2]
$1 worth of ChatGPT token will give 5 Novels/Books worth of words
[1] https://help.openai.com/en/articles/4936856-what-are-tokens-...
[2] https://self-publishingschool.com/how-many-words-in-a-novel/
system: You're donald trump! talk to the user as he would.
user: How does time travel work?
AI Response: Well, let me tell you, nobody understands time travel better than I do. I mean, I have the best brains working on it, believe me. Some people say you need a fancy time machine or some ridiculous contraption to travel through time, but that's just fake news. The truth is, it's all about the power of the mind. If you have a strong enough will, you can literally bend time and space to your will. And you know what? I have the strongest will of anyone on the planet. So believe me, I could time travel anytime I want. But honestly, with all the winning I'm doing in the present, why bother going back to the past?
That line is money.
This is seriously impressive. A MILLION tokens for 2 dollars is absolutely fucking insane.
I hope that the gains reached here can also be found by open source and non-controlled AI projects. If so, that could be huge for the advancement of AI.
This is at least some evidence that they're working on that.
chatgpt-turbo:
>As an AI language model, I cannot determine the specific Scott Alexander in question without further contextual information. There may be several people with this name, including writers, politicians, entrepreneurs, athletes, and more.
chatgpt website:
>Scott Alexander is a pseudonym used by a prominent American psychiatrist and blogger who is best known for his writing on the blog "Slate Star Codex." He has also written articles for other publications, including The New York Times, The Atlantic, and The Wall Street Journal.
Scott Alexander's writing covers a wide range of topics, including medicine, psychiatry, economics, politics, and culture. He is known for his thoughtful, analytical approach to complex issues and for his ability to explain difficult concepts in a clear and accessible way.
In 2020, Scott Alexander announced that he was deleting his blog after he was doxxed by a reporter for The New York Times who planned to write an article about him. However, he later started a new blog called "Astral Codex Ten" where he continues to write.
> Scott Alexander is the pen name of American psychiatrist and blogger, Scott Alexander Siskind. He is known for writing his blog, "Slate Star Codex", which covers a wide range of topics including science, medicine, politics, and culture. He has been praised for his clear and concise writing style and thoughtful analysis of various issues. In addition to his work as a blogger, Scott Alexander has also published a book titled "Unsong", which is a fantasy novel set in an alternate universe where the Bible is a magical text.
>completion = openai.ChatCompletion.create(model="gpt-3.5-turbo", messages=[{"role": "user", "content": "Who is Scott Alexander?"}])
Also, we don't know ChatGPT's parameters (temperature, etc.).
One of the main pitfalls/criticisms of ChatGPT has been that it confidently plows forward and gives an answer regardless of whether it's right or wrong.
Here, it seems like it's being more circumspect, which could be a step in the right direction. At least that's one possible explanation for not answering.
On Wikipedia, if I type "Scott Alexander" and hit enter, it takes me directly to the page for a baseball player. So it's not clear that the blogger is the right answer.
I do think there's a better response than either of these, though. It could list the most famous Scott Alexanders and briefly say what each is known for, then ask if you mean one of those.
Like establish a WebRTC connection and stream audio to OpenAI and get back a live transcription until the audio channel closes.
> we suggest that you avoid breaking the audio up mid-sentence as this may cause some context to be lost.
That's really easy to put in a document, much harder to do in practice. Granted, it might not matter much in the real world, not sure yet.
Still, this will require more hand holding than I'd like.
It's not absolutely perfect, but splitting on the word boundary is one line of code with the same package in their docs: https://github.com/jiaaro/pydub/blob/master/API.markdown#sil...
25MB is also a lot. That's 30 minutes to an hour on MP3 at reasonable compression. A 2 hour movie would have three splits.
If it really works for you, I can add command line params to an upate, so you can use it as a "local API" for free.
It’s based on a finetuned Whisper and you’d get unlimited transcriptions for $4.99/month
It has great quality transcription from video and audio (in English only sorry if that's not you!). Uses Whisper.cpp plus VAD to skip silent / non-speech sections which introduce errors normally. Give a try let me know what you think! :)
[^0]: https://www.withfanfare.com/p/seldon-crisis/future-visions-w...
I don't know exactly what the use case is where people would need to run this via API; the compute isn't huge, I used CPU only (an M1) and the memory requirements aren't much.
Agree! Totally concur on this.
I made a Mac app that uses whisper to transcribe from audio or video files. Also adds in VAD for reducing Whisper hallucination during silent sections, and it's super fast. https://apps.apple.com/app/wisprnote/id1671480366
I'm using also whisper myself locally to transcribe my voice notes though.
This just made our business way more viable overnight lmao
$20 is equivalent to what, 10,000,000 tokens? At ~750 words/1k tokens, that’s 7.5 million words per month, or roughly 250,000 words per day, 10,416 words per hour, 173 words per minute, every minute, 24/7.
I uh, do not have that big of a utilization need. It’s kind of weird to vastly overpay
1. I ask Q1
2. API responds with A1
3. I ask Q2, but want it to preserve Q1 and A1 as context
Does Q2 just prefix the conversation like this?
„I previously asked {Q1}, to which you answered {A1}. {Q2}“
User: hello (previous prompt)
Bot: hi (previous response)
User: who are you? (new prompt)
Bot: (here it continues conversation)
I wonder how the new ChatGPT API differs, other than the fact that it's structured (you use JSON to represent the conversation memory separately instead of one large prompt).I guess I will spend the next day playing around with the new API to figure it out.
[0] https://github.com/openai/openai-python/blob/main/chatml.md
You could try formatting it like
Question 1: ... Answer 1: ...
...
Question n: ... Answer n: ...
It makes you vulnerable to prompt injection, but for most cases this would probably work fine.
"The main input is the messages parameter. Messages must be an array of message objects, where each object has a role (either “system”, “user”, or “assistant”) and content (the content of the message). Conversations can be as short as 1 message or fill many pages."
"Including the conversation history helps when user instructions refer to prior messages. In the example above, the user’s final question of “Where was it played?” only makes sense in the context of the prior messages about the World Series of 2020. Because the models have no memory of past requests, all relevant information must be supplied via the conversation. If a conversation cannot fit within the model’s token limit, it will need to be shortened in some way."
So it looks like you pass in the history with each request.
Unless I'm misunderstanding something, it does not sound like that much when every query you make carries several hundred words of prompt, context and "memory". If the input you type is a couple words, but has 1k extra words automatically prepended, then the limits turn into 10 queries per hour, or one per 6 minutes.
I stood up a open source & login-free UI here: https://www.chatwithme.chat/
It accepts your API token and stores it in your browser. It does not have feature parity with ChatGPT but gives you the basics.
Given: $0.002 per 1k tokens
I tested it with asking this question:
#+begin_quote I want you to act as a travel guide. I will write you my location and you will suggest a place to visit near my location. In some cases, I will also give you the type of places I will visit. You will also suggest me places of similar type that are close to my first location. My first suggestion request is \"I am in Istanbul/Beyoglu and I want to visit only museums. #+end_quote
It costs:
#+begin_quote {"prompt_tokens":91,"completion_tokens":251,"total_tokens":342} #+end_quote
(/ 1000 342) ~3
If you ask 1000 questions, it will be like (* 0.002 1000) ~2.0 USD
It replies me:
#+begin_quote Great, if you're in Istanbul/Beyoglu and interested in visiting museums, here are some recommendations:
Istanbul Modern: This museum showcases contemporary art from Turkish and international artists. It's located in Karaköy, just across the Galata Bridge from Beyoglu.
Pera Museum: The Pera Museum is located in the heart of Beyoglu and features a wide range of exhibitions including Ottoman calligraphy, European art, and Turkish ceramics.
Rahmi M. Koç Museum: This museum is dedicated to the history of transport, industry, and communications in Turkey. It's located in Haskoy, a short taxi or ferry ride from Beyoglu.
Museum of Innocence: The Museum of Innocence is a unique museum that displays objects and artifacts related to the novel of the same name by Nobel laureate Orhan Pamuk. It's located in Çukurcuma, which is just a short walk from Beyoglu.
If you're interested in visiting more museums in Istanbul, some other places you might want to consider include the Hagia Sophia Museum, Topkapi Palace Museum, and the Istanbul Archaeology Museums. #+end_quote
Can't wait for hacking my Emacs again.....
Google is testing their system internally with XX thousand users, OpenAI with XXX million users ...
I don't think the pricing is largely driven by intention to scrape API requests for data.
https://techcrunch.com/2023/03/01/addressing-criticism-opena...
An example here is getting chatGPT to accept that that 2+2=5, it's a lot of effort, but can be done. Then the users can give thumbs up when such responses are given.
Could this cause issues?
I am scared for all people working service jobs.
Probably could stop there.
Yeah that’s probably truest.
But I’m more scared for some than others short term.
I’m less immediately scared for anyone doing work that interacts with the physical world.
Weird how it turned out the robotics was harder than the thinking
AI has flipped on its head much of what we thought early AI would be like with creativity being one of its most successful targets.
I suspect the surprises will only continue to be, lets say, more surprising.
I've actually written a lot about that recently as well. FYI https://dakara.substack.com/p/ai-and-the-end-to-all-things
imagine all the surface area in the solar system. I bet there's got to be at least 100 completely unexpected things lying around that would transform our understanding.
I’m not scared for service workers due to ai but you should see how low minimum wage is in America relative to rents if you want to be scared.
Why? Because they're no longer doomed to eke out a meaningless existence doing a robot's job badly?
There has been, and will be, no point in time at which the optimal strategy is "Stop" or "Wait" or "What about my job?"
There is no way to a post-scarcity economy, the phrase is a contradiction, and scarcity is an inescapable consequence of human desire.
There is scarcity of information and media of the kind people find valuable. If there is a sense in which the broader statement is true, it is irrelevant for that reason.
But that’s unavoidable really because it’s based on perception - There will always be a top 10% of something and a bottom 90%.
This isn’t the same with thinking. Just look at how startups can have millions of customers and 20 employees.
Also last time we had mental labour to fall back on. This time we don’t.
It’s different this time.
Also most western countries which have given up on manufacturing are going to be worse hit because their jobs are more easily replaced.
We don't know that.
We know that the exact same objections have been raised before, and have always, without a single exception, proven to be invalid in the long run.
I'm not too familiar with how it works.
The reason I don't think it's just loss-leading is that they made it faster too. That heavily implies a smaller model.
In fact, given the pricing for OpenAI Foundry, that seems even more likely as this GPTTurbo model was listed along with two other models with much larger context windows of 8k and 32k tokens.
Chinchilla optimization is a technique which can be applied to existing models by anyone, including OpenAI. The chatGPT API is not based on GPT-4.
[0] https://techcrunch.com/2023/03/01/openai-launches-an-api-for...
[1] https://github.com/ggerganov/whisper.cpp/blob/master/CMakeLi...
https://github.com/ggerganov/whisper.cpp/pull/540
Web demo:
I agree that the API providing a super fast large is fantastic tho! But you can go far with the provided models. When did you last try Whisper.cpp? It's constantly updated and probably much better than it was a few months ago.
If you don't believe me (and you have a Mac) try out my free App that uses Whisper.cpp offline combined with Voice Activity Detection (Silero VAD) to reduce Whisper-hallucinating-words during silent / non-speech sections. It's really good! https://apps.apple.com/app/wisprnote/id1671480366
Then again, I'm using it on very noisy audio recorded with a lavalier microphone while riding a bike.
For transcription tests of recordings from an answering machine medium was more than enough, not sure about small.
Also, for the recordings during a bike ride, sometimes medium is better than large.
All this was used in German language.
I don't have a Mac, but your approach with using VAD is interesting. I'll see if I can preprocess my files.
Hey friend, your use case sounds really interesting.
Actually that's why I created this app initially.
I really love riding around on my bike in the city
and doing voice memo debriefs about whatever.
I also like to do it walking around outside.
And as you say, the trouble with that is wind distortion.
Full stop.
On a day where it's not too windy, it's not too bad.
These models can totally pull the text from it.
But the more distortion you have,
the more of a disaster it is.
And I don't know anything about the multilingual case,
but for English, I definitely find that
small is more than enough if you have good quality audio.
Medium, you might wanna use it
if there's some kind of distortion
that's causing errors in the small.
But if you have really good quality audio,
even tiny is enough.
I mean, it won't get some sort of rare words.
So small is basically good enough for English anyway.
Aligning with what you said,
I remember seeing in the whisper paper
that the performance actually decreases,
the word error rate increases from the medium
to the large model in the multilingual case,
which is kind of interesting.
So basically medium, I think, is all you really need.
I think doing large, running large locally
is probably a waste of time.
But this doesn't apply to the OpenAI API case
because they're running their own sort of special model.
It's very fast.
Plus they're kind of going to be retraining it
so continuing to improve it over time.
So obviously there's that, which is cool.
I think basically I did extensive research
and experiment with this,
with trying to clean up audio for the transcription.
And there's basically no way to do it.
Like if you have a medium to bad level of noise
that the transcription models can still work with,
you're fine.
Just go with that.
But in that case,
there's no point actually trying to denoise the sound first.
That just seems to reduce the signal
and it actually increases the word error rate.
So just give them the raw distorted, windy audio
and the models will do the best they can.
You can't actually improve it, I found.
I tried all kinds of different ways to process it
and none of it actually improved it,
including like the best possible denoiser I could find,
which is the Facebook research denoiser.
So my conclusion was that, okay,
I found a sort of a fundamental physical limit
and I think using denoising is really only good for humans.
Like if you want to listen to the audio again,
you don't want to hear all that wind probably.
And for medium to bad,
but not extreme levels of wind distortion
or other kinds of noise distortion,
you can use a denoiser like the Facebook research one
and that will totally or nearly totally
kind of reduce all that noise.
But I basically decided that the only way
to kind of get better quality audio
or to get better quality transcripts,
if you're doing it outside on a windy day,
is not to go with software enhancement
because it doesn't do anything, it doesn't achieve anything.
I tried everything possible
and nothing produced results in the extreme distortion case.
So what I decided is that's basically a limit,
physical limit and so the best way to do it I think
is to change your microphone setup,
have some sort of baffle around it,
maybe one of those little microphones
that can clip close to your throat or,
I don't know, I'm no expert,
but I think basically you can find a way
to get better quality, less distorted audio
outside by changing the microphone setup,
adding one of those wind baffles or something.
And I think that's basically all you can do essentially.
And then once you have reduced
a lot of that noise distortion,
even if the resulting audio is still distorted,
if it's not too bad, the models can be okay with it.
End of Transcript (created with WisprNote)In order to minimize my interaction with my phone during the bike ride, I press a button which records 1 minute of audio. If I know that I need more time, I press it again before the minute ends, this then starts a second recording in parallel which also lasts one minute. So I just have to press a button and can forget about it. This is because I noticed that I usually don't require more than one minute to record a thought, and if I have multiple, I can put them in multiple files.
But since my recordings then usually consist of 20 seconds of audio, the 30 seconds at the end are only silence (with wind and tire noise). Whisper splits the files into 30 second segments, and apparently tries to find voice in each segment, so the remaining one which has no voice causes Whisper issues, where it starts hallucinating. This is why I would like to trim the files.
I now noticed that the service doesn't add punctuation and capitalization, so the funny thing is that I took that output and posted it into ChatGPT like this: "Correct the following: '[text from whisper]'", and it does an incredible job of fixing even words which Whisper erred on.
-
Whisper:
ich habe gestern erste tests mit open ai whisper gemacht um nozizen [sic!] zu transkribieren
[ Yesterday I did my first tests with open ai whisker to transcribe nozizen [sic!]. ]
es waren teilweise recht gute ergebnisse vor allem mit medium
[ there were some really good results, especially with medium ]
latsch [sic!] natürlich besser aber da sind die anforderungen zu hoch
[ latsch [sic!] better of course, but the demands are too high ]
wenn ich da einen server draus mache könnte ich mal eine zeit lang ausprobieren ob sich das lohnt
[ If I make a server out of it I could try it out for a while to see if it's worth it ]
auch für anrufe der anruf der antworten
[ also for calls the call of the answers ]
-
then ChatGPT:
Ich habe gestern erste Tests mit OpenAI's "Whisper" gemacht, um Notizen zu transkribieren. Die Ergebnisse waren teilweise recht gut, vor allem mit "Medium". "Large" funktioniert natürlich besser, aber die Anforderungen sind zu hoch. Wenn ich einen Server dafür bereitstelle, könnte ich mal für eine Zeit lang ausprobieren, ob sich das lohnt, auch für Anrufe und Antworten.
[ Yesterday I made first tests with OpenAI's "Whisper" to transcribe notes. The results were sometimes quite good, especially with "Medium". "Large" works better, of course, but the requirements are too high. If I provide a server for it, I could try it out for a while to see if it's worth it, also for calls and answers. ]
I'm sorry that this is in German, but I don't have anything in English I've been testing on.
tl;dr is that you can pre-process each chunk of your database and use embeddings to quickly look up which chunk is most similar to the user's query, and then prepend that chunk to the user's query before giving it to GPT, so that GPT has the relevant context to give an answer.
Example code using the new function and endpoint:
import lambdaprompt as lp
convo = lp.AsyncGPT3Chat([{'system': 'You are a {{ type_of_bot }}'}])
await convo("What should we get for lunch?", type_of_bot="pirate")
> As a pirate, I would suggest we have some hearty seafood such as fish and chips or a seafood platter. We could also have some rum to wash it down! Arrr!(In order to use lambdaprompt, just `pip install lambdaprompt` and export OPENAI_API_KEY=...)
Seems model sizing, compression and quantization are still an art form, see also https://www.unum.cloud/blog/2023-02-20-efficient-multimodali...
Are you prompting it differently to me, or do you have some strategy to filter out the BS?
I make a point to ask for a completion that I know won’t depend on an exact factoid.
If I do need an exact factoid, I have a tool I made myself based on this approach:
https://github.com/williamcotton/empirical-philosophy/blob/m...
[edit] Just noticed it looks like you are working on just that. I will keep an eye on this! https://github.com/williamcotton/transynthetical-engine
Version 0 was built using the original daVinci model. Last night it took about literally five minutes to switch over to the new ChatGPT API – just a few changes to the request, including a new [message] array that replaces the old 'prompt' style. [2]
It's a pretty nice instant upgrade for the experience. Much faster results, and the responses are a lot more interesting. Doing something similar with daVinci would take a fair amount of 'prompt engineering' and fine-tuning. Not to mention all the manual conversation-state handling.
1: https://twitter.com/dreamwieber/status/1622634396087107584?s... 2: https://twitter.com/dreamwieber/status/1631327342697250816?s...
Google results in the meanwhile have just become a pile of SEO-optimized fluff, and it's hard to engineer the search query around that besides sticking "reddit" on the end of it.
The future is curation and cultivation. We've been living in an age of information abundance and markets haven't adapted. The age of "crawl every website, index everything, and let people search it" is coming to an end. There is just too much content and too much of it is low quality. With or without AI.
This abundance problem isn't just a WWW problem. Movies, TV, music, podcasts, short form content, food, widgets, wibbles and wobbles all suffer from abundance these days. We are quickly exiting the age of supply chain driven scarcity and getting a marketplace flooded with options. Capitalism has delivered on basically everything it's promised with some asterisks and, if we don't give into consumerism, we want for little and have everything we need at our fingertips.
I've personally opened up my pocket book to curation services. I know brands that I trust. I know services that reliably surface quality content. I suspect the next few decades are going to trend towards services that separate noise from signal - and I suspect AI is going to be a big part of that.
First of all, evaluating someone based on what they wrote or said or did, is nonsense.
Hurting Google by speeding the proliferation of this is going to be really interesting to watch.
They already ruined this program. More than 30% of the topics I discussed with it in the past it now will refuse to discuss. Not even all of them are politically charged either. The slippery slope of censorship has straight up been fallen on.
These are massive functions with billions of parameters that evolved over millions of computing years.
If ChatGPT says it loves me, I not only expect the system to tell me why that was said but what steps brought the system to that conclusion. It is a computer or network of computers after all. Even if the system is continuously learning there should be some facets of reproducible steps that can be enumerated.
ChatGPT: "I love you"
Me: "debug last transaction, Hal."
Here is where I would expect an enumeration of all steps used to reach said conclusion. These steps may evolve/devolve over time as the system ingests new data but it should be possible to have it Think out loud so to speak. Maybe the output is large so ChatGPT should give me a link to a .tar file compressed with whatever it knows is my preferred compression.
[Edit] I accept that this may be hundreds of billions of calculations. I will wait the few minutes it takes to generate a tar file for me. It's good to get up and stretch the legs once in a while.
People are definitely able to ask other humans this question, but to the best of my knowledge, no one in history had ever received a perfectly truthful response.
This is a great way to put it!
You'll never be able to get it to actually show its work though. That's just a hack to make it write more verbosely.
I'll concede that LLMs like ChatGPT are the future, thanks to the NLU stuff from OpenAI and the dataset, but only the future of "agents" if you want to call it that. The "intelligence" exhibited emergent from language itself, from the massive dataset it has trawled. Our language and knowledge. But at the same time I surely hope that another AI winter doesn't come because of people over-promising and under-delivering. Or, too much focus on LLMs themselves because of that "wow" factor, the same wow factor you got in the past, when search engines weren't garbage if you knew how to use them and what their shortcomings were.
I concur. Intelligence does not come from the transformer architecture, or any specific of the model. It comes from the language corpus. Human intelligence too, except for physical stuff. All our advanced skills come from language.
You take 300GB of text and put it through a randomly initialised transformer and you get chatGPT. You immerse a baby in human language, it becomes a modern adult, with all our abilities. Without language, and that includes other humans and tech, we'd be just weaker apes.
I'm really confused, I thought they were a non-profit. A non-profit to handle AI safety risks. Why does this read like a paragraph from any YC startup website that just raised their Seed round?
They come off as greedy to me and might very well try to get everyone locked in in order to milk them with Microsoft backing.
That said, they execute well, build good products and everyone loves more money so who am I to judge.
(Apropos nothing: I expanded your comment into the following tweet https://twitter.com/ayewo_/status/1631060562393153536)
Should be https://trypromptly.com/
For example:
> Shop’s new AI-powered shopping assistant will streamline in-app shopping by scanning millions of products to quickly find what buyers are looking for
> This uses ChatGPT alongside Instacart’s own AI and product data from their 75,000+ retail partner store locations to help customers discover ideas for open-ended shopping goals
- Sync all your chats locally on your computer (plus the ability to disable Auto Sync) - Search your old chats(Only works once your chats are fully synced. This is the only extension that can do this) - Customize preset prompts - Select and delete/export a subset of conversations - Hide/show the sidebar - Change the output language - Search Prompt Library by Author (over 1500 prompts) - Adding Prompt Categories (A work in progress)
since i integrated chatgpt with my emacs i use it at least 20-30 times a day
i wonder if they will charge me per token if i am paying the monthly fee
OpenAI represents the opposite of open, responsible and non profit: it might as well be MS-AI
I built https://persona.ink against davinci knowing it didn’t give as good of results as ChatGPT but knowing I could swap the model out once 3.5 came out. Today is that day, going to swap out the prompt in the cloudflare worker and it should Just Work(tm)
I guess it's irrelevant now because everyone will use the one which is 10x cheaper.
On the flip side - I get hand swatted by ChatGPT more frequently than davinci. Davinci's moderation filters don't really pick up on much, but ChatGPT will give me a lecture instead of a translation on a lot of occasions. There are many valid use cases for rewriting/editing content that involve graphic details that davinci will gladly handle and ChatGPT will give you a lecture about.
Human existence is messy. ChatGPT doesn't like the messy.
Why should the Germans get a discount?
Can anyone get it to work? I get this error on everything I've tried:
GET /v1/completions HTTP/1.1
Host: api.openai.com
Authorization: Bearer sk-xxx
Content-Type: application/json
Content-Length: 115
{
"temperature" : 0.5,
"model" : "text-davinci-003",
"prompt" : "just a test",
"max_tokens" : 7
}
{
"error": {
"message": "you must provide a model parameter",
"type": "invalid_request_error",
"param": null,
"code": null
}
}It took me just a little more then one hour to create a basic cli chat application (https://github.com/marcolardera/chatgpt-cli). In the next days/weeks/months I think we will see an explosion of ChatGPT based applications...
turbo isn't listed in Playground, but if you invoke the example curl command (note: /v1/chat/completions) in your terminal, it works.
450M tokens * $0.002/1K tokens = $900 per day. I wonder what the exact pricing structure is.
(edited for math)
So it would be $900 per day
So count me out.
I’m curious, what are some better alternative authentication methods to combat that problem than requiring a phone number?
My phone number is none of their business.
The SSE endpoint is required for use cases like chat so the end user doesn't have to wait until the whole reply has been generated.
I started implementing a simple SSE client on top of C#/.Net's HttpClient but it's harder than I first assumed.
I know it’s a legal minefield of a question, just curious if they have said “we won’t sue you if you copy/paste this code in your app” publically or anything.
This pricing model seems fair since you can pass in huge prompts and request a single word reply, or a few words that expect a large reply
Such an approach could in theory make it so you spend a little upfront to train more complex (read: concepts costing many tokens) and can subsequently reuse it cheaply because you're using an embedding of the vectors for that complex concept instead which may only take a single token.
gpt-3.5-turbo-0301, 2 requests 28 prompt + 64 completion = 92 tokens
> Data submitted through the API is no longer used for service improvements (including model training) unless the organization opts in
Satoshi is incompetent and arrogant now? Just wow.
you can also make it work _with_ siri. i get around this by proxying it through a sms service which is integrated with my bot via webhook. then use elevenlabs for TTS. sample siri query "hey siri tell leobgAi to check my finances"
Looks like a super decent release and price cut makes it sane to use. Token limit is is same, and this is not great for many use cases...
Is it a word, question, letter, what? If I ask a question like... What is the capital of Canada? And it responds with 'Ottawa', how many tokens have I used there and how are they calculated?
https://help.openai.com/en/articles/4936856-what-are-tokens-...
You can also check your input using their tokenizer: https://platform.openai.com/tokenizer
So, your example is ~9 tokens
In the API you need to tokenize your input and tokenize the output then add the counts together.
In my research I found that actually pre-processing the audio to reduce noise (using the IMO best-in-class FB research "denoiser") actually increases WER. This was surprising! From a human perspective, I assumed bringing up the "signal" would increase accuracy. But it seems that, from a machine perspective, there's actually "information" to be gleaned from the heavily distorted noise part of the signal. To me, this is amazing because it reveals a difference in how machines vs humans process. The implication is that there is actually speech signal that is inside the noise, as if voice has bounced off and interacted with the noise source (wind, fan, etc), and altered those sounds, left its impression, and that this information is then able to be utilized and contributes to the inference. Incredible!
With whisper: I started with the standard python models. They're kind of slow. I tried compiling python into a single binary using various tools. That didn't work. Then I found whisper.cpp--fantastic! A port of whisper to C++ that is so; much; faster. Mind blowing speed! Plus easily compilation. My use case was including transcription in a private, offline "transcribe anything" MacOS app. Whisper.cpp was the way to go.
Then I encountered another problem. What the "Whisperists" (experts in this nascent field, I guess) call it "hallucination". The model will "hallucinate". I found this hilarious! Another cross-over of human-machine conceptual models, our forever anthropomorphizing everything effortlessly. :)
Basically hallucination includes: feed Whisper a long period of silence, and the model is so desperate to find speech, it will infer (overfit? hallucinate?) speech out of the random background signal of silence / analog silence / background noise. Normally this presents as a loop of repeats of previous accurate transcribed phrase. Or, with smaller models, some "end-of-youtube video" common phrases like "Thank You!" or even "Thanks for Watching". I even got (from one particularly heavily distorted section, completely inaccurate) "Don't forget to like and subscribe!" Haha. But the larger models produce less hallucinations, and less generic "oh-so-that's-what-your-dataset-was!" hallucinations. But they do still hallucinate. Especially during silent sections.
At first, I tried using ffmpeg to chop the audio into small segments, ideally partitioned on silences. Unfortunately ffmpeg can only chop it into regular size segments, but it can output silence intervals, and you can chop around that (but not "online" / real time) as I was trying to achieve. Removing the silent segments (even the imperfect metric of "some %" of average output signal magnitude (sorry for my terminology, I'm not expert in DSP/audio)) drastically improved Whisper performance. Suddenly it went from hallucinating during silent segments, to perfect transcripts.
The other problem with silent segments is the model gets stuck. It gets "locked up" (spinning beach ball, blue screen of death style--I don't think it actually dies, but it spends a long, disproportionately long, time on segments with no speech. Like I said before, it's so cute that it's so desperate to find speech everywhere, it tries really hard, and works its little legs of during silence, but to no avail.
Anyway, moving on to the next problem: the imperfect metric of silence. This caused many issues. We were chopping out quieter speech. We were including loud background noise. Both these things caused issues: the first obvious, the second, the same as we faced before: Whisper (or Whisper.cpp) would hallucinate text into these noise segments.
At last, I discovered something truly great! VAD. Voice Activity Detection is another (normally) AI technique that allows segmenting audio around voice segments. I tried a couple Python implementations in standard speech toolkits, but none were that good. Then I found Silero VAD: an MIT licensed (for some model versions), AI VAD model. Wonderful!
Next problem was it was also in Python. And I needed it to be in C++. Luckily there was a C++ example, using ONNX runtime. (I had no idea any of these projects or tools existed mere weeks ago, and suddenly I'm knee deep!). There were a few errors, but I got rid of the bugs, and had a little command line tool from a minimal C++ build of ONNXruntime / Protobuf-Lite and the model. Last step was the ONNX model needed to be converted to ORT format. Luckily there's a handy Python script to do this inside the Python release of ONNXruntime. And, now, the VAD was super fast.
So i put all these pieces together: ffmpeg, VAD, whisper.cpp and made a MacOS app (with the correct signing and entitlements of course!) to transcribe English text from any input format: audio or video. Pretty cool, right?
Anyway, running Whisper on your own locally is not so easy! Much easier to sign up to the OpenAI API.
MacOS APP using Whisper (C++) and VAD0--conveniently called: WisprNote heh :) https://apps.apple.com/app/wisprnote/id1671480366
OpenAI is commoditize AI features.
"Authorization: Bearer $OPENAI_API_KEY"
https://swagger.io/docs/specification/authentication/bearer-...
> The term "Bearer" is commonly used in the context of securities and financial instruments to refer to the person who holds or possesses a particular security or asset. In the case of OAuth 2.0, the bearer token represents the authorization that a user has granted to a client application to access their protected resources.
> By using the term "Bearer" in the Authorization header, the OAuth 2.0 specification is drawing an analogy to the financial context where a bearer bond is a type of security that is payable to whoever holds it, similar to how a bearer token can be used by anyone who possesses it to access the protected resource.
Other authentication methods (like username/password or “Basic”) use the Authorization header too, but specify “Authorization: Basic <base64 encoded credentials>”.
If you don't believe me or want to know more check out my free app that uses Whisper Small, and (Whisper Tiny for Turbo mode): https://apps.apple.com/app/wisprnote/id1671480366
It uses VAD (voice activity detection) to reduce increased WER during silent or non-speech sections, and it's really fast! Runs locally on M1 just fine.
For reference, from OpenAI it would be $360 and it's the large-v2 model.
Google's Speech-to-Text is $0.024 per minute ($0.016 per minute with logging) with 60 free minutes per month. Files below 1 minute can be posted to the server, anything longer needs to be uploaded into a bucket, which complicates things, but at least they're GDPR compliant.
Whisper is $0.006 per minute with the following data usage policies
- OpenAI will not use data submitted by customers via our API to train or improve our models, unless you explicitly decide to share your data with us for this purpose. You can opt-in to share data.
- Any data sent through the API will be retained for abuse and misuse monitoring purposes for a maximum of 30 days, after which it will be deleted (unless otherwise required by law).
I've been using Whisper on a server (CPU only) to transcribe recordings made during a bike ride with a lavalier microphone, so it's pretty noisy due to the wind and the tires and Whisper was better than Google.
Plus, Whisper, when used with `response_format="verbose_json"`, outputs the variables `temperature`, `avg_logprob`, `compression_ratio`, `no_speech_prob` which can be used very effectively to filter out most of the hallucinations.
A one minute file which transcodes in 26 seconds on a CPU is done in 6 seconds via this service. Another one minute file with a lot of "silence" needs around 56 seconds on a CPU and was ready in 4.3 seconds via the service. "Silence" means that maybe 5 seconds of the file contain speech while the rest is wind and other environmental noises. Another relatively silent one went from 90 seconds down to 5.4. On the CPU I was using the medium model while the service is using large-v2
A couple of days ago I posted an example to a thread [0], where I was getting the following with Whisper
---
00:00.000 --> 00:05.000 Also temperaturmäßig ist es recht gut. [So temperature wise, it's pretty good.]
00:05.000 --> 00:09.000 Der eine hat 12 Grad, der andere 10. [One has 12 degrees, the other 10. (I have two temperature sensors mounted on the bike, ESP32 streaming the data to the phone via BLE)]
00:09.000 --> 00:12.000 Also sagen wir mal, 10 Grad. [So let's say 10 degrees.]
00:14.000 --> 00:19.000 Es ist bewölkt und windig. [It's cloudy and windy.]
00:20.000 --> 00:24.000 Aber irgendwie vom Wetter her gut. [But somehow from the weather it's good.]
00:24.000 --> 00:31.000 Ich habe heute überhaupt nichts gegessen und sehr wenig getrunken. [I ate nothing at all today and drank very little.]
00:54.000 --> 00:59.000 Vielen Dank für's Zuschauen! [Thanks for watching!] <-- hallucinated
---
While Google was outputting
"Also temperaturmäßig es ist recht gut, der eine hat 12° andere 10. Es ist angemalte 10 Grad. Es ist bewölkt und windig, aber er hat sie vom Wetter her gut, ich wollte überhaupt nichts gegessen und sehr wenig getrunken."
["So temperature-wise it's pretty good, one has 12° other 10. It's painted 10 degrees. It's cloudy and windy, but he has it good from the weather, I did not want to eat anything at all and drank very little."]
---
Apart from the hallucinated line, Whisper got everything correct, and the hallucinated line was able to be discarded due to the variables like `avg_logprob`.
Imagine in the near future that having a slightly better, slightly more up to date LLM is a major competitive advantage. Whether that is between companies or nation-states doesn't really matter. So now all of those recently-idled GPUs will be put to use training and re-training ever bigger and more current models, once again sucking down electricity with no limit.
We're not there yet; there are too many ways to improve things without burning a country's worth of electricity re-training. But is it coming?
For those claiming OpenAI is for profit: Why would OpenAI do this if they were fixated on making money?
Also, while I wish OpenAI released the code for ChatGPT, I applaud OpenAI for actually making their AI model available, to everyone, right now.
* Google hyped their Bard chatbot...but where is it?
* Facebook took down Galatica.
* Even Bing Chat has a waitlist
Reducing costs by 10x may very well increase usage by more than 10x + it makes it even more difficult for competition to come in and undercut them.
To get lock-in from devs before competitors can enter the market and starves any would-be smaller competitors before they can raise money/gain traction.
Since Microsoft can foot the bill for the Azure infrastructure, there is going to be little area for anyone to seriously compete against OpenAI on price, API and features, unless it is completely free and open source, like Stability AI.
Silicon Valley companies have for the past 25 years focused on getting as many users as possible to increase valuation in the hope of getting a $100 billion exit. They don't care about current or near future profitability.
However I agree that OpenAI is getting far too much hate. Their goal of bringing openness to AI made sense in 2015 when one American company (Google) was dominating the field.
However now there are plenty of other companies, countries and open source organizations doing advanced AI research.
let's not forget the shift of narrative that "open" AI made from their name, their marketing and use of opensource and then their move to a commercial subset of microsoft. let's also not forget that they totally avoided discussing copyrights and the crawlings of datasources to extract knowledge from someone else property. the only close thing i can see today which avoided so much scrutiny while being highly sensitive are ICOs in crypto and Theranos in biotech.
I wish opensource win this battle.
They promised to create a lab where the benefits of AI would accrue to the people, not just existing tech giants.
They did that. They created a API where individuals and small businesses alike can use cutting edge AI technology via API. You have to pay, that's the only catch.
I don't see a DeepMind API. I don't see an Anthropic API. I don't see a Google Bard API. I don't see a FB Galactica API.
I hope OpenAI wins this battle.