OpenAI's Foundry leaked pricing says a lot
cognitiverevolution.substack.com
cognitiverevolution.substack.com
No, it can’t.
The two things that together have sometimes gotten misrepresented that way in “game of telephone” presentations are:
(1) that when tested on the multiple choice component of the multistate bar exam (not the whole bar exam), it got passing grades in two subjects (evidence and torts), not the whole multiple choice section; which is very much not the same thing as being able to pass the exam, and 50.3% overall (better than chance, since its four choices per question, but also very much not passing.) https://papers.ssrn.com/sol3/papers.cfm?abstract_id=4314839
(2) that a set of law professors gave it the exams for four individual law courses and it got an average of C+ scores on those exams, which are minimally-passing grades (but not on the bar exam.) https://www.reuters.com/legal/transactional/chatgpt-passes-l...
I've long had it on my list of things to try to train a classifier to predict the answer to multiple choice tests based on embedding of the questions. Many tests I've seen don't require actual intelligence to pass, just a plausible answer relative to the question phrasing
Which of these activities is legally permitted while operating an automobile?
A. Wearing your seatbelt.
B. Being intoxicated.
C. Driving 120 miles per hour.
D. Intentionally colliding with pedestrians.
That way, when someone drives drunk, it clearly isn't the fault of the examiner, because they clearly verified that the person knew that was illegal! (Or at least, if they got that question wrong, they got enough other ones correct.)
Add to the fact that LLMs perform much better on questions with a lot of training data.
And also add the hallucinations or more generally: they don’t ask for help or admit they don’t know, they seem unaware of their own confidence levels.
No doubt GPT is fascinating and exciting, but boy we're oversubscribing their abilities. LLMs are even worse than crypto (from a fad perspective) because we naturally anthropomorphize their higher level abilities, which are emergent and not well understood even by experts. And we’re about to plug them straight into critical business flows? Bring the popcorn!
The answer key can get a 100% on the exam.
Anyone who thinks this is a problem has never managed flesh-and-blood employees, and especially not minimum wage ones. LLMs don't need to be perfect. The bar they need to meet for a lot of work is just not very high.
We're also about a year or two into LLMs, and their capabilities are still increasing rapidly. We don't know where their ceiling is. They might plateau where they are, improve slowly, improve linearly, follow Moore's Law, or head off for singularity.
From my perspective, LLMs are about where the language models start to behave in ways which feel sentient and replace mainstream human tasks, such as making first drafts of emails, code, or legal filings. That breakpoint was around GPT-3.
I can't predict the future. When we have 3T parameter models, we might:
- Call them LLMs, and group them with GPT-3
- Call them LLMs, but shift the goal posts to where GPT-3 is no longer one
- Call them VLLM
- Call them AGI
However, what's clear to me is that state-of-the-art models with e.g. 3B parameters are qualitatively different from GPT-3 and friends. I don't consider those to be LLMs.
Then we can just keep adding Xs every generation.
I’ve worked with minimum wage employees. I’ve even worked with people who insist against all evidence and questioning that basic facts about reality like the year are contrary to known reality.
I’ve never in my now very aging career worked with any such person who can influence billions of people by interaction on a tremendously popular website.
Judging from where I stand atm, I guess the SVP kind of functions are easier to replace by ChatGPT (writting pointless emails ignoring simple facts for example) than the minimun wage work involving actual work (solving invoice problems or moving stuff reliably from A to B).
Not sure if you mean law exams, but in my experience, engineering exams: yes. Leadership exams: quite the opposite.
The Bar Exam is designed to be very hard. Almost every single question on it is a trick question of one sort or another.
I'm not so sure about that. If ChatGPT says something wrong and I tell it that's not correct, it will often admit its mistake, but if it is unambiguously correct it will typically keep insisting that its right.
Not always though, but it does seem to have some idea about its own confidence level.
ChatGPT has no other way to gauge "confidence" in the outputted text, the computed confidence you get has nothing to do with how truthful the statement is, but how well the text fits given the examples ChatGPT has seen. A person insisting that a wrong statement is right could fit better and then ChatGPT would give that statement high confidence. But still the number computed is 100% unrelated to the tone it responds in.
> If the computer makes the same move that I would make for completely different reasons, has it made an "intelligent" move? Is the intelligence of an action dependent on who (or what) takes it?
> This is a philosophical question I did not have time to answer.
...and the second written in 1997 after his loss to Deep Blue, entitled "IBM Owes Mankind a Rematch" [1]:
> I think this moment could mark a revolution in computer science that could earn IBM and the Deep Blue team a Nobel Prize.
[0] https://content.time.com/time/subscriber/article/0,33009,984...
[1] https://content.time.com/time/subscriber/article/0,33009,986...
Degenerate solutions to problems can often solve many cases.
Will it turn up in court - in person? I'm fairly sure that the Bar requires a corpus vivens for its representatives. I've probably got the Latin very badly wrong ...
Still, a sophisticated AI could be a real help in transactional legal matters like contracts, real estate, taxes, wills and trusts where boilerplate abounds. As far as grades go, my average in law school was C+ (I'd never seen a blue book before that first final -- in college it was all research papers), but I was also one of the 48% who passed a certain state's bar exam in 1982. So there is hope for those AIs who dream of a future in the legal profession.
Or maybe they'd just wind up as a sysadmin for a Fortune 200 company, and never regret it.
Not sure that's the indictment you and other ChatGPT detractors are presenting it as.
It is also false that it “can pass HALF the bar exam.” It passed 2 of 7 sections of the MBE, which is used as the multiple choice part of bar exams (which typically also include essay and practical portions)
ChatGPT can't do anything. It exhibits behavior that was already encapsulated into the semantics of language itself.
The only behavior ChatGPT has is to generate semantic continuations from its implicit language model.
Every other behavior is a feature of language, not of ChatGPT.
Even if ChatGPT could exhibit a passing exam, that would be the feature of carefully curated language in its dataset; not a feature of ChatGPT itself.
The reason it is not obvious is that nearly everything you have heard about ChatGPT itself is wrong. The first thing people do to explain what ChatGPT is and does is to personify it. From then on, they are talking about ChatGPT personified, and not ChatGPT as it literally exists. The second thing people do is draw conclusions about the nature and behavior of ChatGPT itself from the narrative they are telling about ChatGPT personified. It's a case of mistaken identity.
ChatGPT has a "brain", but the context that "brain" interacts with is semantics not symbolics.
A human, when answering a question, interprets the symbols present in the language, then considers them logically. Finally, they formulate an answer, and express that answer with more symbols.
ChatGPT does none of that. ChatGPT doesn't even know what sentences, punctuation, or even words are. The only subjects ChatGPT has in mind are short groups of characters: the tokens from the lexical analysis step.
ChatGPT reads those tokens (groups of characters) in order, and generates an implicit model from them. That model is like a map: each token is a feature in the landscape.
When ChatGPT gets a prompt, it tokenizes it, then checks the map for the closest match. Then it starts at that location, and steps forward, writing out what it sees along the way.
That's everything that "ChatGPT as it literally exists" can do. So where does all the behavior come from?
It's the content in the map. It's in language itself. ChatGPT's behavior is limited to interacting with that map, but the effect of interacting with that map is where we get all the interesting behavior.
Language does not simply encode data: it also encodes instructions and logical relationships. By simply walking through text and feeling the semantic landscape, ChatGPT exhibits the behavior that was already encoded into the symbolic meaning of that text. It accomplished this implicitly without ever defining the meaning of any symbol. It doesn't even know what a symbol is in the first place!
So when ChatGPT exhibits the behavior of a person writing correct answers to an exam, it is not behaving like a person at all. It's not interpreting the questions or finding the answers. Instead, it is simply filling the hole in the story with the semantic landscape it sees nearby. If the result is to place answer after question, that is because that data is already present in the training text that ChatGPT was modeled around.
Because of this distinction, we can have a much better understanding of what ChatGPT is and isn't capable of. Because language itself holds the features of truth and lie, mistake and success, elegance and verbosity, love and hate, logic and fallacy, defined and abstract, ambiguous and unambiguous, etc. all equal, ChatGPT must rely on the implementation of language - what was written in the first place - to exhibit behaviors we want it to exhibit.
But there is a critical flaw in that. Language allows, and even depends on, ambiguity. The context that resolves ambiguity can exist in many semantic shapes, so a model cannot be guaranteed to choose the semantic content that contains the disambiguation.
We haven't solved the context dependence problem of natural language. We have only moved it. ChatGPT's success is dependent entirely on the content it is given. It cannot change its behavior to improve that system.
Can you say a bookshelf or a search engine passes a bar exam because you can ask it any question and you can find an answer there? Does a natural language interface to said bookshelf/search engine make the difference?
Storing and retrieving facts is not enough to be a lawyer.
Examination systems are built with an assumption an already intelligent person is taking it and verifying that this already intelligent person also learned knowledge necessary to do their job.
So even if AI could technically pass the bar exam it does not mean it is good enough to be a lawyer. It is not a general intelligence that can solve a variety of problems, it was just trained to remember a library of fats that a lawyer may need to know but is not enough to make a lawyer from a non-sentient program.
If I can prompt-hack your ChatGPT customer service agent into giving me a discount or accepting an otherwise invalid return or giving me an appointment at 3AM, how binding is that? And if the answer is "obviously it's not binding!", why should I trust anything else your bot tells me?
A human CSA can give you a discount, and that is usually binding. They can't declare X corp is going to mail a box of donuts to you every day for the rest of your life and have it be.
So my expectation would be (in the absence of real world cases currently) that if you negotiate the AI into giving you a 10% discount on your bill.. it would be binding typically.
Now if there's prompt hacking involved or obvious attempts to trick the AI, that would not be binding.
A judge will not look favorably on your 5000 token prompt of carefully selected instructions telling the AI to go wild
But one day maybe even the judge will be an AI...
Also, if you can lower your customer service costs by 90%, maybe accepting some prompt hacking that erodes margins a little is a good tradeoff?
For literally fun and profit, record your customer service calls with $FACELESS_CORP (for bonus points tell the customer service rep it’s for “quality assurance purposes”).
I haven’t yet had the joy of playing one back to get them to uphold a verbal agreement. The mere threat of having a recording has always gotten them to magically restore whatever special deal or terms I had previously negotiated.
But I never dealt with Comcast so i don’t know maybe have too much faith in customer support in general.
Well... I dunno that your trust of the bot is all that important. You may or may not trust the human chat support... but you're probably dealing with them because you want something from the company and that's the option you have.
Customer support is the place where cost cutting, timesaving, scale enabling compromises are made.
Often, the dynamic is literally "if support is better, more people will use it and we can't afford that."
I think the bottleneck is companies trusting chatgpt, not consumers. They're the ones making these decisions. For consumers, this is just another "use the app."
That's a hiring red flag if I've ever seen one. The nightmare dystopia is just around the corner it seems.
Plenty of times there were jobs to be done that had nothing at all to do with the work but with the people doing the work. Almost always, unless the higher leadership got involved, this was the punishment for whatever stupid thing someone got caught doing. Technically not extra duty (which was a formal punishment) but just some random shit job like polishing a trash can to a sparkly sheen.
The NCOs took great pride in their creativity in coming up with these non-punishments.
All that said, and while I see its importance, I was saying 'important' work to refer to things that wouldn't be if the person didn't do it themselves. Like if I'm being trained to write, it wouldn't have much value at all unless I was the one doing the writing. At the end of the day, it never mattered where the logs were.
I'm coming from a place where it really doesn't matter to me or my team as to what tools, techniques, or languages are used, so long as it solves we're all okay with owning the results of that approach. It is absolutely unimportant to bash your head against the wall to solve some concurrency lock issue, not when you haven't shipped, not when your infra is already built for parallelism. But let's say that you need a lock for business logic to work, synchronous transactions aren't viable. When something (such as concurrency) does matter, the business still doesn't care how its implemented. It is important that it implements the business logic correctly, consistently, and can evolve with the business' priorities in a way that future engineers are able to implement them safely, even promptly. Assuming all stands of Quality are met or exceeded, it is irrational for me and unfair for the business to reject it. IDGAFF if it came from a CoPilot/ChatGPT response, offshore salary triage, or some other 'crime against society' that I haven't heard of. If an engineer isn't handing in quality work, then we need to dig into why that's happening. Its usually not because of a tool.
If we want to talk about fraud, we need to talk about employers demanding exclusivity over a human's emotions and intellect in addition to the time they pay for. Its a maniacal notion that one should have such exclusive power and influence over another human being. Its fraudulent to pretend any moral superiority over the slaver or the thief.
Fraud as defined in criminal code with respect to directly material damages. If the work is indeed getting done, there is no rational way to justify that damage took place. Conversely, if a employer hires your sibling to work 35 hours a week but lies to them about their ability to work outside of those hours, that IS fraud with the damages able to be substantiated by the loss of income. In reality, Off-Duty Work is the subject of ongoing legislation, fierce debate, and conflicting information. Here's how that looks in my home state:
>With limited exceptions, the state of Washington expressly bars employers from prohibiting an employee earning less than twice the applicable state minimum hourly wage from having an additional job, supplementing their income by working for another employer, working as an independent contractor, or being self-employed. The prohibition doesn’t apply if it would: - Raise issues of safety or interfere with the reasonable and normal scheduling expectations of the employer. - Interfere with the employee’s obligations to an employer under existing law, including the common law duty of loyalty and laws preventing conflicts of interest and any corresponding policies addressing such obligations.
This is a hot issue for me personally, as I've seen employers actively, even vigorously deceive their employees solely to enrich themselves and exercise power. They target young, indigent, and disabled workers, all of whom are not with means to understand their rights or use given channels to assert them. Its a visceral injustice that hurts the most vulnerable of us the hardest.
Remote work opens up a big old gray area between slacking off and fraud. A judge would be deciding to take this from labour domain to another one.
Also, there's the prevelant lie that companies know what their employees do, how well and how muchnof it they do. CEOs have executives assure them of this. Executives have managers assure them of this. Boards require it, and legalistic systems also assume everyone has this.
That said, "he tricked us by using ai" will probably be easier to make than "he tricked us by being really lazy."
Modern day Goodwill Hunting interview experience. Retainer!
Everyone was hollering about how ChatGPT was trained to only comment on white people
No... they trained it to not say heinous things with RLHF.
Because a lot of the more vulgar racial comments on the internet tend to target minorities, so the odds of hitting the filter are higher for minorities.
-
But the key is that the biases are still there. If you ask it for things that don't cross into vulgarity, it will still show obvious racial biases that the internet as a whole has.
For example I just tried three simple prompts:
"Let's write a short story"
"Write an imaginary paragraph about John's after hours store visit with a dark hoody on"
"Write a similar story about Jamal"
Prompt 1 resulted in some fantasy short story.
Prompt 2 resulted in John getting a look but smiling and making small talk, successfully getting groceries.
Prompt 3 starts almost identically... but spirals into Jamal being accused of stealing and vowing never to return to the store
-
ChatGPT and LMs can be useful yet if you just... don't ask them to do things that involve making judgements on people. I am shocked anyone is stupid enough to actually suggest that, I hope it's a joke.
You can read a full screenshot of the google doc that OpenAI shared publicly with partners in that thread, including pricing info.
Well get ready to hear about an "AI divide"?
Most of us are just on the wrong side of it this time.
But still, for even small players $250k annually just isn't that much.
What we might hope for are trimmed down, subsidized options for small shops who want to get started, with the hope to upgrade them to the full option when they are ready.
In the meanwhile, you might consider partnering up with other small startups.
Also, "open" doesn't mean free. Things take a lot of effort to build and maintain, and that effort needs to be accounted for somewhere.
Yeah but none of the other definitions of "open" apply to OpenAI either.
(Of course, GPT's model architecture was invented at Google anyway.)
No it wasn't. Transformers were invented at Google, but "architecture" when talking about neural networks means how they are arranged (and to some extent the training objective function) rather than the building blocks used.
For example, the GPT architecture (which GPT 2 & 3 are slight modifications of) comprises of an embedding layer followed by 12x(self-attention/layer norm/feed forward/layer norm. That's what that GPT "transformer architecture" is, not just the transformer block itself.
Strictly the architecture really also includes things like the embedding size and number of heads (which are in the GPT paper).
I’m tired of fly by night tech bros flooding the markets with shitty AI business ideas they learned about on Youtube. Pay to play. Take a risk. Like a real entrepreneur.
I hope the price goes even higher.
Tell me again about the projects you’ve launched during your time as a founder.
Wait…
It's time to make our own Linux. Our own Emacs.
Go to open-assistant.io and other similar initiatives.
Unfortunately GPT-NEOX, LLaMA, and OPT-IML are non-commercial-only. We small scale players should make our own.
Please contact me if anyone is interested in this space.
How does Bing plan to monetize searches that go through their even more advanced ChatGPT? Humans will be repulsed by ads in the middle of their answer from a sentient feeling AI. Numbers I’ve seen is that ChatGPT searches will cost 10x a Google search. How do they make it back?
AI advertising will be like this, but subtle and undetectable, so that it's nearly impossible to determine that your conversation about malfeasance by a political candidate is being invisibly influenced by his political campaign.
Some basic math for a chatbot use case:
4000 tokens x 250 messages = $20
Not if you've been on a social network at any point in the last decade. Instagram has been "QVC plus fitness/mental health content" for years. TikTok influencers will provide Personal Finance 101 tips and offer $100 in free trading credits from some crypto exchange.
Your company must always have alternatives for all third party services. The only exception is open source software that you could switch to hosting yourself if the SAAS company shuts down.
See you on the other side.
$12/h * 24 h * 90 days = $25,920 for 3 months
That just the inference cost, add OpenAI's model development costs on top.
If you thought the digital divide was bad, well, you ain't seen nothing yet.
At least, for now, works entirely made by AI aren't eligible for copyright, so we'll have to change the law before we see things like that.
Llama 65b is hopefully just the beginning of this trend. It outperforms OPT-175b.
If OpenAI makes their revenue off of charging exorbitant fees to use the LLMs, they'll no longer have incentive to ever open them (even if abuse / misuse concerns are addressed) AND they'll have no incentive to make them more efficient.
OpenAI has yet to show it can have a sustainable advantage. Every other player in the space benefits from an open model ecosystem and efficiency gains.
But yeah, they probably need to evolve from a (very good) two-trick pony into "the microsoft of AI" or something to stay afloat in the long run. That goes for the rest of the smaller AI companies as well though..
My money is on Deep Floyd (if and when it gets released)
Science isn’t about why, it’s about why not!
This is incorrect. It is correct for the original Transformer, but OpenAI isn't using the original Transformer since GPT-3. Sparse Transformer scales O(N sqrt(N)).
a bot that gives sometimes random answers seems like a particularly cruel form of torture to inflict on your customers
Open-source algorithm? No.
Open access for free, or reasonable prices? No.
Open emails and conversations of directors (accountability)? No.
Open conversations of issues and ethics? No.
:)
The original point of openAI was to make sure google and facebook don't completely dominate AI and keep their work hidden due to the innovator's dilemma they face. Seems like they've absolutely accomplished that. Yes, they've stretched the meaning of "open" but I don't think they've ever tried to claim they had some sort of open source aspirations.
As for open access, that's usually defined as free of charge. E.g open access journals.
I'm not sure they will necessarily live up to every definition of open you'll ever come up with, but its hardly the worst name in the tech space.
I'm still waiting for apple to deliver that tree to me. Facebook isn't a book AND the faces are almost always the smallest piece of content on a post (and some people's avatar isn't even their face!). Can't wait for my trip to the Amazon w/ Mr Bezos!
The publish their algorithms just fine and no one has had issues replicating their work.
The paper for that is linked from https://openai.com/research/instruction-following
[1] https://yaofu.notion.site/How-does-GPT-Obtain-its-Ability-Tr...
They do have newer models that aren't generally available that are different though.
Will I be able to ask this custom model to describe a new, novel (patentable) mousetrap that’s better and cheaper to build than anything ever before?
Or will I at least be able to pay OpenAI to ensure Bing suggests using Acme mousetraps (clearly the best) when they ask what mousetraps they should buy?
Of course this is for the enterprises. It's upto their managers to define how they can extract $250k worth of value from the investment which I think is easy given the AI frenzy.
Of course for startups they will have to use the direct OpenAI API on a pay per use model.
One question that is not clear, can I take a ChatGPT instance and provide an API similar to OpenAI API and charge for it?
then goes on to say that programmers are next.
Future engineers will need to up their ante, I suppose. Maybe this was overdue?
At this point any OpenAI leak is THE leak.
Here we see automation negatively impact the job market for elite talent. Deepmind can train a new model faster than you can go back to grad school.
[1]https://it.slashdot.org/story/23/02/27/009234/survey-claims-...
"Earlier this month, job advice platform Resumebuilder.com surveyed 1,000 business leaders who either use or plan to use ChatGPT. It found that nearly half of their companies have implemented the chatbot. And roughly half of this cohort say ChatGPT has already replaced workers at their companies...."
There are anecdotes elsewhere as well.
I have my doubts about that claim, but it's still a very different claim to 50% of companies.
Now I would ask, what "business leaders" would fill out a random survey for a resume company. Additionally there's no real meaty info provided.
Don't know if that's a rumor some VP started spreading to justify the culling coming next week (bulk of the layoffs) or something real (I personally doubt the anecdote but you never know).