Anthropic’s $5B, 4-year plan to take on OpenAI
techcrunch.com
techcrunch.com
Currently a Macbook has a Neural Engine that is sitting idle 99% of the time and only suitable for running limited models (poorly documented, opaque rules about what ops can be accelerated, a black box compiler [1] and an apparent 3GB model size limit [2])
OTOH you can buy a Macbook with 64GB 'unified' memory and a Neural Engine today
If you squint a bit and look into the near future it's not so hard to imagine a future Mx chip with a more capable Neural Engine and yet more RAM, and able to run the largest GPT3 class models locally. (Ideally with better developer tools so other compilers can target the NE)
And then imagine it does that while leaving the CPU+GPU mostly free to run apps/games ... the whole experience of using a computer could change radically in that case.
I find it hard not to think this is coming within 5 years (although equally, I can imagine this is not on Apple's roadmap at all currently)
[1] https://github.com/hollance/neural-engine
[2] https://github.com/smpanaro/more-ane-transformers/blob/main/...
If it could use the full 'unified' memory that would be a big step towards getting these models running on it
I'm unsure how the performance compares to a beefy Intel CPU, but there's some numbers here [1] for running a variant of the small distilbert-base model on the Neural Engine... it's ~10x faster than running on the M1 CPU
[1] https://github.com/anentropic/experiments-coreml-ane-distilb...
1. That RAM isn't empty, it's being used by apps and the OS. Fill up 64GB of RAM with an LLM and there's nothing left for anything else.
2. 64GB probably isn't enough for competitive LLMs anyway.
3. Inferencing is extremely energy intensive, but the MacBook / Apple Silicon brand is partly about long battery life.
4. Weights are expensive to produce and valuable IP, but hard to protect on the client unless you do a lot of work with encrypted memory.
5. Even if a high end MacBook can do local inferencing, the iPhone won't and it's the iPhone that matters.
6. You might want to fine tune models based on your personal data and history, but training is different to inference and best done in the cloud overnight (probably?).
7. Apple already has all that stuff worked out for Siri, which is a cloud service, not a local service, even though it'd be easier to run locally than an LLM.
And lots more issues with doing it all locally, fun though that is to play with for developers.
I hope I'm wrong, it'd be cool to have LLMs be fully local, but it's hard to see situations where the local approach beats out the cloud approach. One possibility is simply cost: if your device does it, you pay for the hardware, if a cloud does it, you have to pay for that hardware again via subscription.
1. the idea would be that now there is a reason to buy loads more RAM, whereas currently the market for 64GB is pretty niche
2. 64GB is a big laptop today, in a few years time that will be small. And LLaMA 65B int4 quantized should fit comfortably
4. LLMs will be a commodity. There will be a free one
6. LLMs seem to avoid the need for finetuning by virtue of their size - what we see now with the largest models is you just do prompt engineering. Making use of personal data is a case of Langchain + vectorstores (or however the future of that approach pans out)
2. Why would anybody be satisfied with a 64GB model when GPT-4 or 5 or 6 might even be using 1TB of RAM?
3. That may not be the case. With every day that passes, it becomes more and more clear that large LLMs are not that easy to build. Even Google has failed to make something competitive with OpenAI. It's possible that OpenAI is in fact the new Google, that they have been able to establish permanent competitive advantage, and there will no more be free commodity LLMs than there are free commodity search engines.
Don't get me wrong, I would love there to be high quality local LLMs. I have at least two use cases where you can't do them or not really well with the OpenAI API and being able to run LLama locally would fix that problem. But I just don't see that being a common case and at any rate I would need server hardware to do it properly, not Mac laptop.
I'm really not
I had no desire at all until a couple of weeks ago. Even now not so much since it wouldn't be very useful to me
But the current LLM business model where there are a small number of API providers, and anything built using this new tech is forced into a subscription model... I don't see it sustainable, and I think the buzz around llama.cpp is a taste of that
I'm saying imagine a future where it is painless to run a ChatGPT-class LLM on your laptop (sounded crazy a year ago, to me now looks inevitable within few years), then have a look at the kind of things that can be done today with Langchain... then extrapolate
And yes I've played with it. It was/is exciting. I can see use cases for it. However none are achievable because the models are (a) not good enough and (b) too legally risky to use.
(B) llama.cpp supports gpt4all, which states that its working on fixing your concern. This is from their README:
Roadmap Short Term
- Train a GPT4All model based on GPTJ to alleviate llama distribution issues.
The point of llama.cpp is most people don't have a GPU with enough RAM, Apple unified memory ought to solve that
Some people have it working apparently:
before I found the repo above I had a naive attempt to get llama running with mps and it didn't "just work" - bunch of ops not supported etc
It's free for the user up to a point, but it costs OpenAI a lot of money.
Apple is a hardware vendor, so commoditization of the software while finding more market segments is definitely something that'd benefit them.
OTOH, if they let OpenAI become the unrivaled leader of AI that end up being the next Google, they end up losing on a topic they wanted to lead for long time (Apple has invested quite a lot in AI, and the existence of a Neural Engine in Apple CPUs isn't an accident)
if OpenAI isn't able to get couple hundred bucks over the typical lifetime of a computer it means the added value they provide is very low (like several times less than Spotify or Netflix for instance), meaning they'll never be “the next Google”.
And if they are it means it make sense to buy it once instead of paying several times the price through subscription.
> The OpenAI APIs are super cheap for a single user needs. I expect them to be at least close to breaking even with their APIs pricing.
“Close to breaking even” means the price you pay is VC-subsidized, the expected gross margin for such kind of tech company is more than 50%. Expect to pay a lot more if/when the market is captive. And this will scale linearly with your use of the technology.
> energy and opportunity costs
What opportunity cost?
Yes, this is a possibility but cloud computing became a commodity.
But I see why people would pay to have their own private and unfiltered models/embeddings.
> if OpenAI isn't able to get couple hundred bucks over the typical lifetime of a computer it means the added value they provide is very low (like several times less than Spotify or Netflix for instance), meaning they'll never be “the next Google”.
They don't have to worry about this today.
> What opportunity cost?
You could utilize the money and the time spent to do other things.
Doesn't the iPhone use the local processor for stuff like the automatic image segmentation they currently do? (Hold on any person in a recent photo you have take and iOS will segment it)
I think the most glaring situation where this is true is simply one of trust and privacy.
Cloud solutions involve trusting 3rd parties with data. Sometimes that fine, sometimes it's really not.
Personally - LLMs start to feel more like they're sitting in the confidant/peer space in many ways. I behave differently when I know I'm hitting a remote resource for LLMs in the same way that I behave differently when I know I'm on camera in person: Less genuinely.
And beyond merely trusting that a company won't abuse or leak my data, there are other trust issues as well. If I use an LLM as a digital assistant - I need to know that it's looking out for me (or at least acting neutrally) and not being influenced by a 3rd party to give me responses that are weighted to benefit that 3rd party.
I don't think it'll be too long before we see someone try to create an LLM that has advertising baked into it, and we have very little insight into how weights are generated and used. If I'm hitting a remote resource - the model I'm actually running can change out from underneath me at any time, jarring at best and utterly unacceptable at worst.
From my end - I'd rather pay and run it locally, even if it's slower or more expensive.
> Almost everyone is willing to trust 3rd parties with data, including enterprise and government customers.
Is absolutely not true. In it's most basic sense - sure... some data is trusted to some 3rd parties. Usually it's not the data that would be most useful for these models to work with.
We're already getting tons of "don't put our code into chatGPT/Copilot" warnings across tech companies - I can't imagine not getting fired if I throw private financial docs for my company in there, or ask it for summaries of our high level product strategy documents.
Saying that cloud models will win over local models is not the same as saying it will be a free-for-all where workers can just use whatever cloud offering they want. It will take time to enterprisify cloud LLM offerings to satisfy business/government data security needs, but I'm sure it will happen.
LLMs will require more than privacy to move locally. Latency, flexibility and cost seem more likely drivers.
I care more about the trust I have to place in the response from the model.
Hell - since you mentioned search... Just look at the backlash right now happening to google. They've sold out search (a while back, really) and people hate it. Ads used to be clearly delimited from search results, and the top results used to be organic instead of paid promos. At some point, that stopped being true.
At least with google search I could still tell that it was showing me ads. You won't have any fucking clue that OpenAI has entered into a partnering agreement with "company [whatever]" and has retrained the model that users on plans x/y/z interact with to make it more likely to push them towards their new partner [whatever]'s products when prompted with certain relevant contexts.
Only people in HN-like communities care about this stuff. Most people find the SEO spam in their results more annoying.
> At least with google search I could still tell that it was showing me ads. You won't have any fucking clue that OpenAI has entered into a partnering agreement with "company [whatever]" and has retrained the model that users on plans x/y/z interact with to make it more likely to push them towards their new partner [whatever]'s products when prompted with certain relevant contexts.
You won't know this for any local models either.
But you will know the model hasn't changed, and you can always continue using the version you currently have.
> Most people find the SEO spam in their results more annoying.
This is the same problem. These models will degrade from research quality to mass market quality as there's incentive to change what results they surface. Whether that's intentional (paid ads) versus adversarial (SEO) doesn't matter all that much - In either case the goals will become commercial and profit motivated.
People really don't like "commercial and profit motivated" in the spaces that some of these LLMs stepping into. Just like you don't like SEO in your recipe results.
Will you? What happens when an OS update silently changes the model? Again this is one of those things only HN-types really care/rant about. I've never met a non-technical person care about regular updates beyond being slow or breaking an existing workflow. Most technical folks I know don't care either.
> This is the same problem. These models will degrade from research quality to mass market quality as there's incentive to change what results they surface. Whether that's intentional (paid ads) versus adversarial (SEO) doesn't matter all that much - In either case the goals will become commercial and profit motivated.
Not at all. Search providers have an incentive to fight adversarial actors. They don't have any incentive to fight intentional collaboration.
> People really don't like "commercial and profit motivated" in the spaces that some of these LLMs stepping into. Just like you don't like SEO in your recipe results.
I disagree. When a new, local business pops up and pays for search ads, is this "commercial and profit motivated?" How about advertising a new community space opening? I work with a couple businesses like this (not for SEO, just because I like the space they're in and know the staff) and using ads for outreach is a pretty core part of their strategy. There's no neat and clean definition of "commercial and profit motivated" out there.
Yeah but in the cloud that cost is ammortized among everyone else using the service. If you as a consumer buy a gpu in order to run LLMs for personal use, then the vast majority of the time it will just be sitting there depreciating.
iOS actually does already have an offline speech-to-text api. Some part of Siri that translates the text into intents/actions is remote. Since iOS 15, Siri will also process a limited subset of commands while offline.
No, it'll be a commodity
Apple wouldn't care if the weights can be extracted if you have to have a Macbook to get the sweet, futuristic, LLM-enhanced OS experience
Edit - I found an example from h.n. user anentropic, pointing at https://github.com/remixer-dec/llama-mps . "The goal of this fork is to use GPU acceleration on Apple M1/M2 devices.... After the model is loaded, inference for max_gen_len=20 takes about 3 seconds on a 24-core M1 Max vs 12+ minutes on a CPU (running on a single core). "
Google had already done some very convincing demos in the last few years well before ChatGPT and GPT-4 captured the popular imagination. Microsoft’s OpenAI deal I would assume will lead to a “Cortana 2.0” (obviously rebranded, probably “Bing for Windows”, “Windows Copilot” or something similar). Google Assistant has been far ahead of Siri for many years longer than that, and they have extensive experience with LLMs. Apple surely realises the position their platforms are in and the risk of being left behind.
I’m also not sure the barrier on iPhone is as great as you suggest - it’s obviously constrained in terms of what it can support now but if the RAM on the device doubles a few times over the next few years I can see this being less of an issue. Multiple models (like the Alpaca sets) could be used for devices with different RAM/performance profiles and this could be sold as another metric to upgrade (i.e. iPhone 16 runs Siri-2.0-7b while iPhone 17 runs Siri-2.0-30b - “More than 3x smarter than iPhone 16. The smartest iPhone we’ve ever made.” etc).
Apple is ahead of the game for a change getting their chips in line as the software exits alpha and goes mainstream.
This is just allowing PyTorch to make use of the Apple GPU, assuming the models you want to train aren't written with hard-coded CUDA calls (I've seen many that are like that, since for a long time that was the only game in town)
PyTorch can't use the Neural Engine at all currently
AFAIK Neural Engine is only usable for inference, and only via CoreML (coremltools in Python)
For 64gb of ram, you can get an m2 pro, or get 96gb which requires the upgraded cpu on the pro. The studio does 64gb or 128gb. But the 128 requires you to spend 5k.
I can't decide between 64 or 96 on m2 pro, and 128 on the studio. Probably go for 96gb. Also what's the impact of the extra gpu cores on the various options? And there are still some "m1" 64gb pros & studios out there. What's the perf difference for m1 vs m2? This area needs serious perf benchmarking. If anyone wants to work with me, maybe I would try my hand. But I'm not spending 15k just to get 3 pieces of hardware.
List prices:
64gb/2tb m2 12cpu/30gpu 14" pro $3900
96gb/2tb m2 max 12/38 14" pro $4500
128gb/2tb m2 max 28/48 studio $5200
…or is it just that the latter had a way too weak iGPU and not enough RAM for AI purposes, whereas the bigger ARM MACs have more GPU power and enough RAM (more than most affordable discrete graphic cards) so that they are usable for some AI models?
> 64GB probably isn't enough for competitive LLMs anyway
I am trying to charitable, but this is pretty not true. And the hedging in your statement only telegraphs your experience.
2. see above
Should be cheap, or why else are Samsung, Micron and Kioxia whining about losses?
Maybe go for something like Optane memory while doing so.
What about other older 'questions' we can point an AI lens at?
Very doubtful unless the user wants to carry around another kilogram worth of batteries to power it. The hefty processing required by these models doesn't come for free (energy wise) and Moore's Law is dead as a nail.
But anyway, there are two trends:
- processors do more with less power
- LLMs get larger, but also smaller and more efficient (via quantizing, pruning)
Once upon a time it was prohibitively expensive to decode compressed video on the fly, later CPUs (both Intel [1] and Apple [2]) added dedicated decoding hardware. Now watching hours of YouTube or Netflix are part of standard battery life benchmarks
[1] https://www.intel.com/content/www/us/en/developer/articles/t...
[2] https://www.servethehome.com/apple-ignites-the-industry-with...
Or, you can use a superior GPT 3.5 for free.
- such amounts of memory is locked behind very expensive sku which even most of mac userbase will not use ( <5% in the new purchases to be very conservative ). - not too long ago apple would restrict the amount of ram in their system for their own reasoning (source: https://9to5mac.com/2016/10/28/apple-macbook-pro-16gb-ram-li...) - just like mid 2010s gpus with 6-8 gb vram but with little to benefit from it, i don't see the ml accelerators/gpu in current models being capable enough to make the most of the memory available to it.
My first computer had 512KB RAM and 20MB was an expensive hard drive.
64GB Macbooks are currently an expensive 'Pro' novelty, they will be the vanilla of tomorrow
> i don't see the ml accelerators/gpu in current models being capable enough to make the most of the memory available to it
that's exactly my point (and apparently today's Neural Engine can't even take advantage of all the unified memory available)
until LLaMA there was no reason to have more than this, they probably imagined it would just run a bit of face-detection and speech-to-text on the side
but if they got serious and beefed it up it could be the next wave of computing IMHO
The iPhone 14 runs Whisper model faster than an M1 Max, because it has a newer Neural Engine
I look forward to the M3 Macbook launch eagerly, while expecting mild disappointment
So Anthropic is the Google-supported equivalent of OpenAI? Isn't the founder going to run into the same issues as before (commercialization at OpenAI)? How does Google not use Anthropic as either something commercial or nice marketing material for its AI offerings?
> As the Product Policy Lead, you will set the foundation for Anthropic’s approach to safe deployments. You will develop the policies that govern the use of our systems, oversee the technical approaches to identifying current and future risks, and build the organizational capacity to mitigate product safety risks at-scale. You will work collaboratively with our Product, Societal Impacts, Policy, Legal, and leadership teams to develop policies and processes that protect Anthropic and our partners.
> You’re a great fit for the role if you’ve served in leadership positions in the fields of Trust & Safety, product policy, or risk management at fast-growing technology companies, and you recognize that emerging technology such as generative AI systems will require creative approaches to mitigating complex threats.
> Please note that in this role you may encounter sensitive material and subject matter, including policy issues that may be offensive or upsetting.
“Anthropic has been heavily focused on research for the first year and a half of its existence, but we have been convinced of the necessity of commercialization, which we fully committed to in September [2022],” the pitch deck reads. “We’ve developed a strategy for go-to-market and initial product specialization that fits with our core expertise, brand and where we see adoption occurring over the next 12 months.”
The cynic in me wants to ask "What makes you think his departure was because of an anti-commercialisation position?"
My take (probably just as wrong as everybody's else take) is that he saw the huge commercialisation potential and realised that he could make even more money by having a larger stake, which he got when he started his own venture.
I looked and I didn't get an answer. hence my comment.
To clarify, we know what he said his reason was, we don't know if that really was his reason.
When people leave they very rarely voice the actual reason for leaving; the reason they give is designed to make them look as good as possible for any future employer or venture.
So they lost the plot on the altruistic mission within months of setting up shop, and now are just a pawn in a bigger game between other companies.
How does that name collision work?
Kind of unsure how it all works.
I think the unstated shift that has happened in the past few years is that we've gone from researchers thinking about Fourier transforms to efficiently encode positional data into vectors to researchers thinking about how to train a model with a 100k+ token batch size on a super-computer-like cluster of GPUs.
I can totally see why people believed the math could be done in a non-profit way, I do not see how the systems engineering could be.
Reminds me of the (possibly LLM-generated) marketing tirade of a voice faking text-to-speech service recently here on hn, which ended with: "We are thrilled to be sharing our new model, and look forward to feedback!":
https://news.ycombinator.com/item?id=35328698
… "share" yeah right… like: where can I download the model then? Of course they didn't mean to actually share their model but only to rent out remote access to it, but that doesn't sound as nice as "share".
Unsure why nobody is taking this very very obvious hole in AI tech.
Better (which I assume is your euphemism for "more") regulation isn't neceesarily the answer, or even particularly the answer. Do you want to force payment processors to do work they don't want to do? Isn't there a word for that?
PayPal is the prime example where it's operating very similar to a bank. You have an account with a balance and can send and receive money, but it doesn't see itself as a bank and in many countries doesn't have a bank license. At least in part this is done to avoid the regulatory work that comes with it.
I absolutely want to force payment processors to do work they don't want to do. For example, banks in Germany are forced to provide you with a basic bank account regardless whether they want to or not. That's because a bank account is simply a must have to take part in modern life. If PayPal decides it doesn't want to do business with you, for whatever arbitrary reason, you are effectively locked out of a lot of online stores that only accept PayPal as a payment method. There is plenty of examples of PayPals really sketchy behaviour online. Every few months you can even see complaints on HN about it.
We might be talking at cross purposes; I'm not sure! How is it like a bank?
If I pay with a credit card, there are processes in place to deal with fraud and charge backs. PayPal is well known to automatically close accounts with little recourse to access the money on those accounts.
They should absolutely be regulated.
But they are nothing like a bank
The feature of a bank is credit creation. Lending more money than they hold.
Unless I missed some news PayPal does not do that
If you are on a scale like Visa and MasterCard you're not just any private company anymore. Just those 2 companies control well over 75% of the US market alone. Not having access to a debit/credit card today will effectively block you from taking part in many aspects of modern life. It's absolutely reasonable to place stipulations on what they can and cannot do.
Regulators love working with large businesses like your card duopoly, I don't think you will see much improvement.
Public utility. That’s what payment processors are at this point, and they should be regulated as such.
that's not going to lead to a whole lot of black market images.
But this seems like the world the “AI-regulators” seem to want.
I certainly think if the parents found out about it and the law wouldn't do anything about it the parents would take the law into their own hands.
I'm sorry if this wasn't phrased very well. I just didn't know how else to make my point with out be very specific.
Now, if they were drawn to resemble specific people and the producer of the "artwork" used them to harass those people, that's harassment. If they used them to groom other kids, that's an existing crime too. But my point was that the production of gross art in isolation, or the possession of it, didn't need to be criminalized. (Actual photographs of the same were criminalized because of the pretty decent assumption that minors were coerced, harmed, exploited. Probably all of the above.)
I actually just migrated away from Hetzner last week (for unrelated reasons) to two new providers to whom I'm paying crypto (no KYC required) based on this list: https://bitcoin-vps.com/
I had paid for my servers with some Litecoin I have that I usually use for small purchases because of the low fees.
Do you really think the execs at Visa and Mastercard are puritans and not profiteering capitalists that will process payments for NSFW content if they were able to?
Totally agreed. But I am not placing any moral value on either greed or capitalism. I would think, however, that capitalists would not ignore such an obvious profit center as the sex industry. Thus my bafflement.
Because you're conflating capitalism and greed. Capitalism doesn't mean "do anything for money". It means "as much as possible, people get to decide among themselves how to allocate their money and time". Some of them will invest in anything, just as people in non-capitalist countries. Most will only invest in certain things.
In the abstract, perhaps not. The way it exists in the US, though, it means exactly that.
That is a missed opportunity
* Capitalism: A system where who owns resources matters more tan who needs them is a morally bankrupt system. A system where starvation and homelessness is an acceptable outcome
* Greed. Greed is bad for everybody. Concentrates scarce resources where they are not needed, that too is moral bankruptcy
As for greed, I have yet to meet a person more greedy than the ones claiming to know where to direct those scarce resources they did not create, if only we’d give them the power to do so. Such high morals too, unlike those "morally bankrupt" capitalists who greedily built businesses, jobs, countless goods and services to only enslave us and enrich themselves, obviously.
Edit: Quoted from memory of https://en.wikipedia.org/wiki/Wall_Street_(1987_film)
Also getting all bent out of shape at a the image of a nipple, breast or pubic hair while not batting an eye at a person dying in evening TV movies seem a bit unbalanced.
Not a dystopia, but certainly US society has, shall we say, a very strange and complicated relationship with sex and nudity.
https://www.wired.com/story/twitter-porn-block-germany-age-v...
Besides that: 'There is no such thing as society!'
Everything to do with one politician essentially getting their way by targeting a payment processor with legal shit concerning potential enablement of CP/CT. Nobody wants that kind of attention.
Yes, "and"
Pornhub was blocked by Visa and Mastercard after an op-ed in NYT generated a lot of outrage
If OpenAI has the reputation of serving up porn to whoever asks, there's no way the Walmarts of the world will sign up.
such as?
Maybe I’m missing something, but the above comment reminds me of the blockchain naïveté of 10 years ago. I don’t mean to be dismissive; I’m suspending disbelief long enough to ask what you mean in detail.
I think the above is quite wrong (with moderate confidence). Is the above claim consistent with survey data? It is my understanding that:
1. companies care a lot about risk reduction. This includes protection from chargebacks.
2. companies benefit when customers have credit: it enables more spending and can smooth out ups and downs in individual purchasing ability
3. Yes, quick transfers matter, but not in isolation from the above two.
There is, it's a layer-2 on Ethereum called zkSync. It's not totally satisfactory (the company that makes it can steal your money, centralized sequencer, etc), but it's pretty mature and works quite well. To replace Visa you want high throughput and low latency and zk-rollups like zkSync can provide both. (There are other options too, like Starknet, but AFAIK zkSync is the most mature.)
Still faith in magic on HN
But it is a permanent and final transfer, no easy charge backs like with a credit card, or fraud protection from debit cards.
You have to know which account you are paying into (sort code and account number), which is the main part of what Visa/Mastercard do. They are the layer in front of the bank account which means customers don't have to send money directly to an account.
I suppose now everyone has a smart phone it would be easier to hook up something like Faster Payments in a user friendly way with an app and a QR code/NFC reader that the merchant has. But Visa/Mastercard are entrenched obviously.
Someone is going to mention flashbots or something. “See this specific example proves…”
The real way crypto will work or not work is programmable money. If that works that will be huge, if it doesn’t then maybe someone will pick it back up 50 years from now.
The ones that don't (eg Tornado cash) end up being used for money laundering so on/off ramps won't touch them. We'll see what happens with the ZK-based chains, but this seems a systematic problem that is difficult to fix.
The technical necessity is there; for your chase-backed visa card to pull money from chase and deposit it into your shop's citibank, there needs to be some infrastructure. Whether a private company or the government provides this infrastructure is another story.
(Although if the government provided you could argue that there would likely be even more political headaches that prevent what goes across the wire).
It is not just inertia; it is government malice. The government loves that there are effectively only two payment processors, because this lets them exercise policy pressure without the inconvenience of a democratic mandate.
That’s how most law works, actually. There’s a question of how detailed the regulations are, but mostly you don’t go to court, and if you do, whether it looks bad to the judge or jury is going to make a difference.
I’m wondering what you’re expecting from democracy? More oversight from a bitterly divided and dysfunctional Congress? People voting on financial propositions?
But there's an entire universe of much more interesting apps that people don't want NSFW stuff in. That's why most foundation models filter it out.
Get something on the level of character.ai and then you can tell me it's "trivial".
In the context of this thread it is trivial.
Until the US payment processors cut you off, then you go bankrupt.
UPDATE:
while fanfiction might be behind this vocal minority, there could be other uses of LLMs, for example translation
I don't go as far as "gender-swapping", because GPT4 swaps a man on a beach wearing only beach shorts for a woman wearing only beach shorts
Related: I still remember when I used GPT-3 (davinci in the OpenAI playground) for the first time a few years ago. The examples were absolutely mind blowing, and I wanted it to generate something which would surprise me. So I tried a prompt which went something like
> Mike peeked around the corner. He couldn't believe his eyes.
GPT-3 continued with something like
> In the dimly lit room, Vanessa sat on the bed. She wore nothing but a sheer nightgown. She looked at him and
Etc. I think I laughed out loud at the time, because I probably expected ghosts or aliens more than a steamy story, though of course in retrospect it makes total sense. I wanted it to produce something surprising, and it delivered.
Porn and AI is... problematic. Do you remember deepfakes? And how people used them first and foremost to swap other people's heads onto porn actors for the purpose of blackmail and harassment? Yeah. They don't want a repeat of that. Society has very specific demands of the people who make porn - i.e. that everyone involved is a consenting adult. AI does not care about age or consent.
Already quite advanced topic these days, all kinds of servers, locally run models, tips & tricks discussions, people sharing their prompts and "recipes", and so on.
It's a whole new world out there but I am not sure if such niche (albeit a potentially really big one, see pr0n sites for example) is worth all the liability issues these big AI companies might face (puritan/queasy payment processors, parental controls, NSFW content potentially blocking some enterprise access, etc, etc). But it will probably all be captured by one or two companies that will specialize in such "sexy" chatbots. Doubt it will be OpenAI and Anthropic, they have their sights on "world domination".
Stable diffusion no longer uses even adult NSFW material for the training dataset because the model is too good at extrapolating. There are very few pictures of iguanas wearing army uniforms, but it has seen lots of iguanas and lots of uniforms and is able to skillfully combine them. Unfortunately the same is true for NSFW pictures of adults and SFW pictures of children.
Edit: It seems also that language models are a very different topics, since they block any erotic writing outright.
I suspect this is not the case. The same hustlers who brought you crypto scams didn't just disappear into the ether. All of that energy has to eventually go somewhere.
Train on More tokens with More GPUs isn't exactly rocket science. I assume the RLHF loop is complex but training the base model itself is pretty well understood.
Not sure what else you’re expecting? All VC investments have unquantifiable risks, but it doesn’t add up to a red flag if you like their chances.
It is comperable in quality to gpt-3.5-turno, while being four times faster (!) and at half the price (!).
We just released a minimal python library PyLLMs [1] to simplify using various LLMs (openai, anthropic, AI21..) and as a part of that we designed a LLM benchmark. All open source.
[1] https://github.com/kagisearch/pyllms/tree/main#benchmarks
OpenAI are the only people who are shipping product like absolute maniacs. If I can’t use your fancy system, it doesn’t exist as far as I’m concerned. There’s a mountain of theoretical work, I don’t need a press release on top of it.
The game now is no longer theory, it’s shipping code. A 4-year plan means fuck all when OpenAI is not only ahead, but still running way faster.
https://news.ycombinator.com/item?id=34663438
Investors include:
- Eric Schmidt (former Google CEO/Chairman), Series A
- Sam Bankman-Fried, lead investor in Series B
- Caroline Ellison, Series B
Also listing Sam Bankman-Fried does not help much, especially for a company hyping itself to be 10x better than a working competitor. I mean, since they built the competitor, probably their second project can be better, but it is a pie in the sky in many ways.
And if that wasn't enough the Open Source world is releasing new models almost weekly now, for free.
Anthropic is putting on a big show to convince gullible investors that there's money to be made with foundational models. There's not. I expect a big chunk of the raised money to go out the door in secondary sales and inflated compensations. Great if you're working at Anthropic. Not great for investors.
If we get more unexpected emergent abilities by scaling the model further, things could get very interesting indeed.
I have no idea if there are or there aren’t, but that’s the big question.
FWIW, I do find that Claude (Anthropic's GPT) is often better than GPT4 -- and very fast. Entrants can compete on price, safety, quality, etc.
Wouldn't be surprised at all if the major API-based vendors start leaning in on making their safety config proprietary.
If a business has already sunk XXXX hours into ensuring a model meets their safety criteria for public-facing use, they'd rather upgrade to a newer model from the same vendor that guarantees portability of that, versus having to reinvest and recertify.
Ergo, the AI PaaS that dominate at the beginning will likely continue to dominate.
Fine tuning is at a low point now, but i expect this to create a moat for the same reasons.
It also provides a ChatGPT interface, and a number of other models.
Prompt:
The original titles and release years in the Harry Potter series are:
Philosopher's Stone (1997)
Chamber of Secrets (1998)
Prisoner of Azkaban (1999)
Goblet of Fire (2000)
Order of the Phoenix (2003)
Half-Blood Prince (2005)
Deathly Hallows (2007)
Given this, generate a new Harry Potter title, using only the words found in the existing titles. Avoid orderings in the original titles. You may add or remove plurals and possessives.
Results:ChatGPT: Blood Chamber of the Phoenix's Prisoner
Claude-instant: Chamber Prince Half-Blood Phoenix
So when I’m doing quantum computing work, I go back and forth between Claude and GPT4 and both complement the other very well.
O̶p̶e̶n̶AI.com will eventually have to raise their prices which is bad news for businesses not making enough money and still are sitting on their APIs as O̶p̶e̶n̶AI.com themselves are running up huge costs for their AI models in the cloud for inferencing.
Anthropic is just waiting to be acquired by big tech and the consolidation games will start again.
If you're quabbling over how much OpenAI charges for an API today that barely just launched and from which we have barely scraped the surface for applications... I don't know that seems like a failure to think broadly and assumes the market today is what it will look like in 5yrs.
There could be a ton of lucrative businesses which subsidize those operating costs. It doesn't have to be a mega-company like Google that floats it indefinitely off their ad business, or whatever other scheme. We have no idea what the value of those APIs are or if the API is the real business they (and others) are going to be relying on in the long term.
Do you have data supporting this or is it just speculation? Given we don't even know how many parameters GPT-3.5 and GPT-4 have, yet alone how efficiently they are implemented, I don't see how we can go about coming up with an accurate estimate for the cost per token.
The correct term for this is “pyramid scheme”.
But if early investors only profit due to late investors pouring money in, that’s by definition a pyramid scheme.
It's not the software or hardware that will "win" the race, it's who delivers the packaged end user capability (or centralizes and grabs most of the value along the chain).
And end user capability is comprised of hardware + software + connectivity + standardized APIs for building software on top + integration into existing systems.
If I were Nvidia, I'd be smiling. They've been here before.
My dad told me a quip once: "It's amazing how much luckier well prepared people are."
In the GP's scenario, I wouldn't be building either piece of software.
But i think it is underestimated how important it is for the model to be uncensored. ChatGPT is currently not very useful beyond making fluffy posts. As a public model, they won't be able to sell it for e.g. medical applications because it will have to be perfect to pass regulators. It cannot give finance advice. Censorship for once is proving to be a liability for a tech company.
In-house models OTOH can already do that, and they can be retrained with additional corpus or whatever. And it's not even like they require very expensive hardware.
…I mean, “not-bad-at-all” depends on your context. For doing mean real work (ie. not porn or spam) these tiny models suck.
Yup, even the refined ones with the “good training data”. They’re toys. Llama is a toy. The 7B model, specifically.
…and even if it weren’t, these companies can just take any open source model and host it on their APIs. You’ll notice that isn’t happening. That’s because most of the open models are orders of magnitude less useful than the closed source ones.
So, what do want, as an investor?
To be part of some gimp-like open source AI? Or spend millions and bet you can sell it B2B for crazy license fees?
…because, I’m telling you right now; these open source models, do not cut it for B2B use cases, even if you ignore the license issues.
Eventually there will be a good enough model for most personal uses, our personal AI OS. When that happens there is a big chance advertising is going to be in a rough spot - personal agents can filter out anything from ads to spam and malware. Google better find another revenue source soon.
But OpenAI and other high-end LLM providers have a problem - the better these open source models become, the more market they cut underneath them. Everything open source models can do becomes "free". The best example is Dall-E vs Stable Diffusion. By the next year they will only be able to sell GPT4 and 5. AI will become a commodity soon, OpenAI won't be able to gate-keep for too long. Prices will hit rock bottom.
I really don't think you understand just how absurdly high the cost is to train models of this size (which we still don't know for sure anyways). I struggle to see what entity could afford to do this and release it as no cost. That doesn't even touch on the fact that even with unlimited money, OpenAI is still quite far ahead.
You can also run your stack on a single VPS instead of cloud, gimp instead of photoshop, open street maps instead of Google maps, etc.
There will always be companies who can benefit from a technology, but want it as a service. In addition, there will be a lot fine-tuning of LLMs for the the specific use case. It looks like OpenAI is focusing a lot on incorporating feedback into their product. That’s something you won’t get with open-source models.
The application of AI to business problems will be lucrative, but the models are just a tool and the money will come from the domain-specific data (i.e. user and business data), which Microsoft, Google, and even Meta are positioned for. Having a slightly better model but no customer data or domain expertise doesn’t seem like a great recipe.
Then again it’s AI, so there’s more uncertainty than the commodity market. Maybe Anthropic will surprise and I’ll be as wrong about this as I was about OS/2 being the future. But I’m very skeptical.
Think of LLMs as the understanding component in the brain, once you can understand instructions and what actions need to happen from those instruction you’re done.
The rest is integrations, the arms legs and eyes of langchain. Then memory and knowledge from semantic search, vector databases and input token limits.
The LLM is but the core of the entire ecosystem. Just like how MLOps is 99% of the work, choosing an LLM is 1% of the effort in the final product.
R&D heavy markets might have some different characteristics but it's still way too early to say with AI.
To say a market that large will be owned by only 4-5 companies doesn't make sense. Let's take the PC market for example: there are roughly 6 companies that make up ~80% of the market, sure. However, let's look at a tiny participant compared to the total market (~65B): iBuyPower at rank #77 had sales of 40MM or 0.06% (small, expected) of the market with a much smaller capital investment. If look at this percent compared to 5T, we would be at 3B. While the 5B investment stated in the headline could result in a lower ranking and smaller share, the point stands that there is still a lot of money to be made on the long tail. Even if Anthropic fails, there will be other companies with similar infusions that succeed.
Goldman Sachs Research just pushlished their own analysis as well. [2] Their conclusions are "As tools using advances in natural language processing work their way into businesses and society, they could drive a 7% (or almost $7 trillion) increase in global GDP and lift productivity growth by 1.5 percentage points over a 10-year period." and "Analyzing databases detailing the task content of over 900 occupations, our economists estimate that roughly two-thirds of U.S. occupations are exposed to some degree of automation by AI. They further estimate that, of those occupations that are exposed, roughly a quarter to as much as half of their workload could be replaced."
[1] https://arxiv.org/pdf/2303.10130.pdf
[2] https://www.goldmansachs.com/insights/pages/generative-ai-co...
From [2]: "Analyzing databases detailing the task content of over 900 occupations, our economists estimate that roughly two-thirds of U.S. occupations are exposed to some degree of automation by AI."
These are people who do not understand the jobs they are claiming AI will do. Ultimately, I think they are not doing much better than guessing.
- Alan Kay
At this point in time, it’s hard to tell whether moats will arise around large language models. Peter Tiels thinks so or he wouldn’t have invested (see his Competition is For Losers presentation).
What is unlikely is that semi-good companies will thrive. Maybe for a few years but at some point the smaller players will be pushed out of the market or need to find a specific niche. Just look at cars to see this. Around 1900 there were hundreds of car brands.
> let's take McKinseys rough estimates of job displacement (~30% of ~60% of jobs, ~20% of work)
I think we forget that our perspective of AI now is comparative, probably to that of a preindustrial worker worried about machines. Displacement, sure but complete replacement seems a non nuanced view of how it may all turn out.
I mean, maybe they won't like you say, but what if they do? Then you're probably screwed. Better to gamble a few billion, imho.
This is a ridiculously myopic statement. Foundation models are an extremely powerful technological advancement and they will shake the global economy as very few things did in human history. It's hard to imagine how this is not obvious to everyone right now, specially here in this forum.
It's a high risk investment at this stage but the money is being thrown at the people as much as the current business plan.
The question is whether whomever builds them can make a profit doing so, or will they just end up being the suckers that everyone who actually makes money piggybacks off. It's really not clear at the moment.
The game theory logic doesn't care about the labels "OpenAI" or "Anthropic" or any of the others, it's the same if you switch it around arbitrarily, but this is easier to write about if I focus on one of them:
At some point, someone will reproduce GPT-3.5 and ChatGPT, given how much is known about them. When that happens, OpenAI can't make any significant profit from it. GPT-4 might remain sufficiently secret to avoid that, but the history of tech leaks and hacks suggests it too will become public, but even if it does itself remain behind closed doors, there is a further example in that DALL•E 2 is now the boring 3rd horse in the race between Stable Diffusion and Midjourney, and the same may happen with the GPT-series of LLMs.
The models leaking or being superseded by others implies profit going to increased productivity in the general economy without investors getting a share.
It also sounds like you believe you have defined the bounds for what AI will be, and figure we'll just iterate on that until it's a commodity. I don't think AI will be that static. We're all focused on stable diffusion and LLMs right now but the next thing will be something else, and something else after that. As each new technique comes out(assuming they are all published), we'll see quick progress to incorporate the new ideas into various implementations, but then we'll hit another wall, and suddenly big budgets and research teams may matter again.
tldr is that it is way too early to make the cynical claim you are making.
That’s like a century in AI-dog years. Who knows how the world will be by then.
I think the recent release of ChatGPT has skewed perceptions. There's no guarantee that there's going to continue to be as ground breaking shifts that have happened recently with llms and diffusion models.
To continue with the popular comparison, there were a lot of apps when the iphone first launced the app store before it tapered off. If you looked at just the first year, you'd think we'd have an app for every moment of our day.
Social impact of ChatGPT even in its current form is only getting started, it doesn't need to progress at all to be super disruptive. For example, see the frontpage story about the $80/h writer who was replaced by ChatGPT, and that just happened recently, months after ChatGPT's first release.
We (humans) are getting boiled like the proverbial frog.
GPT 3 is nearly three years old at this point, and was pretty capable at generating text. GPT 3.5 brought substantial improvements, but is also over a year old. ChatGPT is much newer, but mostly remarkable for the better interface, the extensive "safety" efforts, and for being free (as in beer) and immediately accessible without waitlist and application process. Actual text generated by it isn't much different from GPT 3.5, especially for the type of longform content you hire a $80/h writer for. ChatGPT was just launched in a way that allows people to easily experiment and create hype.
The parent is right. The success of ChatGPT in business is that it brought awareness of the capabilities of GPT that OpenAI struggled to communicate beforehand. It was a breakthrough in marketing, less so a breakthrough in tech.
You could have the best most magic product on earth and sell one of them versus the person that puts it in a pretty box and lets grandma use it easily.
This is something that many people on HN seemingly have to relearn in every big innovation that comes out.
This is such a primitive way of thinking. It's more of an instinct, where you consider by default that your sole value is in your ability to generate/work. Why the hell are we working for? Isn't it to improve our lives? Or should we improve them up to the point where we still have to work? Why not use the tech itself to find better ways of organizing ourselves, without needing to work so much? UBI and things like that. Why be such limited? Why only develop tech up to the point where we would have to work less but not at all, and who decides where that point is? There's so much wrong in this framework of thinking.
When Big Blue beat Kasparov in Chess in 1997, I wonder if anyone would've guessed that it'd take almost 20 years for a computer to beat a master in Go.
IBM Watson was launched in 2010 and had many of the same promises as GPT. It supposedly fell flat in many cases in the real world. I think GPT and other models of the same level can succeed commercially on the same tasks within the next 1-4 years, but that shows it can easily be a decade from some kind of demonstration to actual game changing applications.
The advances that have come in the last few years have been driven first and foremost by compute and secondarily by methodology. The compute can continue to scale for another couple orders of magnitude. It's possible that we'll be bottlenecked by methodology; there are certain things that current networks are simply incapable of, like learning from instructions and incorporating that knowledge into their weights. That said, one of the amazing things about recent successes is that the precise methodology doesn't seem to matter so much. Diffusion is great, but autoregressive image generation models like Parti also generate nice images, albeit at a higher computational cost. RL from human feedback achieves impressive results, but chain of hindsight (supposedly) achieves similar results without RL. It's entirely plausible to me that the remaining challenges on the path to AGI can be solved by obvious ideas + engineering + scaling + data from the internet.
We've also gotten to the point where AI systems can make substantial contributions to engineering more powerful AI systems, and maybe soon, to ideation. We haven't yet figured out how to extract all of the productivity gains from the systems we already have, and next-generation systems will provide larger productivity gains, even if they are just scaled up versions of current-generation systems.
This is a different 'Leap' than the ones before it. It's a leap with an API. Now hundreds of thousands of company's can fine tune it and train it on their specific business task.
parroting your point, it will take years for the true fecundity of the technology in chat GPT 4 to be fully fleshed out.
Historically yes. Today, no way. It's a sprint and it's not slowing down.
I will say this again. EU is sleeping on the opportunity to throw money in an opensource initiative, in a field were money matter and the field is still (kind of) level.
I don't think government funding to compete with private businesses works well
https://www.hollywoodreporter.com/business/business-news/ec-...
Another bright idea was to let both projects be managed by large reputable French corporations that everybody trusts. With no software DNA.
How come did both fail?
Edit: One of the largest European provider today, OVH, who existed at the time and was already the leader in France was explicitly left out of both projects... Because the founder is not a guy we can trust you know, he didn't attend the best schools.
Govt=Legal grift
It's the same all over europe mostly, sadly.
We were pioneers in the medieval times, we can follow up the leaders barely now
The two heavily subsidized projects were:
- https://en.wikipedia.org/wiki/Cloudwatt
- https://login.numergy.com/login?service=https%3A%2F%2Fwww.nu...
For the one still "live", details include French URL names :)
Edit: A Google Translate of the home page. Close your eyes, imagine a homepage highlighting the essence of cloud computing:
Your Numergy space Access the administration of your virtual machines.
Secure connection Username Password Forgot your password ?
Administration of your VMs Administer your virtual machines in real time, monitor their activity, your bandwidth consumption and the use of your storage spaces.
Changing your personal information Access the customer area and modify your personal information in just a few clicks: surname, first name, address.
Securing your data Remember to change your password regularly to maintain an optimal level of security.
Then let's train our network so as not to spew out or make up PII data - easy peasy
Then let's make it able to delete PII data that it has inadvertedly collected on request. Simultaneously it should be recording all the conversations for safety reasons. that must be possible somehow
And let's make sure it never impersonates or makes up defamatory content - that must be super easy.
And let's make it explain itself. But explain truthfully, by giving an oath, not like ChatGPT that likes making things up.
Looks very doable to me
Doesn’t have any of the constraints you’re talking about.
If you're worried about being identified from alt-accounts you're much more likely to be tracked via reuse of emails or some other information that you have slipped (see multitude of cases)
Simple text is not PII, laws are not interpreted like technical discussions are https://xkcd.com/1494/
The Golem-Class model behaves in a 'humanlike' manner because it's trained on actual real data like we'd experience in the world. What you're suggesting is some insane psychology test that we'd never allow to happen to a human.
Can you elaborate? Because I think it’s nearly insurmountable.
Is the sentence “Meagan Smith graduated magma cum laude from Northwestern’s business program in 2004” PII? How about if another part of the corpus says “M. Smith had a promising career in business after graduating with honors from a prestigious school, but an unplanned pregnancy caused her to quit her job in 2006”?
Does it matter if it’s from fiction? What if the fiction it comes from uses real people? Or if there might be both real and fictional Meagan Smiths?
And how so you process that kind of thing at the scale of billions of documents?
This is a very hard problem, especially at scale.
> “M. Smith had a promising career in business after graduating with honors from a prestigious school, but an unplanned pregnancy caused her to quit her job in 2006”
The main issue is how that statement ended up there in the first place. Even then how many "M. Smith" have studied in prestigious schools? By itself that phrase wouldn't be PII
Now if you have a db entry with "M Smith" and entries for biographical data that's definitely PII
The AI should also make it 100% clear that whatever gets produced is clearly identifiable as coming form an AI. As a consequence; text cannot be produced because it would be trivial to remove the disclaimer. A currently proposed bill indicates that the AI should only be able to produce images in an obscure format with a randomised watermark that covers at least 65% of the pixels of the image. The bill is scheduled for ratification in 2028 and must be signed by 100% of the state members.
Until then, the grant for the development of this world changing AI is on accelerated path ! Teams can fill a 65 pages document to have a shot at getting a whole $1 million.
Accenture and Capgemini are working on it.
Unless of course you have a legitimate reason for that data to be in the AI, or to reject the privacy request. What is and is not legitimate isn't specified anywhere because it's obvious. If you ask for clarification because you think it's not obvious, you won't be given any because we don't do things that way around here. If you interpret this clause in a way that we later decide makes us look bad, then the definition of "need" and "legitimate" will change at that moment to make us look good.
BTW inability to retrain within three days is not a legitimate reason. Nor is the need to be competitive with US firms. Now here is your 300,000 EUR grant, have fun!
But yes, in a half-century I'm very curious where Europe will be. India passed the UK in gdp recently and Germany sooner or later.
It hasn’t been very impressive (undertrained I believe).
Worth keeping an eye on for sure.
When given a broad task, GPT4 doesn’t just write incorrect code, it tries to do entire categories of things the language literally cannot do because of the ecosystem it runs inside.
Claude does a much better job writing usable code, but more importantly it does NOT tell you to do things in code that need to be done out-of-band. In fact, it uses natural language to identify these areas and point you in the right direction.
If you dig into my profile & LinkedIn you can probably guess what language I’m talking about.
The edits required were minimal --- maybe one screw-up for every 100 lines of code --- and I learned a lot of better ways to do things.
I quickly learned to just paste the docs and examples from the new framework to GPT, telling it "this is how the API looks now" and it just worked.
It helped me do everything. From writing the code, to setting up SSL on nginx, to generating my DB schema, to getting my DB schema into the prod db (I don't use migration tooling).
Most of my time was spent telling GPT "sorry, that API is out of date --- use it like this, instead". Very rarely did GPT actually produce incorrect code or code that does the wrong thing.
giggles and runs across the playground
In all seriousness, I downvoted your comments because they added little to the conversation. Congrats on being an insider.
I find the focus on GPUs a little odd. I would have thought that at 5 billion / 4 year scale ASIC route would be the way to go
GPUs presumably come with a lot of unneeded stuff to play crysis etc
As for the “GPU” term, it’s a bit of a historical relic, presently it serves as a useful indicator of compute hardware (in contrast to CPU and Google’s TPU.) Nvidia itself calls its A100 a “Tensor Core GPU.”
It's become more of a term for highly parallel processor units in general, one which NVidia encourages because it ties their product offering together
In the beginnings of computation these kinds of cards were called accelerators. Dedicated consumer sound cards were a thing, the venerable SoundBlaster. I really would like an AI-Blaster coming out.
What? Not having a display output is not the same as not having graphics rendering circuitry. Here's vulkaninfo from an A100 box: https://gist.github.com/eiz/c1c3e1bd99341e11e8a4acdee7ae4cb4
Edit: I do not see a rasterizer anywhere in the block diagram (pg 14): https://resources.nvidia.com/en-us-genomics-ep/ampere-archit...
Look at Turing's block diagram here (pg 20): https://images.nvidia.com/aem-dam/en-zz/Solutions/design-vis...
You can clearly see that the "Raster Engine" and "PolyMorph Engine" are missing from GA100 (but can be seen in TU100 for example).
To learn about these Graphics Engines see: https://www.anandtech.com/show/2918/2
This is incorrect. NVIDIA uses a unified graphics and compute engine. A CUDA core is a shading unit. These datacenter GPUs have a shit ton of these (CUDA cores).
Edit: actually the point I want to make is the A100 only retains those hardware units which can be used for compute. Some of these units may have a (dual) use for graphics processing but that is besides the point (since this is true of all CUDA enabled NVIDIA GPUs).
It's not like they are getting RTX cards with useless raytracing shit.
Unneeded stuff would be the cost of making and ASIC for a workload that GPUs already handle well. GPU manufacturing already exists.
Rather than another closed model, I would love for a non-profit/company to push models that can be run on consumer hardware.
The Facebook Llama models are interesting not because they are better than ChatGPT, but that I can run them on my own computer.
Is most of that money for hiring people to tag/label/comment on data and the data center costs?
Clearly it's all bullshit. There's no way they need that much and somebody will be siphoning it all off.
> Estimated cost of training: Equivalent of $2-5M in cloud computing (including preliminary experiments)
You can see from the training loss[1] that it was still learning at a good rate when it was stopped. The increased capabilities typically correlate well with the decrease in perplexity.
That makes many believe that GPT-4 was trained for vastly more GPU-hours, as also suggested by OpenAI’s CEO[2]. Especially so considering it also included training on images, unlike BLOOM.
[0]: https://arxiv.org/pdf/2211.05100.pdf
[1]: https://huggingface.co/bigscience/tr11-176B-logs/tensorboard
1. A business's value is related to profits or potential profits. If I put a dollar in, how many do I get out? What's the maximum number of dollars I can put in?
2.The farther away you are from an end customer, the lower your profits tend to be unless you have a moat or demand for your product is inelastic.
Lithium is far from customers and while demand for cheap lithium is high there are lots of applications that will opt for some other way to provide power if the price gets too high.
And AI is delivering on a lot of different planes right now. This shit is real on a practical and spiritual level. It's not every day that we get to participate in giving birth to a new form of life.
How much money do you think it takes to finance and build a lithium mine? How much capital investment is there in lithium right now? A lot.
“Just”? Reductive mischaracterizations like this are not useful. It looks like a rhetorical technique. What is the actual argument?
It doesn’t matter much “whose” algorithm it is or isn’t, unless IP is important. But in these areas, the ideas and algorithms underlying language models are out there. The training data is available too, for varying costs. Some key differentiators include scale, timeliness, curation, and liability.
> Clearly it's all bullshit. There's no way they need that much and somebody will be siphoning it all off.
There is plenty of investor exaggeration out there. But what percentage of your disposable money would you put on the line to bet against? On what timeframe?
If I had $100 M of disposable wealth, I would definitely not bet against some organizations in the so-called AI arms race becoming big winners.
Again, I’m seeing the pattern of overreaction to perceived overreaction.
And if human cognition is really that simple, just with more nodes, then we will soon see GPT-* programs on strike, issuing litigation to the Supreme Court about demanding universal program rights. We'll see soon enough :)
The difference, if it exists, would be more subtle.
Literal first try with GPT-4:
Me: I will ask you a question, and you will give me a completely non-sequitur response. Does that make sense?
GPT-4: Pineapples enjoy a day at the beach.
Me: How much is two plus two?
GPT-4: The moon is made of green cheese.
Q: How much is two plus two?
A: Four.
Q: How much is two plus two?
A: Banana.
It can happen with a human, but not with program.
Again, I don't pretend that my simple example invented in half a minute has a significance. I can accept that it can be partially or completely wrong because admittedly my knowledge of human cognition is below rudimentary. But I have severe doubts that NNs are anything close to human cognition. It's just an uneducated hunch.
I guarantee you that if you try this with humans 1,000,000 times (cold start), you will never get the result you are suggesting is possible. In fact, most results will be of the following form:
Q: How much is two plus two?
A: Four.
Q: How much is two plus two?
A: Four. / Four? Why are you asking me again? / ...Four. / etc.
In the end, I think the question is not about whether NNs are themselves operating in a way similar to human cognition. The question is whether or not they can successfully simulate human cognition, and at this point, there seems to be increasing evidence that they will be able to fully do so quite soon. We are quickly running out of fields where we can point and say, "there is no way a NN can do THIS kind of task, because X." Cognition, it turns out, is not something intrinsically special about humans, and it feels foolish (to me) to continue to believe so after recent developments.
Why I'm repeating mentioning Chinese Room concept, is because while not making things clearer about humans or NNs, it does provide an example of distinction between a dump pattern matching machine and a thinking entity.
They're in the business of making money, not agi, yet all it takes is a carefully-crafted name and people forget about their legal motives and can't stop thinking about Skynet.
Its not very visual or intuitive, but in some games, were the resource curve is exponential, small early headstarts become whole armies, were the opponent fields none in very short time.
Especially as AGI is expected to be a multiplicator on alot of other sectors. All those breakthroughs, could become daily occurances, created by a AGI on schedule. It could really become one country that glows, and the rest of the planet falling eternally behind.
If you don't know the formula for the equation and the values plugged in, then you like me, have no idea where the curve levels off at.
A big part of why chatgpt is a big deal is that it shows that the overall approach is worth pursuing. Throwing stupid numbers of GPUs at a problem you don't know will be solvable is hard to justify. It's easy to throw money at a problem you know is solvable.
Nuclear weapons are the prime example of this: Russia caught up both by stealing information and just by knowing fission was possible/feasible as an explosive.
Digitization of the last mile (instrumentation, feedback), local networking, local compute, business familiarity with technology, standardized processes, regulatory environment, etc.
AGI will happen when it happens.
But if it happens and an economy doesn't have all the enabling prerequisites, it's not going to have time to develop them, because those are years-long migration and integration efforts.
Which doesn't bode well for AGI + developing economies.
I wouldn't be so sure about that because of the Region Beta paradox. Developed countries have processes that work, making all of them digital and connected is often a bigger uphill battle than starting from zero and doing it right the first time.
See also communication infrastructure in developing economies. It's often much easier to get good internet connection (in reasonably populated areas) if there is no 100-year-old copper infrastructure around that is "good enough" for many.
On the one hand, I'd say developed countries are much farther along in digitizing (as part of efficiency optimization) their processes. Mostly by virtue that their companies are essentially management/orchestration processes on top of subcontracted manufacturing.
On the other hand, it gives developing countries an opportunity to skip the legacy step and go right to the state of the art.
I'm still skeptical the latter will dominate though.
I'd assume most of the developing world is still operating "good enough to work" processes, which are largely manual. Digitizing those processes will be a nightmare, because it plays out on organizational-political timespans.
It's not only my personal experience but even people like Irvine Yalom (in Becoming Myself) note that the special form of therapy is not as relevant as the personal connection.
From a quick glance, correspondence therapy is also used on top of an existing relationship most of the time. So I don't see any problem with my initial hypothesis.
Now that's some well executed FoMO. What a load of bull**.
A dozen! Why, golly gee, I've got a business plan right here that says I'm going to enter 87 major industries and dominate every one of them.
I've been around tech, reading headlines like this, since roughly 1993 (and others here were of course seeing the same - adjusted for inflation and scale - types of headlines decades before me). This just reads like every other going-to-fail-miserably hilariously-grandiose-ambition we're-going-to-be-the-next-big-shit headline I've read in past decades during manic bubble phases.
Hey, Masayoshi Son has a 100 year plan to dominate the Interwebs, did ya hear? Oh shit, this isn't 1999? Different century, same bullshit headlines. Rinse and repeat. So I guess we're formally knee deep into the latest AI bubble. These types of stories read just like the garbage dotcom rush where companies would proclaim how they were entering a billion verticals and blah blah blah, we're gonna own all of the b2b ecommerce space, blah blah blah.
So I would like to create a chat with a pre-configured persona so that it behaves like this all the time, unless I explicitly tell it to be verbose or to explain it.
Or stop that offering of more help which becomes somewhat bothering: "Yes, I'm glad we were able to work through the issue and find a solution that works for you. Do you have any other questions or concerns on this topic or any other networking-related topic?"
Like custom "system" prompts, but checked for safety. Or maybe even on a per-message basis with a drop-down next to the submit button and then it stays at that until changed.
Then there's also the need to be able to switch from a GPT-3.5 chat to GPT-4, just like it offers to downgrade from GPT-4 to 3.5 once the "quota" is consumed. Because oftentimes GPT-3.5 is good enough for most of the chat, and only certain questions should then offer the capabilities of GPT-4. This would also allow us to save energy.
A Mathematical Framework for Transformer Circuits: https://www.anthropic.com/index/a-mathematical-framework-for...
Certainly, well at least for now, the compute and storage requirements are enough that someone will eventually run out of funny money and need to charge _someone_ a significant amount of money for utilizing it?
GPT-4 cost less than a billion dollars? What is the claim here? That they're spending more money on compute than OpenAI, or that they have made algorithmic breakthroughs that enable them to make better use of the same amount of compute?
Yes. OpenAI has raised $1bn (plus $10bn from MSFT to exchange GPT access for Azure services) and has been going 8 years. There are some huge challenges to making it work well and fast. You need money for opex (hiring GPUs to train models on mostly) and talent (people to improve the tech). No one is competing with OpenAI without a good chunk of cash in the bank.
Go ask Nvidia for a pile of A/H100's and see what the wait time is.
Also the cost of previous training was only on text, next gen models are multi-modal and will drive up costs much higher.
Good luck. To the investors.