OpenAI is too cheap to beat
generatingconversation.substack.com
generatingconversation.substack.com
This is Uber/AirBnB/Wework/literally every VC subsidized hungry-hungry-hippos market grab all over again. If you’re falling in love because the prices are so low, that is ephemeral at best and is not a moat. Someone try calling an Uber in SF today and tell me how much that costs you and how much worse the experience is vs 2017.
OpenAI is the undisputed future of AI… for timescales 6 months and less. They are still extremely vulnerable to complete disruption and as likely to be the next MySpace as they are Facebook.
AI models have some GPU constraints but could easily reach a state where the cost to opperate falls and becomes relatively trivial with almost no lowerbound, for most use cases.
You are correct there is a race for marketshare. The crux in this case will be keeping it. Easy come, easy go. Models often make the worst business model.
It is for the consumers of these models, is there even a point to train your own or experiment with OSS!
If you need chatGPT4-quality responses they aren't close yet, but it'll happen.
They trained a custom model for this. Better accuracy, sure, but I was a little surprised to watch how much faster it is than GPT4.
Based on their testing, they’ve become believers in domain specific smaller models, especially for performance.
However, I am skeptical that smaller and more focused models will do well for more general tasks. I’ve found gpt-4 to just be flatly excellent for tasks that involve emitting a custom DSL and customer-provided business context (the latter being something we cannot train for, not without immense expense).
Better beyond a certain point is unlikely to be competitive with the cheaper models.
In that case you might be able to get linguistic competence with a much smaller model that you end up training with a smaller, cleaner, and probably partially synthetic data set.
Add all the downloadable material at archive.org and you've got a formidable corpus.
What I would really want though is an uncensored LLM. OpenAI is basically unusable now, most of its replies are like "I'm only a dumb AI and my lawyers don't want me to answer your question". Yes I work in cyber. But it's pretty insane now.
https://arxiv.org/abs/2305.08377
and shows how LLM technology has a lot more to offer than "ChatGPT". The real takeaway is that by training LLMs with real training data (even with a "less powerful" model) you can get an error rate more than 10x less than you get with the "zero shot" model of asking ChatGPT to answer a question for you the same way that Mickey Mouse asked the broom to clean up for him in Fantasia. The "few-shot" approach of supplying a few examples in the attention window was a little better but not much.
The problem isn't something that will go away with a more powerful model because the problem has a lot to do with the intrinsic fuzziness of language.
People who are waiting for an exponentially more expensive ChatGPT-5 to save them will be pushing a bubble around under a rug endlessly while the grinds who formulate well-defined problems and make training sets will actually cross the finish line.
Remember that Moore's Law is over in the sense that transistors are not getting cheaper generation after generation, that is why the NVIDIA 40xx series is such a disappointment to most people. LLMs have some possibility of getting cheaper from a software perspective as we understand how they work and hardware can be better optimized to make the most of those transistors, but the driving force of the semiconductor revolution is spent unless people find some entirely different way to build chips.
But... people really want to be like Mickey in Fantasia and hope the grinds are going to make magic for them.
I am unconvinced by the idea of trying to redefine Moore's Law to be about MSRP. The NVIDIA H100 has twice the FLOPS of the A100 on a smaller die. That's Moore's Law, full stop. When NVIDIA has useful competition in the AI space, they'll be forced to cut prices, as has reliably been the case for every semiconductor vendor for the last 60 years.
You say that training datasets will win, but this is where OpenAI is currently have a big leg up: Everyone is dumping tons of real data into them, while the LocalLLM crowd is using GPT-4 to try to keep up.
We will see who is faster.
Neural nets don't need fully precise digital computing. Especially with quantization we're seeing that losing a bit of precision in the weights isn't impactful. Now that we're serving huge foundation models with static weights there's an enormous incentive to develop analog hardware to run them.
Mark my words, this will lead to a renaissance in analog computing, and in the future we will be shocked at the enormous waste of having run huge models on digital chips.
Just think, how many multiplications per second is the light refracting through your window right now clocking? More or less than is required to ChatGPT do you think? If only the crystals were configured correctly and the patterns of light coming through could be interpreted...
It really depends where you draw the lines, because you could also say that one single transistor in my electrical CPU is doing a kerjillion calculations for all of the atoms and electrons involved.
You have to create a prompt/function that for a wide set of inputs, generates a token sequence that will perpetually expand in a manner that corresponds to an externally observed truth.
Way too often it feels like you have to shove a universal decoding sequence into a prompt.
“Talk your steps, list your clues, etc.”
Just trying to luck into a prompt that keeps decompressing the model/ generating the next token that ensures the next token is true.*
I recall there was a paper with a relevant title recently… https://arxiv.org/abs/2309.10668
Basically - LLMs don’t reason , they regurgitate. If they have the right training data, and the right prompt, they can decompress the training data into something that can be validated as true
——-
* Also this has to be done in a limited context window, there is no long term memory, and there is no real underlying model of thought.
Fascinating stuff - a synchronous distributed system allows treating 1000 chips as one, knowing exactly when data will arrive cycle-for-cycle and which network paths are open. The compiler can balance loads. No more indeterminism or complexity in optimizing performance (high compute utilization). A few basic operations suffice, with the compiler handling optimization, instead of 100 kernel variants of CONV for all shapes. Of course, it integrates with Pytorch and other frameworks.
For example, I struggle to see DALL-E winning over Firefly if Firefly is integrated into a very rich environment, whereas DALL-E is basically prompt UI only (while DALL-E 3 is a better model IMO).
Both are B2C and have network effects.
It's a race to the bottom on pricing on the provider/infra side. It seems very unlikely that any single LLM provider will achieve a sustained and durable advantage enough to achieve large margins on their product.
Consumers can swap between providers with relative ease, and there is very little stickiness to these LLM APIs, since the interfaces are general and operate on natural language. Versus something like building out a Salesforce integration and then trying to update that to a competitor. Or migrating from Mongo to DynamoDB.
Building the LLMs is where the cool tech lives, but surprised so many are seeing that as a compelling investment opportunity.
But, let's see!
Cloud infra may be a comparable market, since computation is a big share of AI costs. Did consumers win big from competition between AWS, Azure, and GCP? Not sure. I see an uptick in write ups saying “We switched off cloud and reduced costs by 2/3rds.” Not a scientific sample but may leave the question open.
Margin is a function of stickiness/cost of switching (among other things).
I suspect eventually we will enter a world where migrating cloud providers is mostly a click of a button, but we're a long ways off from that. Requires vendor agnostic and portable apis/containers/WASM runtimes everywhere.
Swapping an LLM, at least in the current state, is about as close to updating a pointer as you can get.
Especially given that many orgs use managed or cloud specific solutions that have no 1:1 mapping between vendors
As someone who worked on a similar project recently, I'm getting that idea that you obviously don't know what you're talking about.
You need to plan ahead, sure. Portability was one of my main concerns (including the possibility to go self-hosted). But it's definitely not impossible, nor too hard to do.
And Kubernetes? It goes way beyond containers, you can replace many cloud provider-specific resources with Kubernetes resources. What alternative gives you that?
This will almost certainly lead to a prompt rebuild, to better accommodate new model idiosyncrasies.
If you are unlucky- your use case may be one where Evals require human review.
Unless you are YOLOing it without evals. In that case this is relevant.
Learning aws vs learning how to operate a vps are of comparable complexity
It took me less then a day to setup infra for my startup (more than 10 years ago)
But tell me what third tier cheap provider has managed scale-to-zero-or-infinity functions? Managed storage with S3-like API? Where can I get an API gateway cheaper than Amazon? What about managed databases? These tools allow me to develop insanely scalable software incredibly easily.
Agreed - all the big clouds are very expensive VPS hostings. Don't use it for that.
You can scale from 1 thread 512MB RAM, to 500 threads and 12TB of RAM (off the shelf). Which is good enough for almost everyone who isn't planet scale.
Auto scaling also comes with auto billing. Oops, your accidental infinite loop spawning functions has bankrupted your company. You don't have that risk starting with a VPS.
What's third tier here? Pretty much all of them offer S3 compatibility it's basically table stakes
I try to not use anything else if I can avoid it on a new project.
Ingress gives you the ability to load balance and is threaded and will scale with network transfer and cpu. Database should scale with a bump in VM specs as well, CPU and disk IOPS.
If you keep your app server stateless you can simply multiply it for the number of copies you need for your load.
Systemd can keep your app server running, it you docker it up and use that
This right here so many times over.
I'm not going to spend time worrying about scaling, I'm going spend time figuring out how to make stuff stateless.
For new small systems I suggest you start with a cloud VPS. If traffic is low, cost of downtime is low, and system requirements are high then a cheap mini PC ($150) at the home or office can keep your bill microscopic. If your app server and database are small then you can just throw them on the VPS too.
I run light traffic stuff at home in a closet so it doesn't occupy more costly cloud RAM. Production ready saas offerings I'm trying to sell right now are all in the cloud. Hosting all my stuff in the cloud would cost me hundreds per month. My home SLA is fine for the extras. I don't need colocation at this time, but I have spoken with data centers to understand my upgrade path.
You can run a live backup server at a second location and have pretty good redundancy should the primary lose power or connectivity.
When system requirements elevate (SLA, security, etc) you probably want to move into a data center for better physical security, reliable power and network. Bigger VPS is fine if it is big enough. Can also do a colocation if you don't want to rent, and you contract directly with a data center. I wouldn't look at colocation until your actual hosting needs exceed at least $100/mo and you're ready for a year long commitment.
> Cloud infra may be a comparable market, since computation is a big share of AI costs. Did consumers win big from competition between AWS, Azure, and GCP? Not sure. I see an uptick in write ups saying “We switched off cloud and reduced costs by 2/3rds.” Not a scientific sample but may leave the question open.
(Probably also less reliable than "cloud", but who cares if you're a startup.)
Several very big profile names have recently begun moving back to self-hosted, hybrid, or dedicated hosting solutions.
Cloud computing never was good in terms of value, however, it was only good in terms of scalability. AI solutions built on top of 'the cloud' will always be even worse.
Note that I pay a certain "third tier" cloud provider less than $100/mo total for hosting large websites that would cost me more than 10 grand a month on AWS/Azure/Google...while having better uptime. (the biggest differences? the complete lack of IO and bandwidth charges, and much lower storage charges)
That should tell you all you need about these types of bubbles, but then again, most of us that watched the entire tech field unfold since the pre-internet phase already knew this.
Could you provide some examples?
Some random survey from ESG found 60% of respondents repatriated at least some workloads. Who knows what the N was though.
Source: https://www.sdxcentral.com/articles/thinking-about-cloud-rep...
Frankly, it makes sense that companies are deciding what works best and where. But the death knell of cloud providers is simply not happening.
The prime selling point for cloud for large enterprises was (and still is):
- a signatory that shares blame on several core security issues (iso stuff) - high amount of flexibility for individual teams used to asking for a vm then waiting four weeks for the itops dept to bring it online
Now the vast majority of cloud moves for large enterprises ends up as a shitshow due to poor implementation sure, but the key points for getting it sold are still there.
And ofc you CAN still cost optimize with cloud, its just harder.
Context: Worked in post and pre sales over some years in MSFT in the enterprise segment.
I got out before the downturn and everyone talking about cost, but my approach in selling azure to the c-suite would be fairly similar today I reckon.
If the cloud had not existed, those that claim they saved money switching away from cloud might never have been in business in the first place.
(If you can train your internally deployed LLM on data none of your competitors have, that's an advantage).
This could be one of the more interesting privacy fights of the next decade.
I’m sure there are easy cynical takes about how they will just shrink wrap the EULA, and maybe they will. But in a good privacy environment, users should never be surprised and have control over how their data is used. And I think we’ve made some progress there.
If there's one company that I don't think cares about user permissions or the law, it'd be Twitter.
The EU officially warned Elon about DSA fines and the response was less than serious.
https://www.cnn.com/2023/10/10/tech/x-europe-israel-misinfor...
I will happily take a 1:1 bet on them not going the way of MySpace in, say, 5 years!
I'm willing to bet dollars to doughnuts that Google and Facebook have at least one, possibly 2 or more orders of magnitude more latent training data to work with - not including Googles search index.
My uninformed opinion is that Google and Meta's ML efforts are fragmented - with lots of serious effort going into increasing existing revenue streams with LLMs and the like being treated as a hobby or R&D projects. OpenAI is putting all its effort into a handful of projects that go into a product they sell. The dynamic and headcounts will change if the LLM market grows into billions
It seems more likely that at Google at least they just fell into the classic innovator's dilemma in which they were stuck trying to apply innovation to their current business models in an attempt at incremental innovation instead of seeking an entirely different customer and market.
Both have a large graveyard of failed innovation attempts. Remember Google Buzz? Google Plus?
They're not as good at innovating, that's why they acquire startups all the time. It's a blood transfusion
If OpenAI survives all the legal challenges, they'll just click "Go" and be in business in weeks to months.
If OpenAI gets smacked down, they haven't lost much.
There are probably also some submarine operations that are already doing/have done the training. If OpenAI gets bankrupted for copyright violations, we'll just never hear from those.
Not sure any startup would be ok with sitting on a potentially gamechanging model.
Google/Facebook? Maybe?
Amazon's LLM efforts are pretty much in shambles
For what reason do you doubt that?
It's not clear there's any more of a long-term market for this any more than there is for compilers. I'm kind of scratching my head trying to figure out where this assumption there is one comes from.
That is a beautiful metaphor, seriously.
Yet...what is the biggest U.S. ride share app?
Hungry Hippos is true. But it also works.
Apple did good with the M processors though
I'm not sure, there are 100+ yo companies like Kodak, Nikon, IBM, Panasonic, GSK, Merck, etc. With that in mind, it's not hard for me to imagine that some of the tech giants will also have a presence in hundred years, maybe not as dominant as today, but they could survive.
Coca-Cola, Nestle, and Unilever. Easy bet. Coca-Cola is already 137.
The point of surge pricing is to rebalance supply and demand when demand for rides outstrips supply of drivers, whether by attracting more drivers or by discouraging price sensitive riders.
Some riders at the airport will strongly prefer Uber, for whatever reason, so they're less price sensitive than you. Because you're happy to substitute an Uber for a taxi, you decrease the demand as a response to the surge pricing, preserving the limited supply of drivers for the riders who really want them (or are at least are price insensitive enough to pay for that privilege).
I just took Lyft again to the airport earlier this month same location and I was billed $49 USD, and a $1.30 “Texas Surcharge”.
An inflation calculator says that $38 usd in 2015 is equivalent to $49 in 2023. Color me surprised. I thought the prices had significantly increased since I signed up but it looks like actually no they didn’t.
Trawling back through those old emails I do see constant “50% off all weekday rides” offers from the time I signed up until about March 2016, at which point they stopped. So there were some subsidized incentives when they were early in Austin but it looks like they stopped sometime in early 2016. So if the money train existed, it happened before that, at least in Austin.
I like Uber; it’s convenient and fairly reliable. Five milliseconds after Lyft creates a better experience, I can switch.
Some will scream in horror but I wanted Netflix to be a monopoly. A single place and app and account with all the content I need.
"competition" in streaming space has been nothing but disastrous for me as a consumer. It led to greedy heterogeneous islands of content, with proliferation of crappy apps and pointless restrictions and return to cable package mentality.
Again, Possibly irrationally and ignorantly, my fear is that 5 years from now I'll need a dozen subscriptions to less good services which will hoard their source data and models and be specialized based on which content they got licenses to. I. E. There'll be ai1 with new York times and Wikipedia, and ai2 with Washington post and encyclopedia Britannica, and ai3 with I don't know fox news and RT, and ai4 with mit and Harvard business libraries, and ai5 focused on math with extra subscription to wolfram, and ai6 with rights to stack overflow and JavaScript and so on.
There are many scenarios various writers have posited where we are actually in local maxima lf ll, with data being increasingly closed and or poisoned, and possibly segregated in the near future. :-/
We're not socially ready for one AI winner. The resulting giant would be too powerful and too influential.
If you run it you get some money and higher balances mean better treatment when it takes over.
Distributed evil AI evaluation in exchange for protection from it. You can trade the protection on the free market.
If other vendors LLMs become good enough it will actually be easily to interchange and then the race for the best UX and integration will be upon us (which the other commenter alluded to).
"if other llms are good enough" assumes that in principle they have same opportunities, access to same data, or content. My fear is precisely that this assumption may be taken away - I. E. That news paper publishers or encyclopedia owners or big websites (stack overflow, web Md, etc) will enter into arrangement with specific llm companies - just like Netflix Disney prime etc aren't competing on their app or price or flexibility, but on exclusive underlying content. Nobody WANTS to subscribe to Paramount+ or cbs access... But if they hold enough material hostage some people will 'have to'. I can see a future, not far off, where different llm organizations selling feature is not how good their technology is - to your and overvodys point, THAT moat is likely to even out - but what underlying training data they have legal access to.
I am convinced that's how they will end up.
If anything is going to be based on how good the tech is, it's this.
I hate this. Not the best will win, but the one with the biggest pockets. Nothing that helps with technofeudalism. Proper competition would be good.
I've talked to a good amount of businesses and 90% of custom use cases would also have negligible AI costs. In my opinion, unless you're in a super regulated industry or doing genuinely cutting edge stuff, you should probably just be using the best that's available (OpenAI).
Microsoft is a behemoth at dealing with enterprises and has been doing it for decades. Even old school enterprises are OK with uploading data to Azure.
I don’t understand how people are surprised by this anymore.
So yeah, it’s the best option right now, when the company is burning through cash, but they’re planning on getting that money back from you eventually.
Genuine question, what are some examples of companies in that "hiking the price" camp?
I can think of tons of tech companies that sold or sell stuff at a loss for growth, but struggling to find examples where the companies then are able to turn dominant market share into higher prices.
To be clear, I'm definitely not implying they are not out there, just looking for examples.
One thing I'll add is that it's not always that this ends with higher prices in an absolute sense, but that the tech company is able to essentially cut the knees out of their competitors until they're a shell of their former selves. Then when the prices go "up", they're in a way a return to the "norm", only they have a larger and dominant market share because of their crazy pricing in the early stages.
But yeah the movie thing is not at all unrealistic here. Though at night I usually wave the torch on my phone to attract attention because they don't always see a raised hand.
At the busiest time it's a bit harder but at that time the ride-sharing services are also overloaded so it's still faster to just wait for a green light (free taxi). We don't get Uber as far as I know but we do have a similar thing called Cabify. But it's useless if you need something quick and they've put the prices up too much. I now only use them for scheduled stuff like airport dropoffs.
There's no way I want to ever wait on hold for a dispatcher again, or be mystified as to when my cab is arriving (this always seems to involve standing outside in the snow or pouring rain).
If the cab companies have apps comparable to Uber and Lyft, sure, I'll give them a shot.
- Google is showing more and more ads over time to power high revenue growth YoY
- Unity has just tried to increase its prices
Uber/Lyft did raise prices, but interestingly (at least to me) is that if the strategy was the smother the competition with low prices, it didn't seem to work.
Unity is interesting too, though I'm not sure it would make a good poster child for this playbook. It raised prices but seems to be suffering for it.
As for Unity, they're certainly dealing with a bunch of underperforming PE and IPO-enabled M&A on the one hand (really should have considered that AppLovin offer, folks), but also just a failure to extract reasonable income from their flagship product on the other; I don't think their problems come from raising prices per se (game devs pay for a lot already, an engine fee is nothing new to them) as much as how they chose to do it and the original pricing model they tried to force on their clients. What they chose to do and the way they handled it wasn't just bad, it was "HBS case study bad."
At some point, someone else will make a competitive model, if it’s Facebook then it might even be open source, and the industry will see price competition downwards.
I don't know how they could _not_ incorporate customer usage to improve their models.
There is no such thing as an open source Google because Google’s value is in its vast data centers. Search is hard to train and hard to run.
GPT4 is not that big. It’s about 220B parameters, if you believe geohot, or perhaps more if you don’t.
One hard drive.
Whereas the underlying algorithms behind all these GPTs so far are broadly same. Yes, OpenAI does probably have better data, model finetuning and other engineering techniques now, but I don't feel it's anything special that'll allow themselves to differentiate themselves from competitors in the long run.
(If the data collected from a current LLM user in improving model proves very valuable, that's different. I personally think that's not the case now but who knows).
rephrasing this for LLMs instead of search: "you can create your own model architecture/training method, but you can't crawl the web and serve language query results to billions of worldwide users in a few milliseconds."
that checks out, right? Google/search == """Open"""AI/LLMs still seems like a decent metaphor to me.
Also OpenAI gets significant discount on compute due to favourable deals from Nvidia and Microsoft. And they could design their server better for their homogenous needs. They are already working on AI chip.
People will figure out what OpenAI is doing and duplicate it. There’s many people working at OpenAI, it’s going to leak out.
e.g. As they have a fixed model which they know they would get billions of request to, they could even work with analogue chip which is significantly cheaper and faster for inference. [1] could achieve 10-100x flops/watt for fixed models compared to nvidia for their first gen chip.
That is the moat. For developer platforms, it's all about building mindshare and adoption. The more people who know how to use OpenAI, the stronger OpenAI's position on the market. It doesn't matter if there's equivalent or slightly better models unless they start to fall significantly behind (and they're currently well in the lead).
However, what will be interesting is if the price of delivering ChatGPT-style experiences drops as the industry matures/advances, and their pricing moat erodes.
Unlike Uber - where prices are dictated by factors unlikely to move significantly (Labour / Vehicles / Fuel / etc), The LLM space doesn't have these types of overheads.
Sure, if you want to let a monopoly have all the added value while you get to keep the rest you can do that.
Just make sure you're never successful enough to inspire them though, otherwise you're dead the next minute. Oops.
I suspect that OpenAI doesn't have the bandwidth to build most uses of ai and so is in the bill gates platform land: the ecosystem should be pocketing more money than the owner of the platform is.
It's possible gpt next or next++ ends up making whatever work you do trivial, but it's likely you'll still have customers
Apple may be leading the way here, with Apple Silicon prioritizing AI processing and built into all their devices. These capabilities are free (or at least don't require an extra sub), and just used to sell more hardware.
OpenAI is clearly going to compete in that market with its upcoming smart phone or device [1]. But what revenue model can OpenAI use to compete with Apple's and not get undercut by it? I suppose hardware + free GPT3.5, and optional subscription to GPT4 (or whatever their highest end version is). Maybe that will be competitive.
I also wonder what mobile OS OpenAI will choose. Probably not Android, otherwise they would have partnered with Google. A revamped and updated Microsoft mobile OS maybe, given their MS partnership? Or something new and bespoke? I could imagine Johnny Ive demanding something new, purpose-built, and designed from scratch for a new AI-oriented UI/UX paradigm.
A market for increasingly sophisticated AI that can only be done in huge GPU datacenters will exist, and that's probably where the margins will be for a long time. I think that's what OpenAI, Microsoft, Google, and the others will be increasingly competing for.
[1]:https://www.reuters.com/technology/openai-jony-ive-talks-rai...
So eventually you could be running decent sized models locally (iOS could even provide an API with fine tuning etc)
Would completely change how I use the device.
Siri: "Sorry. Dictation service is unavailable at the moment."
It's past time for excuses. High-level people at Apple need to be fired over this. Hello? Tim? Do your job. Hello? Anybody home...?
My hope is that the upcoming eu rulings allow competition here. Ie force Apple to get out of the way of making their hardware better with better software.
You know what would get Apple to fix this? Forced competition. You know what Apple spends their trillions preventing?
It's not a time for complacency, if only because driver assistance is becoming more important every day. There are good, sound business reasons to put competent people on the Siri team.
It’ll probably take a while though.
What phone are you referring to? A quick google didn’t seem to pull up anything related to OpenAI launching a hardware product?
https://www.yahoo.com/entertainment/openai-jony-ive-talks-ra...
Excuse me, I'm not an english native, you mean like a smart phone? Or do you mean some sort of other new business direction? Where did you get the info thtat they're planning to launch a phone?
https://www.nytimes.com/2023/09/28/technology/openai-apple-s...
In my experience apple's ML on iphones is seamless. Tap and hold on your dog in a picture and it'll cut out the background, your photos are all sorted automatically including by person (and I think by pet).
OCR is seamless - you just select text in images as if it was real text.
I totally understand these aren't comparable to LLMs - rumor has it apple is working on an llm - if their execution is anything like their current ML execution it'll be glorious.
(Siri objectively sucks although I'm not sure it's fair to compare siri to an LLM as AFAIK siri does not do text prediction but is instead a traditional "manually crafted workflow" type of thing that just uses S2T to navigate)
Wasn't that solved about a decade ago. Does anyone suck at that?
Does android even have native OCR? Last I checked everything required an OCR app of varying quality (including windows/linux).
On ios/macos you can literally just click on a picture and select the text in it as if it wasn't a picture. I know for sure on iOS you don't even open an app to do it, just any picture you can select it.
Last I checked the Opensource OCR tools were decent but behind the closed source stuff as well.
Random google result of OCR on android (could be outdated) - https://www.reddit.com/r/androidapps/comments/10te5et/why_oc...
Tesseract? https://github.com/tesseract-ocr/tesseract
To use, go to the app switcher and just start selecting text.
It's funny trying to select text in a video while the video is playing...
AFAIK this is a feature only available on Google Pixel devices. so not on 'Android' in general.
For one, my Oppo device running Android 12 doesn't have it.
Yes, Android fragmentation sucks.
That at least works for Pixel phones. Not sure about Samsung and others.
I'm not sure this is true, even in the short term. For some things yes, that's definitely true. But for other things that are real-time or near real-time where network latency would be unacceptable, we're already there. For example, Google's Pixel 8 launch includes real-time audio processing/enhancing which is made possible by their new Tensor chip.
I'm no fan of Apple, but I think they're on the right path with local AI. It may even be possible that the tendency of other device makers to put AI in the cloud might give Apple a much better user experience, unless Google can start thinking local-first which kind of goes against their grain.
Agreed. Something else I wonder is if local AI in mobile devices might be better able to learn from its real-time interactions with the physical world than datacenter-based AI.
It's walking around in the world with a human with all its various sensors recording in real-time (unless disabled) - mic, camera, GPS/location, LiDAR, barometer, gyro, accelerometer, proximity, ambient light, etc. Then the human uses it to interact with the world too in various ways.
All that data can of course be quickly sent to a datacenter too, and integrated into the core system there, so maybe not. But I'm curious about this difference and wonder what advantages local AI might eventually confer.
GPS, phone orientation, last 5 apps you were in, etc. --> embedding
you might even have like "what time is it?" compressed as it's own embedding.
They will keep pricing the off-the-shelf AI at-cost to keep competitors at bay.
As for competitors, Anthropic is the most similar to OpenAI both in capabilities and business model. I am not sure what Google is up to, since historically their focus has been in using AI to enhance their products rather than making it a product. The "dark horses" here are Stability and Mistral which both are OSS and European and will try to make that their edge as they give the models for _free_ but to institutional clients that are more sensitive to the models being used and where is the data being handled.
Amazon and Apple are probably catching up. Apple likely thinks that all of this just makes their own hardware more attractive. It's not clear to me what Meta's end goal is.
Let me introduce you to the VC business model. Get comical amounts of money. Charge peanuts for an initial product. Build a moat once you trap enough businesses inside it. Jack up prices.
For example, efforts like OpenMOE https://github.com/XueFuzhao/OpenMoE or similar will probably eventually lead to very competitive performance and cost-effectiveness for open source models. At least in terms of competing with GPT-3.5 for many applications.
Also see https://laion.ai/
I also believe that within say 1-3 years there will be a different type of training approach that does not require such large datasets or manual human feedback.
I guess if we ignore pretraining, don't sample-efficient fine-tuning on carefully curated instruction datasets sort of achieve this? LIMA and OpenOrca show some really promising results to date.
This makes a lot of sense. A small model that “knows” enough English and a couple of programming languages should be enough for it to replace something like copilot, or use plug-ins or do RAG on a substantially larger dataset
The issue right now is that to get a model that can do those things, the current algorithms still need massive amounts of data, way more than what the final user needs
OpenAI will murder my solution by quality, by availability, by reliability and by scalability...all for the price of a coffee.
It's a personal project though & partly intended for learning purposes so there is scope for accepting trainwreck level tradeoffs.
No idea how commercial projects are justifying this though.
Sometimes this can be unacceptable. Law,, medicine, finance, all of them would prefer a self-hosted, private GPT.
[0] - https://platform.openai.com/docs/models/how-we-use-your-data
There are some cases where you really can't afford to send Microsoft data for their OpenAI offering... but there are a lot more where some figurehead solidified their power by insisting the company build less secure versions of public offerings instead of letting their "gold" go to a 3rd party provider.
As AI starts to appear as a competitive advantage, and the SOTA of self-hosted lagging so ridiculously far behind, you're seeing that work less and less. Take Harvey.ai for example: it's a frankly non-functional product and still manages to spook top law firms with tech policies that have been entrenched for decades into paying money despite being OpenAI based on the simple chance they might get outcompeted otherwise.
But you just can't. You cannot trust the scrappy startup OpenAI. You can't even trust Microsoft's normal cloud offering, because the people who actually give a fuck about the risk NEED to have granular detail of what data, readable by whom, is stored exactly where and for how long, and how can you make sure, and how do you know that access is scoped to the absolute minimum number of people, and is there a paper trail for that?
For these "figureheads": the buck, stopping, here, etc.
> You cannot trust the scrappy startup OpenAI
Not saying you do: Azure has a dedicated capacity driven GPT-4/3.5 offering that you can stick in your VPC with everything from PCI to HITRUST certs. These are the things that come out if you actually care about delivering solutions vs jumping to deliver the right sounding words for the figureheads like "We'd never trust those scrappy OpenAI guys!!!!"
> Quite the opposite -- every single tech person would love to make our data someone else's problem (and get a big career boost from dealing with cloud tech instead of the dead-end that is local sysadmin!).
You're attracting the least equipped people who tumbled into what you just admitted is a dead end trajectory, usually paying below market rates as a result, and then expecting them to outperform the people paying the most money for competent security outlays with much bigger fish (Azure is working with teams that need FedRAMP, DoD certs, HIPPA compliance, and much more)
The end result is that you end up with a poorly maintained leak sieve of an infrastructure in which Azure would likely be the most secure component you have to lean on in your entire organization.
You say:
> because the people who actually give a fuck about the risk NEED to have granular detail of what data, readable by whom, is stored exactly where and for how long, and how can you make sure, and how do you know that access is scoped to the absolute minimum number of people, and is there a paper trail for that
They don't care about risk, they care about flawed perceptions of risk that don't align with reality. These are the same companies that get pwned for years through some basic social engineering, and all that they ever have to show for it is audit logs that show who ac... ah wait no one ever actually checked the logs and it turns out they're useless because subsystem X Y and Z aren't even connected to it.
This nonsense that you can't trust anyone with your data is completely unfounded
You realize that Microsoft is a publicly traded company that has multiple privacy certifications? They have to subject to detailed data ownership and consumption audits. They most definitely have data ownership and retention logs and you can request for copies of their certification audits to understand how they log/track this. I think the parent is being too optimistic but your answer is so comically simplistic it's silly. I highly suggest you read about the world of HIPAA, PCI, and FedRAMP instead of just thinking "omg the data".
It's the lawyers that you need to convince. Good luck convincing any bigco lawyer that your company's data is safe on openAI because their legal agreement says "we don't train on API calls."
It's “not be used to train or improve OpenAI models”, doesn't mean it's not used to get knowledge about your prompts, your business use case. In fact, the wording of the policy is lose enough they could train a policy model on it (just not the LLM itself).
So its hard to trust.
For legal issues it's a bit more nuanced (eg new york state has guidelines about best practices, but they're honestly fairly sensible and would probably allow SOC 2 or equivalent)
The best thing anybody can do with your data is not store it for very long. Beyond that, they should take sensible measures, like encrypt it at rest, have policies restricting access, etc
...but you'd struggle to get close to even GPT 3.5 let alone 4 for generic tasks.
For custom tunes...yeah sure custom rolls will beat generic openAI. But that's a bit like pitting customed tuned cars against street legal manufacturer cars. It's an apple to oranges comparison
But tasks are only generic in the aggregate, and with local models you can hand-off between different models for different tasks.
The general rule of thumb is, take a model size (7B, 13B, 34B, 70B) and multiply that by 0.5 or 0.625. If that number is smaller than the combined amount of system RAM and VRAM in your system, you can run the model at 4-bit and 5-bit quantization respectively.
* Specific purposes/verticals.
* Security/privacy/hosting requirements.
Really depends what your API input is, using full GPT 4 context will drive you bankrupt with ~a dollar per prompt.
Or is it the deal with Microsoft for cloud services making it cheap?
Or are they just operating at a massive loss to kill off other competition?
Or something else?
1) They hiring too talent to make their models as efficient as possible.
2) They have a sweetheart deal with MS.
3) They’re better funded than everyone else and bringing in substantial revenue.
49% isn't _just_ a share, it's a significant portion of the company.
They might also be operating at a loss afaik, but I suspect they're one of the few that can break even just based on scale, brand recognition, and economics.
I’ve seen 100 to 200 million active users, but nothing about paid users from them. The surveys I saw when doing a quick google search reported much less than 1% of users paying.
https://seekingalpha.com/news/4007459-microsoft-backed-opena...
I’m also very curious how much of that money is VC money coming from startups trying to build their customer base, vs how much actually sustainable.
Bingo.
> the deal with Microsoft for cloud services making it cheap?
It should make it cheaper, but it takes time and engineers to migrate work load from AWS (which is reasonably adept at scaling) to azure, which is not.
I think they run k8s in something like 4k groups of nodes, which is spectacular, because k8s isn't really designed to do that. Running it at that scale is challenging because the traffic required to coordinate is massive (well it was last time I looked into it.)
They've got a landfill of cash.
7 months later, nothing's changed surprisingly. Even open-source models are trickier to get to be more cost-effective despite the many inference optimizations since. Anthropic Claude is closer to price and quality effectiveness now, but there's no reason to switch.
Either there will be some major technological breakthrough that lowers their costs, or they will all eventually start raising prices.
It was roughly $150 for me to build a small dataset with a few thousand quarter-page chunks of text for a data project using GPT4. GPT3 is substantially cheaper but it would hallucinate 30% of the time; honestly a nice fine-tune of LlaMA is on-par with GPT3 and after the sunk cost all it costs is a few $0.01 in electricity to generate the same sized dataset.
You can rent a 3090 at $0.20/h on vast.ai, or $0.40/h on runpod. Using VLLM at 400t/s that's 1440000 generated tokens. Generating that amount of tokens with GPT3.5 would be $2.88.
"some" is doing a lot of heavy lifting here.
Also: don't discount the labor cost of curating a fine tuning dataset, running a FT training run, even if the hardware is cheap.
LLama models only really shine for things that GPTs would refuse to even consider because of corporate RLHF, and if you need to keep your data local I suppose. For the rest they're second rate at best.
I use both OpenAI and Anthropic (I use my own Common Lisp and Racket Scheme client libraries that I implement with similar APIs) and I was amazed last night how well self hosted LLama-based and Mistral LLMs run on a 32G Mac Mini that was delivered to me late yesterday afternoon. This is not expensive hardware. And tools like llama.cpp will keep getting more efficient, etc.
We are going to see unimaginable (at least to me) advances in AI and AGI, at all levels of the tech food chain. And these advances will occur quickly. Place your bets, and remain flexible!
That scares me. I hate moats and actively want out. Running the uncensored 70B parameter Llama 2 model on my MacBook is great, but it's just not a competitive enough general intelligence to entirely substitute for GPT-4 yet. I think our community will get there, but the surrounding water is deepening, and I'm nervous...
this is the thing that scare me.
when do these models stop getting smarter? or at least slow down?
what do you mean? is the API uncensored?
The cost of the GPT-4 API is ballpark around $0.05 / 1000 tokens. If you want to include a rolling context window which you basically HAVE TO DO if you want to maintain a persistent conversation, you will easily meet or exceed 1000+ tokens.
ChatGPT Pro gives you 50 GPT-4 queries every three hours. If you're using it all day you might average about 100 daily queries. Using a dedicated GPT4 API would run you approximately five dollars a day for the same thing - that's $150 a month as opposed to flat cost of $20.
We are in a different category.
I'm forming a new company. I was asked for a single sentence to describe the company as well as a longer paragraph. I wrote the single sentence and then fed it into chat.openai.com ChatGPT 3.5 (before I paid) and asked for a longer version.
What it came up with was a bit too heavy on the adjectives (the tone was a bit too much marketing), but wow, it nailed it in terms general concepts about what I'm building. This was all without knowing anything about the business. I can easily edit it back down to an easier tone for people to digest.
When I fed the same input into ChatGPT4, the results were 100x better.
I also like to use it to summarize my thoughts. I can write down a bunch of unfiltered gibberish about how I'm feeling today and what I'm thinking. Feed it in and it'll give me a great summary of what I just said in a non-threatening and neutral tone. Having gone through a lot of professional therapy (and even being married to one in a past life), it feels a lot like that to me... except a lot less expensive.
OpenAI can't "win" (whatever that means) long term unless they figure out how to collect user data inputs (text based questions) and reward vectors (Thumbs up and down) persistently, at scale, in the extreme long term - which means building something people rely on all day everyday.
As far as I can tell they have no unique distribution avenues to do this today outside of copilot.
Meanwhile, Apple, Amazon and Alphabet are certainly bringing GPT capabilities to Siri/Alexa/Whatever Google's voice thing is called, albeit slower and more carefully, but they have no need to rush at all here.
I'll bet Microsoft will slowly absorb OpenAI given their investment position and integrate it into Bing or something and fade away.
Spending the capex/opex to run a cluster of compute isn't easy or cheap. It isn't just the cost of the GPU, but the cost of everything else around it that isn't just monetary.
Make adoption easy, give a free base tier but charge more could be a very effective model to get start ups stuck on you. It even probably makes adoption by small teams in big companies possible that can then grow ...
Answer these questions, and the equation shifts a bunch!
A full rack with 16 amps usable power and some bandwidth is $400/month in Kansas City, MO. That is enough to power 5x A100s 24x7, so 10k plus $80 per month each, amortized, of course many more A100s would drop the price.
Once installed in the rack ($250 1 time cost) you shouldn't need to touch it. So 10k plus $1250 per A100, per year including power. You can put 2 or 3 A100s per cheapo Celeron based CPU with motherboards.
Of course if doing very bursty work then it may well make sense to rent...
You also left out the data tech costs- probably at least $50K/individual-year in KC (although I guess I'd just work for free ribs).
If you're putting A100s into celeron motherboards... I don't know what to say. You're not saving money by putting a ferrari engine in a prius.
You hire people who are employed by the DC, by the hour for DC work. How much work is there to do once it is screwed into the rack?
The problem though is that getting 2-3MW of power in the US is increasingly difficult and you're going to pay a lot more for it since the cheap stuff is already taken.
Even more distressing is that if you're going to build new data center space, you can't get the rest of the stuff in the supply chain... backup gennies, transformers, cooling towers, etc...
Also that is a 8xA100 system as others have noted, but it is the 40GB one which can be found on eBay for as low as $3k if you go with the SXM4 one (although the price of supporting components may vary) or $5k for the PCI-e version.
You can build a lot of stuff on top of these two.
It's absolutely worth the money when you look at the whole picture. Also lambda labs never has availability. I actually can schedule a distributed cluster on AWS.
That highly depends on many things. If you run a business with a relatively steady load that doesn't need to scale quickly multiple times per day, AWS is definitely not for you. Take Let's Encrypt[1] as an example. Just because cloud is the hype doesn't mean it's always worth it.
Edit: Or a personal experience: I had a customer that insisted on building their website on AWS. They weren't expecting high traffic loads and didn't need high availability, so I suggested to just use a VPS for $50 a month. They wanted to go the AWS route. Now their website is super scalable with all the cool buzzwords and it costs them $400 a month to run. Great! And in addition, the whole setup is way more complex to maintain since it's built on AWS instead of just a simple website with a database and some cache.
S3 and EC2 are priced very competitively.
Your pricing only matters if you have availability.
To train a top model you need hundreds of them in a very advanced datacenter.
You can't just plug gpus into standard systems and train, everything is custom.
The technical talent required for these systems is rare to say the least. The technical talent to make a model is also rare.
I trained a few foundation models with images, and I would NEVER buy any of them. These guys are on a wildly different scale than basically everyone.
And since there are many different industries/specializations with a lot of nuanced, undocumented knowledge which is not available online, it will be difficult for a single large company to acquire all that specialist information. I think the bottleneck isn't going to be hardware costs, but merely putting together the optimal training data. To do this, you need to find the top experts in the world in any given field.
Unfortunately, it's difficult to do right now because top experts are often not given credit these days. Those who are promoted as the top people in any given field are often mostly good at politics and lack the deep nuanced knowledge that would be required to produce top quality training data.
The pricing is too good to be true with you think about it rationally. If they raise prices they seem much, much less attractive than using AWS or Azure.
Amazon seem to have a much better business built around their Bedrock offering. And all their other tools are available there like SageMaker, ec2, integration with MLFlow, etc, etc.
I guess the same goes for Azure, if you are already using it it's much easier to just stick with whatever they are offering for LLM Ops.
OpenAI offering just models doesn't seem like it can last forever, and to compete with AWS or Azure at enterprise level they need to build all the things Amazon/MS have built.
The other side of that coin seems much more realistic.
In what way shape or form?
> If they raise prices they seem much, much less attractive than using AWS or Azure.
They're already significantly more expensive than Azure. OpenAI charges something like $30k a month for dedicated capacity on a "call our sales team" basis: GPT 3.5/ GPT-4 on Azure comes with that for free.
And GPT-4 is already slow and expensive enough that no one just chooses it arbitrarily... they're using it for things no other model can do. They could charge double for GPT-4 and GPT-4 would still be the only model that can do those tasks: you wouldn't get to just switch off to some other GPT-4 equivalent provider.
> Amazon seem to have a much better business built around their Bedrock offering
Amazon is literally doing the same thing with Bedrock! They're offering Anthropic at competitive prices to OpenAI for a model that's no cheaper to run based on their own dedicated capacity numbers.
> OpenAI offering just models doesn't seem like it can last forever, and to compete with AWS or Azure at enterprise level they need to build all the things Amazon/MS have built.
OpenAI is not trying to become Azure: They actively go out of their way to hide the fact they even offer half the things they offer to enterprises, instead relying on Azure absorbing demand as much as possible.
OpenAI wants ChatGPT Plus to be the new Prime, as in no one should be able to afford to not pay OpenAI for their immensely valuable offering.
Except unlike Prime, the offering is software, not commerce: If Amazon could get AWS-like margins from their e-commerce business, AWS would be a footnote.
But there are a ton of use-cases where a 1 to 7B parameter fine-tuned model will be faster, cheaper and easier to deploy than a prompted or fine-tuned GPT-3.5-sized model.
In fact, it might be a strong statement but I'd argue that most current use-cases for (non-fine-tuned) GPT-3.5 fit in that bucket.
(Disclaimer: currently building https://openpipe.ai; making it trivial for product engineers to replace OpenAI prompts with their own fine-tuned models.)
Also, it's important to note that fine tuning produces a vast amount of data about use cases where fine tuning is useful.
[1]: That knowledge cutoff and terrible UX of browse the web is brutal compared to the experience of Bard
If they can pull off what Apple has done with its M line, then possibly they can make themselves even more cost effective than the competition. The competition will mostly be limited to supplies from Nvidia.
I believe in house manufacturing of their own hardware is definitely the way forward. Top down lock down. Own the hardware, own the models, own the user trained datasets.
Not quite accurate; finetuned 3.5 is only 4x cheaper than GPT-4. Cost per million output tokens from https://openai.com/pricing
$ 2 - GPT-3.5
$16 - finetuned GPT-3.5
$60 - GPT 4Renting GPU servers at AWS is just stupid. Did you ever see bitcoin miners calculate the price of mining one bitcoin by looking at renting AWS GPU servers?
1. LLM that talks like ChatGPT 2. Image generator that makes realistic portraits from verbal descriptions.
Are the costs in the data acquisition, human training input, training CPU/GPU hours, hardware, or ??
I think a comparison to Netflix and the media industry is fair. Years ago, people claimed Netflix's tech and infrastructure was their moat, but it turns out it was the cheap easy access to loads of high quality content that was the real motivator.
While it's nice to consume the cheap stuff, it is not good for healthy markets.
huggingface will be more of a winner in the long run with an exit via an MSFT acquisition if I had to call it.
OpenAI gets cheaper by twiddling their thumbs for the next few years, meanwhile they continue to amass more and more data for RLHF.
It's weird that people are trying to drag non-software scaling into a software scaling problem: Lyft, Doordash, Instacart, etc. all relied on VC dollars to scale non-software growth like software. OpenAI really is just stupidly cheap compared to anything those past high CAC plays were aiming to do.
RLHF I looked it up. Is this really useful? The average human has zero general expertise because people are specialized (I know nothing about say, 1960s avant garde french cinema and my responses in a conversation there would be garbage - given the breadth of human knowledge even the most accomplished scholars are useless for over 99% of it). Won't there be a quality decrease? How is this accommodated for?
If the chat systems simply gave the most popular answers it would cease to be useful real fast.
Before RLHF instruct tuning the models could only complete sentences
Technically they still complete sentences, but now they have a strong association for a format where a question is followed by an answer
Maybe I'm missing something crucial here, but why does it dripfeed answers like this? Does it have to think really hard about the meaning of 1e100? Why can't it just spit it out instantly without such a delay/drip, like with the near-instant Wolfram Alpha?
https://ai.stackexchange.com/questions/38923/why-does-chatgp...
Future products would be able to hide some of that, but for now, that’s what the ChatGPT / Bing Assistant product does.
And it is the right business strategy for them. It is MSFT money, it is free money.
The one thing openAI has that's hard to compete with is a metric shit ton of money and brand recognition.
It feels like the biggest investor bait of this year
Will it beat ARM IPO?