I'm sure that in the not too distant future (a few years at most) we will be happily running these on customer level hardware.
I do wander if companies working to develop these type of revenue models truly think it's a long term structure?
I'm sure that in the not too distant future (a few years at most) we will be happily running these on customer level hardware.
I do wander if companies working to develop these type of revenue models truly think it's a long term structure?
Whether Adobe ever decides to let their model run locally or lock it forever into the cloud is a choice they will have to make. A lot of people trust Adobe products, so it's entirely conceivable that some people will always choose to pay for a pay-per-use generative solution from Adobe rather than try to run competing solutions locally. The question is probably whether it generates more revenue than negativity for Adobe. If most Adobe users are running their own models locally and avoiding the feature, then I think Adobe will be more likely to follow suit and move away from the pay-per-use cloud approach.
I would have said the consumer sentiment amongst Adobe users is the exact opposite - that people don't trust Adobe products but they use them because they either have to or because they're currently the best products available.
Software companies already do that. There are all kinds of locally-run advanced features that are only enabled with a more expensive subscription tier even though you already have the code and assets for them.
Sure, most are merely subscription based, but there are others that are per use.
To pay for the privilege of using a very advanced AI model. That's more reasonable than paying to unlock a game character skin that's already on your SSD, and that happens millions of times a day.
> It'd be like charging per use of any other advanced features in locally-run Adobe software.
I don't see anything stopping this.
Enclaves are rarely broken and the people that can are selling it to the CIA, not leaking it on the pirate bay.
They seem like pretty much the perfect fit for cloud - burst compute which would result in very low hardware utilisation if ran locally.
Why would it be better to have a $1,500 GPU that is weak and used infrequently, when you could share a big cluster of better GPUs shared between a big group of people, and have it more heavily utilised?
There is a philosophical argument about owning your own hardware etc, however I think the economics and performance will eventually push this to the cloud for most use-cases (most people will just get better bang-for-buck in the cloud).
This is referring to the Photoshop stuff, which is way better than any type of SD inpainting for removing things from images. Firefly might be slower? I haven’t used it since it first came out.
I agree for just general image generation SD or Midjourney are better options in their own way.
xformers, 1024x1024x diffuser pipeline.
But someone needs to make this possible and maintain such a solution which would cost also money.either you pay adobe what you already do or pay someone else who maintains the model, the infrastructure etc.
Sure some will run it themselves but my guess is that this is a niche group of people as most don't care .
And designing your software to a minimum-spec of a 3080 would be pretty wild.
Energy, partial hardware cost, setup time, fine-tuning time.
…until I tried the same on my RTX 4070 and it made my Mac look like a joke.
For the 30 seconds my Mac would have taken for 1 result, which will probably need revising, the RTX would give me 30 results.
However the RTX was half the cost of my Mac, so it’s not a good investment if I just want to generate some images. I’d rather pay for the cloud if I didn’t have the RTX already.
The actual best image AI, midjourney, is probably a gigantic model under the hood, that takes 8 A100s to run (Aka more than 100GB VRAM). That's why their quality is leaps and bounds above stable diffusion XL, its because the model size simply allows for it.
Model sizes continuously grow to exploit the available hardware to the limit. Midjourney and GPT-4 have both proven that model quality is decisive to success and paying customers, so consumer hardware can never catchup to whatever Nvidia sells to the cloud.
Unet are really expensive to run compare to a regular GPT model and they are compute-bound thx to convolution, a reason why no one has trained a unet that comes close to consume 100GB of VRAM during inference. I doubt that MJ is much bigger than SD XL, it's good but not revolutionary.
An optimized XL model in the hands of an expert beats it, handily.
And Adobe sells to experts, not consumers for the most part.
Do you have any source for this speculation? In my experience image models are always much smaller than language and even the largest llama will fit in a smaller GPU machine than that.
For me it kicks the shit out of Midjourney in flexibility and quality. I can make more images of higher quality, faster and cheaper.
Whether the amounts they pay would make licensing your work sensible or not, Adobe is surely assuming this will ultimately end up as Napster-to-Spotify transition.
If we end up "happily" (means legally as well) running these on customer level hardware, then the question won't be about credits of computation. It'll be about credits to use licensed work.
If this is true (which I kinda doubt), is it going to matter to most people? Like you can't really tell the images used to train a model from the images it generates (if it's trained correctly), so I doubt the majority of people would care, like those who already use MJ for example. Training models on copyrighted data for academic research will be allowed, the models will be published, and good luck enforcing the licence; and here I'm talking about the worst case scenario where a court would find an image generated by AI to be derivative of another image in a pool of billions in a dataset (this goes way beyond any definition of derivative work for now).
However, Stable diffusion already can run on mobile devices. There is already a good iOS app for it (and the dev is here on HN) but the problem seems to be that no one cares. There are 700,000 cloud imagegen apps crowding it out, because thats what's easier and more profitable to spam across the store and web.
Back in the 'Google Daydream' days, Google might have found that they didn't get any more image-generation performance by raising the parameter count - but that's just because the technology at the time couldn't effectively utilise more parameters. It's impossible to know what next-gen models might be able to use, but I suspect we will find ways to allow the models to take advantage of even higher parameter counts.
Stable diffusion can run on mobile devices, but it's painful and image generation takes a fraction of the time via cloud services.
For image quality, sure - language understanding is still an issue. SDXL can generate a beautiful image, but if it doesn’t show exactly what you asked for in the prompt, on the first try, there is still room for improvement. The gap between LLMs and image generators in this regard is huge.
The models themselves will be hoarded as IP. Doesn't matter if they're in the cloud or on devices, they'll be licensed like commercial proprietary software with the same restrictions commercial software has.
Or alternately someone could make a major advance in distributed training and we could all contribute cycles in a distributed effort like Folding@Home. As it stands training requires far too much bandwidth for synchronization and moving model data around. Some approach to sharding training would have to be discovered. It’s an open problem area.
Neural networks are very parallelizable and training is stochastic so my intuition is that it should be possible. Even if it were less efficient than synchronous training you could make up for that by harnessing 100X the compute from a huge crowd.
It's one thing to train on Common Crawl in 2023, but what about when you have to shell out millions of dollars just for access to data sets to train on in the future? Same thing with human reinforcement. The customers for both are willing to pay much more than a crowdfunding campaign would.
Training is expensive now, but data sets can be expensive in the future.
If the dataset contains text from 2022, things that happen in 2023 and later won't be in it. The model will only get you so far, and new data, events, concepts, discoveries, etc will be absent.
If we trained models on all text generated up until 1900, for example, you could get it to produce some impressive results if just generating text is the goal. If the goal is to build something that imitates a more general AI, it wouldn't "know" about antibiotic treatments for common illnesses, modern vaccines, either World War, powered flight, transistors, computers, etc. It would only be useful for so much.
I mean I guess my electricity provider gets paid per compute.
I doubt that is what Adobe will do. This is a new revenue stream for them, why would they remove it?
Gimp will use local generation but Adobe is using a proprietary dataset that they can keep secure in the cloud.
So yeah this is going to be sticking around.
This is the stage of AI that will impress me most. If I can use your AI completely offline on my device on a spaceship orbiting Pluto, then I will say we have achieved an AI capacity that is impressive, even if its got the quirks of chatgpt today.
> If I can use your AI completely offline on my device on a spaceship orbiting Pluto
And the answer was yes. I do not know the exact system requirement of the current ChatGPT, but I am fairly confident
1) ChatGPT no network mode could fit in a half of a server rack, maybe way less.
2) You can fit half a server work on a spaceship that can go to Pluto.
My guess is it's more like 1 server worth. Google tells me GPT3 was 1TB which is a very small laptop.
Sidenote, ChatGPT uses tens of thousands of GPUs to run its architecture, I think it'll take a little more than just some laptop.
Thank you for answering though :) I do think its an invaluable goal to have off-line first AI.
For example - The Tesla self driving AI takes many hundreds (thousands?) of computers to build the model. Then it runs in realtime of 1 "GPU" that lives in my car. It's not sending frames in realtime to a supercomputer to process.
So for a spaceship - same thing. You don't need to send the thing that makes the model. You just send a finished model and a GPU to run it.
people are looking at an extremely limited view of “bigger models on better hardware will always be in the cloud” when that reality simply won’t matter for most use cases
Those models will affect us more than today already and change how we perceive AI.
Than we will start to see AI optimized hardware (much more optimized).
And than perhaps in 10 years we all run a lot more models locally.
Nonetheless or despite this, the normal consumer doesn't run open models and will probably not do that for a very long time. Searching, keeping up-to-date and running models is still effort and the usage model makes a ton of sense. Escpecially in time of SaaS.
Im not running wikipedia locally. And none of my social circle operates infrastructure / server.
People just want to use it.
Besides that, whatever local models or open models will be able to do, AIaaS will have faster models, better models and more convinient models.
I'm just waiting to pay for google assistent if it becomes smart and can manage my emails my calendar and everything else. After all my gmail account already has access (through email and password reset) to most services i use.
I'm more curiuos when we will see AI service integration through much more system to system communication. Machine friendly apis (which partially already exist anyway)
PS: Look at how fast hardware development currently is. Not much change in Memory etc. Models will not just become 100x smaller in just a few years. We are right now at optimizing those models to be cost efficient. Alone this phase will take a few years.
I also run LLMs such as trains of llama2, though LLMs on commodity hardware are not as “there” yet as image generators. It’s a decent question and answer bot and summarizer but isn’t GPT-4 level. I could see another iteration approaching that but I’d probably need more RAM.
Plenty of us already are. SDXL is as good as anything in the cloud.
Unless nVidia changes their monetization model, and for example introduces an App Store for AI, with subscriptions, of course on locked down hardware.
On the contrary. In a few years there wont be a lot of customer level hardware software (especially business software) without a subscription.
I'm sure the workloads will shift locally more and more, if for no better reasons than latency and privacy.
Are the technical requirements driving these monetization schemes or is it the other way around?