The problem has nothing to do with commercializing image gen AI and all to do with Emad/Stability having seemingly 0 sensible business plans.
Seriously this seemed to be the plan:
Step 1: Release SD for free
Step 2: ???
Step 3: Profit
The vast majority of users couldn't be bothered to take the steps necessary to get it running locally so I don't even think the open sourcing philosophy would have been a serious hurdle to wider commercial adoption.
In my opinion, a paid, easy to use, robust UI around Stability's models should have been the number one priority and they waited far too long to even begin.
There's been a lot of amazing augmentations to the stable diffusion models (ControlNet, Dreambooth etc) that have propped up, lots of free research and implementations because the research community has latched onto the stability models and I feel they failed to capitalize on any of it.
It’s a shame because they’re literally just using stable diffusion for all their tech but built a nicer front end and incorporated control net. No-where else has done this.
Controlnet / instantID etc are the really killer things about SD and make it way more powerful than Midjourney, but they aren’t even available via the stability API. They just don’t seem to care.
But for the individuals involved, it might also be
Step 2: Leverage fame in AI space for massive VC injection on favorable terms.
> Step 1: Release SD for free
> Step 2: ???
> Step 3: Profit
That’s not true. He was pretty open about the business plan. The plan was to have open foundational models and provide services to governments and corporations that wanted custom models trained on private data, tailored to their specific jurisdictions and problem domains.
You know, the old tried and true licensed merchandise model. Everybody gets paid.
"Cool with AI" and "sell my likeness so nobody ever needs to hire me again" are too close for comfort on this one.
I'm betting the list of folks who would sign the AI license are pretty small, and mostly irrelevant.
It's just not there yet. GenAI outputs aren't something audiences wants to hang on a wall. It's something that evoke sense of distress. Otherwise everyone's tracing them at least.
> It's just not there yet. GenAI outputs aren't something audiences wants to hang on a wall.
People have a wide range of standards. Last summer I attended the We Are Developers event in Berlin, and there were huge posters that I could easily tell were from AI due to the eyes not matching; more recently, I've used (a better version) to convert a photo of a friend's dog into a renaissance oil painting, and it was beyond my skill to find the flaws with it… yet my friend noticed instantly.
Also, even with "real art", Der Kuss (by Klimt) is widely regarded as being good art, beautiful, romantic, etc. — yet to me, the man looks like he has a broken neck, while the woman looks like she's been decapitated at the shoulder then had her head rotated 90° and reattached via her ear.
[0] This is also why people look at a Google street view image with a ©2017 Google[1] tiled over on a blue sky and say "LOL, Google's trying to own the sky", or why people even on this very forum ask how some new company can trademark a descriptive term like "GPT"[2], seemingly surprised by this being possible even though there's already a very convenient example of e.g. Hasbro already having "Transformers".
[1] https://www.google.com/maps/@33.7319434,10.8655264,3a,77.2y,...
The point is, generative AI images are not widely regarded as good art. They're often seen as passable for some filler use cases and hard to tell apart from human generations, but not "good".
It's not not-there-yet because AI sometimes generates sixth fingers, it's something another level from Gustav Klimt, Damien Hirst, Kusama Yayoi, or the likes[0]. It could be that genAI is leaving something that human artist would filter out, or because images are too disorganized that they appear to us to be encoding malice or other negative emotions, or maybe I'm just wrong and it's all about anatomy.
But whatever the reason is, IMO, it's way too rarely considered good, gaining too few supportive celebrities and artists and audiences, to work.
0: I admit I'm not well versed with contemporary art, or art in general for that matter
> It's not not-there-yet because AI sometimes generates sixth fingers, it's something another level from Gustav Klimt
My point is: yes AI is different — it's better. (Or, less provocatively: better by my specific standards).
Always? No. But I chose Der Kuss specifically because of the high regard in which it is held, and yet to my eye it messes with anatomy as badly as if he had put 6 fingers on one of the hands (indeed, my first impression when I look closely at the hand of the man behind the head of the woman, is that the fingers art too long and thumb looks like a finger).
wait what? Isn't that missing the point of expressionism? Klimt's Judith I is basically a photo, surely he can draw sh*t if he wanted to?
But myriad predecessors such as Vermeer, Rembrandt, Van Gogh, da Vinci, et al., have done enough in realism, and also photography was becoming more viable and more prevalent, that artists basically started diversifying? Isn't that what lead to various forms of early 20th century arts like surrealism(super-real -ism), cubism, etc?
I don't mean offense but that's just, surely that level of understanding can't be basis of policy decisions when it comes to moral rights and licensing discussions and "artists should just use AI" and such???
I am asserting here that the AI is (at its best) more competent, not any of the other things.
I suspect that the law will follow the economics, just as it often has done for everything else before — you're communicating with me via a device named after the job that the device made redundant ("computer").
But I said "often" not "always", because the business leaders ignoring the workers they were displacing 200 years ago led to riots, and eventually to the Communist Manifesto. I wouldn't discount this repeating.
--
I've just looked up "Judith I" (I recognise the art, just not the name), and I don't even understand why you're holding this up as an example of "basically a photo".
As for the other artists demonstrating realism: photography made realism redundant despite being initially dismissed as "not real art". Artists were forced to diversify, because a small box of chemistry was allowing unskilled people do their old job faster, cheaper, and better. Photography only became an art in its own right when people found ways to make it hard, for example by travelling the world and using it to document their travels, or with increasingly complex motion pictures.
I suspect that art fulfils the same role in humans as tails fulfil in peacocks: an expensive signal to demonstrate power, such that the difficulty is the entire point and anything which makes it easy is seen as worse than not even trying. This is also why forgeries are a big deal, instead of being "that's a nice picture", and why an original painting can retain a high price despite (or perhaps because of) a large number of extremely cheap prints being plastered onto everything from dorm rooms to chocolate wrappers.
Automatic1111, ComfyUI, Oobabooga. There's more value within these 3 projects than within at least 1 billion dollars worth of money thrown around on yet another podunk VC backed firm with no product.
It appears that no one is even trying to seriously compete with them on the two primary things that they excel at - 1. Developer/prosumer focus and 2. extension ecosystem.
Also, if you're a VC/Angel reading my comments about this, I would very much love to talk to you.
The dream
AI startups need not an insignificant amount of startup capital , you cannot just spend weekends to build like you would a saas app . Model training is expensive so only wealthy individuals can even consider this route
Companies like that have no oversight or control mechanisms when management inevitably goes down crazy paths, also without external valuations option vesting structures are hard to ascertain value.
20% investors 70% founders 2-3% employees (1% emp1, 1% emp2, 0.5% emp3, 0.25% emp4) 7% for future employees before next funding round
I was a heavy user since the beginning but my usage has dropped to almost 0
That's why I asked that question to see if others notice something similar or if that's just me
The more I think about the AI space the more I realize that open sourcing large models is pointless now.
Until you can reasonably buy a rig to run the model there is simply no point in doing this. It's no like you will be edified by setting the weights either.
I think an ethical business model for these business is to release whatever model can fit into a $10,000 machine and keeping the rest closed source until above machine is able to run them.
Also, things like this are in the works:
https://news.ycombinator.com/item?id=39794864
Which will put the system RAM of the new 24-channel PC servers in range of the Nvidia H100 on memory bandwidth, while using commodity DDR5.
The medium sized models like GPT3 and Grok are 185b and 314b respectively.
There is no way for _anyone_ to run these on a sub $50k machine in 2024, and even if you can the token generation speed on CPU is under 0.1 tokens per second.
2 x 32GB: $142
2 x 64GB: $318
8GB: $16
2 x 16GB: $64
2TB of 128GB DDR4 ECC: $9,600 (https://www.amazon.com/NEMIX-RAM-Registered-Compatible-Mothe...)
> Servers that support that much are actually cheap (~$200)
What does this mean? What motherboards support 2TB of RAM at $200? Most of them are pushing $1,000. With no CPU.
It may not hit $50K, but it's definitely not going to be $2K.
https://www.ebay.com/itm/176298520843
Here are 128GB LRDIMMs for $98:
https://www.ebay.com/itm/196305803969
For 2TB and the server you're at $1698. You can get a drive bracket for a few bucks and a 2TB SSD for $100 and have almost $200 left over to put faster CPUs in it if you want to.
That's stinking Optane, would work if you're desperate. Normal 128GB LRDIMMs cost more than other DDR4 DIMMs. You can, however, get DDR4 RDIMMs for ~$1/GB:
https://www.ebay.com/itm/186345903230
With 32GB RDIMMs that machine would max out at 768GB, which could still run a 1T model at q4 or grok at FP16. And then it would cost less than $1000.
Or find a quad-socket system with 48 memory slots and then use 64GB LRDIMMs ($1.12/GB):
https://www.ebay.com/itm/176299295509
The quad socket systems aren't $200, but you can find them for $550 or so:
https://www.newegg.com/hp-proliant-rack-mount/p/2NS-0006-3E5...
Maybe less if you shop around (they're not as common).
I believe there is some research on how to distribute large models across multiple GPUs, which could make the cost less lumpy.
And "depending on the task" is the point. There are systems that would be uselessly slow for real-time interaction but if your concern is to have it process confidential data you don't want to upload to a third party you can just let it run and come back whenever it finishes. And releasing the model allows people to do the latter even if machines necessary to do the former are still prohibitively expensive.
Also, hardware gets cheaper over time and it's useful to have the model out there so it's well-optimized and stable by the time fast hardware becomes affordable instead of waiting for the hardware and only then getting to work on the code.
The withdrawn paper: https://arxiv.org/abs/2310.17680
The wrong source: https://www.forbes.com/sites/forbestechcouncil/2023/02/17/is...
The discussion: https://www.reddit.com/r/LocalLLaMA/comments/17jrj82/new_mic...
As models advance, they will become - not just larger - but also more efficient. Hardware advances. Large models will run just fine on affordable hardware in just a few years.
And while AIs may become more compute-efficient in some respects, the tasks we ask AIs to do will grow larger and more complex.
Sure you might get a good image locally but what about when the market moves to video? Sure chat GPT might give good responses locally, but how long will it take when you want it to refactor an entire codebase?
Not saying that local compute won’t have its use-cases though… and this is just a prediction that may turn out to be spectacularly wrong!
You get all the benefits of academics and open source folks pushing your model forward, and a vastly improved hiring pool.
But it doesn't stop you launching a commercial offering, because 99.99% of the world's population doesn't have 48GB+ of VRAM.
I see it this way to be honest:
- companies will aggresively try to use AI in the next 2-3 years, downsizing themselves in the meantime
- the 3-5 year launch mark will show that downsizing was an awful idea and took too many hits to really be worth it. I don't know if those hits will be in profits (depends on the company) but it will clearly hit that uncanny valley.
- 6-8 year mark will have studios hiring like crazy to get the talent they bled back. They won't be as big as before, but it will grow to a more sane level of operation.
- 10-12 year mark will have the "Apple" of AI finally nail the happy medium between efficiency and profitability (hopefully without devastating the workers, but who knows?). Competitors will follow throw and properly usher the promises AI is making right now.
- 15 year mark is when AI has proper pipelining, training, college courses, legal lines, etc. established and becomes standard faire, no stranger than using an IDE.
As I see it, companies and AI tech alike are trying to pretend to be the 10 year mark all the while we're currently in legal talks and figuring out what and where to use AI to begin with. In my biased opinion, I hope there's enough red tape on generative art to make it not worth it for large studios to leverage it easily (e.g. generative art loses all copyright/trademarkability, even if using owned IPs. Likely not that extreme, but close).
They will not downsize, they will train their workforce or hire replacements that are willing to pick up these more powerful and efficient tools. In the hands of a skilled professional there will be no uncanny valley.
This will result in surplus funds, that can be invested in more talent, which in turn will keep feeding AI development. The only way is up.
Not allowing copyright on AI generated work is a ridiculous and untenable decision that will be overturned eventually.
Sure, the smart companies will use it as a tool, but most companies aren't smart, or just don't care. It'll vary by industry. There is already talks of sizing down VFX/Animation for a mix of outsourcing and AI reliance, for example. And industry that already underpays its artists.
>Not allowing copyright on AI generated work is a ridiculous and untenable decision that will be overturned eventually.
Maybe, once the dust settles on who and what and how you copyright AI. It'll be a while, though. But I get the logic. No one can (nor wants to) succinctly explain what sources were used in a generative art work right now, and that generative process drives the art a lot more than the artist for most generative art. Even without AI there is a line between "I lightly edited this existing work on photosshop" and "I significantly altered a base template to the point where you can't recognize the template anymore" where copyright will kick in.
Still, my biased hopes involve them being very strict with this line. You can't just give 2 prompts and expect to "own" an artwork.
I disagree completely. But I should note I was referring to medium to large scale companies. Nothing in those companies happens in "months" these days.
Maybe some startups rise from being first to market much faster, but given the huge legal issues I'm not seeing it. Microsoft et al. can afford A lengthy legal battle and make backup plans. A startup can't.
And that becomes part of the problem because it's hard to sell unreliable technology unless you design the product in a way that plays well with the current shortcomings. We will get there, but it's still a few iterations away.
They're the new WYSIWYG/low-code. Everyone that doesn't fully understand the problem space thinks they're some ultimate solution that is going to revolutionise everything. People that do are responding with a resounding 'meh'.
Stable Diffusion is a great example. Something that can generate consistent game assets would be an absolute game changer for the entire game industry and open up a new wave of high tech indie game development, but despite every "oh wow" demo hitting the front page of HN, we've had the tech for a couple of years now and the only thing that's come out of it is some janky half solutions (3D meshes from pictures that are unworkable in real games, still no way to generate assets in consistent styles without a huge amount of complex tinkering) and a bunch of fucking hentai lol.
I look at Copilot and it’s been the same for me. I’m either working on a huge codebase and most of the time, it means tweaking and refactoring, which is not something I trust a LLM with. Or it’s a greenfield project and I usually write only the necessary code for a task and boilerplate generation is not a thing for me. Coding for me is like sculpting and LLM-based solutions feel like trying to do with bricks attached to my feet. You can get something working if you’re patient enough, but it’s make more sense and it’s more enjoyable to just use your fingers.
Whatever Kinbaku is... haha
I think a key part that's missing currently is the agent training approach: https://youtu.be/v3UBlEJDXR0?si=8w4Jt0bNEBfIXkZl
It's very easy to create a "Stochastic Parrot" but I'm quite sure these models are capable of learning underlying information such as correct layout of a knot - given the right data and curriculum of course. Maybe slight architecture tweaks.
I'm sure this is the reason we're starting to see a normal amount of fingers or ability to write text. Proof of concept was 2015 until 2022 now we're starting to see interesting things come out of the workshops.
Now things are pretty mature, but it took decades to get there but there is still a whole bunch of hacks upon hacks behind the scenes. Same story will repeat with each new problem domain.
I don't think the limiting factor here is the software; it looks like we got AI-generated art pretty much as soon as consumer graphics cards could handle it (10 years ago it would have been quite hard). I'd be measuring progress in hardware generations not years and from that perspective Stable Diffusion is young.
Obviously, current AIs cannot generate game rulesets because the game feel is an internal phenomenon that cannot be represented in the material domain and therefore AIs cannot train on it.
Creating the art, on the other hand, goes from a week per card to half a day per card, or something similar.
GPT-4 cost $100M+ to train (Altman), but Dario Amodei has said next-gen models may cost $1B to train, and $10B models are not inconceivable.
I'd guess OpenAI's payroll is probably $0.5B (770 highly paid employees + benefits, not to mention hundreds of contractors creating data sets).
They're doing what they should: growing the customer base while continuing to work on the next generation of the core technology, and developing the support code to apply what they have to as broad a cross-section of problems as they have the potential to offer a solution for.
Didn't stop them from being extremely successful.
The point is, OpenAI can afford to have free-loaders as long as their deals from enterprise, governments are paying for the service.
Midjourney doesn't have a free plan so no free-loaders there and they're making $200M+ with no VCs.
Stability.ai will always suffer from free-loaders due to their fully open source AI.
Stability's recent models (SD3, SV3D, StableLM2, StableCode, and more) are neither open licensed nor planned for release as open licensed.
The reason is that Stability.ai gave away everything for free. Until recently, they didn't even attempt to charge money for their models.
I've heard the only reason they're not already closed up is that they're reselling all of the rented GPU quota they leased out years ago. Companies are subletting from Stability, which locked in lots of long term GPU processing capacity.
There's no business plan here. This is the Movie Pass of AI.
They have raised 110M in October and they say that training a particular model costs them hundreds of thousands $ in compute costs.
They don't lack money.