Emad Mostaque resigned as CEO of Stability AI
stability.ai
stability.ai
I’m a huge fan of the tech, but as reality sets in things are gonna get quite rough and there will need to be a painful culling of the AI space before sustainable and long term value materializes.
-
This is most likely the reason being the SAMA firing, to be able to re-align to the MIC without terrible consequence, from a PR perspective.
no criticism here aside from the fact that we will see the AI killer fully autonomous robots will be here and unfettered by 'alignments' much sooner than we expected...
And the MIC is where all the unscrupulous monies without audits will come from.
What exactly do you mean with this sentence? That less woke/regulated companies will suddenly leapfrog the giants now? What timeframe are we talking here?
And do you mean for example non public mil applications from US/China or whatever or private models from unknown players?
One thing i've been wondering is that if GPT-4 was 100 mil to train, then there's really a lot of plutocrats, despots, private companies and states for that matter that could in principle 10x that amount if they really wanted to go all in, and maybe they are right now?
The bottleneck is the talent pool out there though, but i'm sure there's a lot people out there from libertarians to nation states that don't care about alignment at all, which is potentially pretty crazy / exciting / worrying depending on view.
What kind of "leapfrog" do you think is necessary to produce a "killer fully autonomous robot"?
We've actually had "autonomous killer robots," machines that kill based on the outcome of some sensor plus processing, for centuries, and fairly sophisticated ones have been feasible for decades. (For example it's trivial to build a "killer robot" that triggers off of face or voice recognition.)
The only thing that's changed recently is the kind of decisions and how effectively the robot can go looking for someone to kill.
This video is from 7 years ago: https://www.youtube.com/watch?v=ecClODh4zYk
The interesting thing is who's actively working on this, any despots, private companies, foreign nations, who from the talent pool, criminal orgs or even western military which mostly works for the western elite classes.
Aren’t they? Pretty sure that most tactical and strategical decisions are automated to the bottom. Drones and cameras with all sorts of CV and ML features. You don’t need scary walking talking androids with red eyes to control battlefields and streets. The idea of “Terminator” is similar to a mailman on an antigrav bicycle.
Kinda. In my experience, the bigger issue is the skillset has largely been diffused.
Overtraining on internal corpora has been more than enough to enable automation benefits, and the ecosystem around ML is very robust now - 10 years ago SDKs like Scikit-learn or PyTorch were much less robust than they are now. Implementing commercial grade SVM or <insert_model_here> is fairly straightforward now.
ML models have largely been commodified, and for most usecases, the process of implementing models fairly straightforward internally.
IMO, the real value will be on the infrastructural side of ML - how to simply and enhance deployment, how to manage API security, how to manage multiple concurrent deployments, how to maximize performance, etc.
And I have put my money where my mouth is for this thesis, as it is one that has been validated by every peer of mine as well.
In my experience, there really hasn't been that significant of a flood of money in this space for several years now, or at least not to the level I've seen based on discussion here on HN.
I think HN tends to skew towards conversations around models for some reason, but almost all my peers are either funding or working on either tooling or ML driven applications since 2021.
-------
I've found HN to have a horrible noise to value signal nowadays, and people with experience (eg. My friends in the YC community) deviating towards Bookface or in person meetups instead now.
There was a flood of new accounts in the 2020-22 period (my hunch is it's LessWrong, SSC, and Reddit driven based on the inside jokes and posts I've been seeing recently on HN) but they don't reflect the actual reality of the industry.
Deep down I fundamentally believe it's a solution searching for a problem. That said, if people want to put money into it, who am I to judge.
That said, IMO crypto as a growth story is largely done. Now that Coinbase has IPOed and the industry consolidated or wiped out (eg. Kraken, FTX, OpenSea), it's not as attractive an industry anymore unless some actual fundamental problems are found that can't be remediated by existing financial infrastructure (legal and illegal).
This is plain hystronics, but regulation is absolutely coming and will only help the larger existing players to consolidate.
Coinbase, FTX, etc have all seen the writing on the wall and are working with regulators on this.
The Wild West days are definetly over now.
If I'd invest in the space, I'd probably look at KYC and AML product opportunties such as Chainalysis, but even that space has consolidated
The magic is finding which subsegments might have even higher margins than others.
Crypto used to have fairly high margins, but all the easy gains have been claimed by larger firms as it's a much more mature industry now, but portions of the ML space still have much more opportunity for growth, so it makes sense to deploy capital there (this is a very high level view so take with a grain of salt - B2C and SMB B2B and Enterprise B2B have entirely different GTM motions and path to profitability).
But at least for me, I can't justify participating in a Series A round for a crypto startup compared to an MLOps startup today.
Not wanting to be part of that is just having some morals. Grabbing money where you can isn't a sign of a successful life.
There are definitely still homeopathic treatments that work, but have not received the investment necessary to go through clinical trials (or haven’t caught on due to lack of recognition from insurance companies).
Story of my life.
I have a severe HN addiction that I haven't been able to shake off.
Not just training but inference too, right? They have to make money off each query.
A challenge at the moment is a lot of the AI movement is led by folks that are brilliant technologists but have little to no experience running viable businesses or building a viable business plan. That was clearly part of why OpenAI has its turmoil in that some where trying to be tech purists where others knew the whole AI space will implode if it’s not a viable business. To some degree that seems to behind a lot of the recent chaos inside these larger AI companies.
This is saying nothing about "technologists" (or as they're starting to become derided as "wordcels": people that communicate well, but cannot execute anything themselves).
It would be... not trivial, but straightforward to map out the finances on everything involved, and see if there is room from any standpoint (engineering, financial, product, etc.) to get queries to breakeven, or even profitable.
But at that point, I believe the answer will be "no, it's not possible at the moment." So it becomes a game of stalling, and burning more money until R&D finds something new that may change that answer (a big if).
if the Good Lord's willing and the creek don't rise
They played their cards so damn right when deep learning was taking off.
Recently, during an interview [1], when questioned about OpenAI's Sora, Shantanu Narayen (Adobe CEO) gave an interesting perspective on where value is created. His view (paraphrased generously)..
GenAI entails 3 'layers': Data, Foundational Models and the Interface Layer.
Why Sora may not be a big threat is because Adobe operates not only at first two layers (Data and Foundational model) but also at the interface layer. Not only Adobe perhaps knows better than anyone else what is need and workflow of a moviemaker, but I guess most importantly they already have moviemakers as their customers.
So product companies like Adobe (& Microsoft, Google etc.) are in better position to monetize GenAI. Pure-play AI companies like OpenAI are perhaps in B2B business. Actually, they maybe really in api business, they would have great data, would be building great foundational models and giving results of those as APIs; which other companies who are closer to their unique set of customers with their unique needs would be able to monetize and some part of those $$ flows back to pure-play AI companies
[1] At 5 mins mark.. https://www.cnbc.com/video/2024/02/20/adobe-ceo-shantanu-nar...
I only ever heard creatives complain about Adobe and their UI/UX and how they don’t understand their customers.
Never really used any of their products myself though. Maybe they still are best-in-class. I can’t tell.
Figma could build an illustrator killer in 6 months if they wanted to and it would be obliterated.
If they actually tackled this task people would be kicking themselves for putting up with the shambles that is illustrator for this long.
Statements like this are almost always wrong, if for no other reason that a technically superior alternative is rarely compelling enough by itself. It that weren’t the case you would see it happen far more often…
The scenario I was calling out is more: Company A is great at X, so clearly that technology could be used to easily be great at Y, and doom company B in 6 months. The problem with that thinking is that often the technology is the easy part (and sometimes superficial similarity, also).
Every user hates using microsoft products, and don't get me started on SAP. But these are gigantic companies with wildly successful products aimed at enterprise customers.
Only because they've never had a chance to experience the competition.
Having worked in IBM and had to use the Lotus Office Suite I can tell you Microsoft won fair and square. And I'm not even talking about the detestable abomination that is Lotus Notes.
This one was so obscure to find because it seems to exist in a weird space of being Microsoft Office but not.
Every major change in the last 6 years has either been weird window dressing changes to welcome panels or new document panels, in all cases building sluggish jank heavy interfaces, try navigating to a folder in the premier one and weep as clicks take actual seconds to recognize.
Or just silly floating tooltips like the ones in Photoshop that also take a second to visible draw in.
All tangible tool changes exist outside the interface or you jump to a web interface in a window and back with the results being passed between in a way that makes it very obvious the developers are trying to avoid touching the core tools code.
Very clear Narayens outsourcing and not being a product guy has lead to this
There are tools no one uses, and there are tools people constantly complain about.
They’re terrible products for the most part. But they buy the competition (Substance) or they develop some half baked substitute (Adobe XD) that people might use since it comes with the subscription.
Where start-ups like Stability need to be rising to compete will have to be AI-native e.g. products re-thought of from the ground up like an AI image editor or as foundation-level AI research companies, agents or AI infrastructure companies.
There's no reason Stability can't play in both B2B and API if planned and strategized well and OpenAI can definitely pull it off with their tech and talent. But Stability has a few important differentiators from OpenAI where I believe if they launch an AI-native product in the multimodal space, they stand to differentiate significantly: - People join because they believed in Emad's vision of open source so it is their job to figure out a commercial model for open source. They can retain AI talent by ensuring a commitment to open source here. If they need to ensure their moat is retained and can commercialize, they should delay releasing model weights until a product surrounding the weights has been released first. Still open source and open weights but give them time to figure out a commercial strategy to capitalize their research. However because of this promise, they will not be able to license their technology to other companies. - Stability's strong research DNA (unsure about their engineering) is so badly fumbled by a lack of a cohesive product strategy that it leads to sub-par product releases. In agreement to the 3 'layers' argument, that's exactly Stability's greatest strength and weakness. Their focus on foundational models is incredibly strong and has come at the cost of the interface layer (and ultimately the data layer as it has a flywheel effect).
The company currently screams a need for effective leadership that can add on interface and data layers to their product strategy so they can build a strong moat outside of a strong research team which has shown it can disappear at any moment...
Obviously this is the polite way to send him off given the latest news about his leadership, but this rationale doesn't track.
Instead of wasting all the compute on bitcoin we pretrain fully open models which can run on people's hardware. A 120b ternary model is the most interesting thing in the world. No one can train one now because you need a billion dollar super computer.
https://www.cnbc.com/amp/2024/01/18/mark-zuckerberg-indicate...
Model training is unlike that. It's large state that is constantly updated and updates require full, up to date state.
This means you cannot distribute it efficiently over slow network with many smaller workers.
That's why NVIDIA is providing scalable clusters with specialized connectivity so they have ultra low latency and massive throughput.
Even in those setups it takes ie. a month to train base model.
Converted to distributed setup this same task would take billions of years - ie. it's not feasible.
There aren't any known ways of contributing computation without access to the full state. This would require completely different architecture, not only "different than transformers" but "different than gradient descent", which would be basically creating new branch in machine learning and starting from zero.
Safe bet is on "ain't going to happen" - better to focus on current state of art and keep advancing it until it builds itself and anything else we can dream of to reach this "mission fucking accomplished".
You gradient descend on your state.
Each step needs to work on up to date state otherwise you're computing gradient descend from state that doesn't exist anymore and your computed gradient descent delta is nonsensical if applied to the most recent state (it was calculated on old one, direction that your computation calculated is now wrong).
You also can't calculate it without having access to the whole state. You have to do full forward and backward pass and mutate weights.
There aren't any ways of slicing and distributing that make sense in terms of efficiency.
The reason is that too much data at too high frequency needs to be mutated and then made readable.
That's also the reason why nvidia is focusing so much on hyper efficient interconnects - because that's the bottleneck.
Computation itself is way ahead of in/out data transfer. Data transfer is the main problem and going in the direction of architecture that dramatically reduces it by several orders of magnitude is just not the way to go.
If somebody solves this problem it'll mean they solved much more interesting problem – because it'll mean you can locally uptrain model and inject this knowledge into bigger one arbitrarily.
There is nothing stopping you from distributing assembly/machine code for CPU instructions, yet nobody does it because it doesn't make sense from performance perspective.
Or amazon driving truck from one depo to other to unload one package at a time to "distribute" unloading because "distributing = faster".
"Never underestimate the bandwidth of a station wagon full of tapes hurtling down the highway."
It's only a question of scale and not of any principles.
But given that they don't seem to have worked on it since, I guess it wasn't too successful. But maybe there is a way
But if we think of mixture of experts models outperforming "monolithic" models, why not? Maybe instead of 8 you can do 1000 and that is easy to paralellize. It sounds worth exploring to me.
Base model was still trained in usual, non distributed way (by far the most cost).
Fine tunes were also trained in usual, non distributed way.
Proposed approach tries out several combinations to pick one that seems to perform better (where combination means ie. adhoc per layer operation).
Merging is not distributed as well.
There is not much distribution happening overall beyond the fact that fine tunes were trained independently.
Taking weight averages, weighted weight averages, trimming low diffs, doing arithmetic (subtracting base model from fine tune) etc. are all ad hoc trials throwing something on the wall and seeing what sticks the most. None of those work well.
For distributed training to work we'd have to have better algebra around this multidimentional/multilayer/multiconnectivity state. We don't have it and it has many problems, ie. evaluation is way too expensive. But solving "no need to rerun through whole training/benchmark corpus to see if my tiny change is better or not" problem will mean we solved problem of extracting essence of intelligence. If we do that, then hyper-efficient data centers will still keep beating out any distributed approach and it's all largely irrelevant because that's pure AGI already.
In the case of the paper, they are using OPT-6.7b as the seed LM which requires 8xV100 GPUs for fine-tuning each expert. That's a combined total of 256GB of VRAM for a single expert while the 3090 only has 24GB of VRAM and is still one of the most expensive GPUs out there.
Maybe we could use something like PEFT or QLoRA in combination with this technique to make each expert small enough for the community to fine-tune and make a worse Mixtral 8x7b, but I don't know enough to say for sure.
Or maybe it turns out we can make a good MoE model with thousands of smaller experts. Experts small enough for a separate member of the community to independently fine-tune on a normal GPU, but idk.
To have both a performant and distributed LLM trained from scratch, we still need a completely different architecture to do it, but this work is pretty cool and may mean that if nothing else, there is something the community can do to help move things forward.
Also, I was going to say the MoE routing on this technique was lacking, but I found a more recent paper[0] by Meta which fixes this with a final fine-tuning stage.
During MoE training you still need access to all weights.
1k experts would mean 30 TB of state to juggle with on 7B params. Training and inference is infeasible at this size.
If you'd want to keep the size while increasing number of experts, you'd end up with 7b -> 56m model. What kind of computation can you do on 56m model? Remember that expert model in MoE runs the whole inference without consulting or otherwise reusing any information from other experts. Thin network at the top just routes it to one of experts. But at this small size those are not "experts" anymore, it'd be more like Mixture of Idiots.
To put it in other way, MoE is optimization technique with low scaling ceiling that is more local maximum solution that global one (this idea works against you quickly if you want to go more that direction).
You can host a local horde and do "Seti at home" for stable diffusion.
so if these days someone is talking about decentralized anything i'd bet it involves coinshit again
Read his Wikipedia page and tell me he doesn’t sound like your run of the mill crypto scammer.
> He claims that he holds B.A. and M.A. degrees in mathematics and computer science from the University of Oxford.[7][8] However, according to him, he did not attend his graduation ceremony to receive his degrees, and therefore, he does not technically possess a BA or an MA.[7]
At the start of 2021, according to their website, it was a 55 year old bank. By the end of 2021, it was a 70 year old bank!
The bank's website is a WordPress site. And their customers must be unhappy - online banking hasn't worked for nearly two years at this point.
Anyway, their Deputy CEO gave this hilarious interview from his gaming rig. A 33 year old Deputy CEO, who by his LinkedIn claimed to have graduated HEC Lausanne in Switzerland with a Master of Science at the age of 15... celebrating his graduation by immediately being named Professor of Finance at a university in Lebanon. While dividing his spare time between running hedge funds in Switzerland and uhh... Jacksonville, FL.
The name of his fund? Indepedance [sic] Weath [sic] Management. Yeah, okay.
In this hilariously inept interview, he claimed that people's claims about Deltec's money movements being several times larger than all the banking in their country was due to them misunderstanding the country's two banking licenses, the names of which he "couldn't remember right now" (the Deputy CEO of a bank who can't remember the name of banking licenses), and he "wasn't sure which one they had, but we might have both".
Once the ridicule and all this started piling on, within 24 hours, he was removed from the bank's website leadership page. When people pointed out how suspicious that looked, he was -re-added-.
The bank then deleted the company's entire website and replaced it with a minimally edited WordPress site, where most of the links and buttons were non-functional and remained so for months thereafter.
I mean fuck it, if the cryptobros want to look at all that and say "seems legit to me", alright, let em.
And that becomes part of the problem because it's hard to sell unreliable technology unless you design the product in a way that plays well with the current shortcomings. We will get there, but it's still a few iterations away.
They're the new WYSIWYG/low-code. Everyone that doesn't fully understand the problem space thinks they're some ultimate solution that is going to revolutionise everything. People that do are responding with a resounding 'meh'.
Stable Diffusion is a great example. Something that can generate consistent game assets would be an absolute game changer for the entire game industry and open up a new wave of high tech indie game development, but despite every "oh wow" demo hitting the front page of HN, we've had the tech for a couple of years now and the only thing that's come out of it is some janky half solutions (3D meshes from pictures that are unworkable in real games, still no way to generate assets in consistent styles without a huge amount of complex tinkering) and a bunch of fucking hentai lol.
I don't think the limiting factor here is the software; it looks like we got AI-generated art pretty much as soon as consumer graphics cards could handle it (10 years ago it would have been quite hard). I'd be measuring progress in hardware generations not years and from that perspective Stable Diffusion is young.
Obviously, current AIs cannot generate game rulesets because the game feel is an internal phenomenon that cannot be represented in the material domain and therefore AIs cannot train on it.
Creating the art, on the other hand, goes from a week per card to half a day per card, or something similar.
Whatever Kinbaku is... haha
I think a key part that's missing currently is the agent training approach: https://youtu.be/v3UBlEJDXR0?si=8w4Jt0bNEBfIXkZl
It's very easy to create a "Stochastic Parrot" but I'm quite sure these models are capable of learning underlying information such as correct layout of a knot - given the right data and curriculum of course. Maybe slight architecture tweaks.
I'm sure this is the reason we're starting to see a normal amount of fingers or ability to write text. Proof of concept was 2015 until 2022 now we're starting to see interesting things come out of the workshops.
I look at Copilot and it’s been the same for me. I’m either working on a huge codebase and most of the time, it means tweaking and refactoring, which is not something I trust a LLM with. Or it’s a greenfield project and I usually write only the necessary code for a task and boilerplate generation is not a thing for me. Coding for me is like sculpting and LLM-based solutions feel like trying to do with bricks attached to my feet. You can get something working if you’re patient enough, but it’s make more sense and it’s more enjoyable to just use your fingers.
Now things are pretty mature, but it took decades to get there but there is still a whole bunch of hacks upon hacks behind the scenes. Same story will repeat with each new problem domain.
The reason is that Stability.ai gave away everything for free. Until recently, they didn't even attempt to charge money for their models.
I've heard the only reason they're not already closed up is that they're reselling all of the rented GPU quota they leased out years ago. Companies are subletting from Stability, which locked in lots of long term GPU processing capacity.
There's no business plan here. This is the Movie Pass of AI.
They have raised 110M in October and they say that training a particular model costs them hundreds of thousands $ in compute costs.
They don't lack money.
The point is, OpenAI can afford to have free-loaders as long as their deals from enterprise, governments are paying for the service.
Midjourney doesn't have a free plan so no free-loaders there and they're making $200M+ with no VCs.
Stability.ai will always suffer from free-loaders due to their fully open source AI.
Stability's recent models (SD3, SV3D, StableLM2, StableCode, and more) are neither open licensed nor planned for release as open licensed.
Didn't stop them from being extremely successful.
GPT-4 cost $100M+ to train (Altman), but Dario Amodei has said next-gen models may cost $1B to train, and $10B models are not inconceivable.
I'd guess OpenAI's payroll is probably $0.5B (770 highly paid employees + benefits, not to mention hundreds of contractors creating data sets).
They're doing what they should: growing the customer base while continuing to work on the next generation of the core technology, and developing the support code to apply what they have to as broad a cross-section of problems as they have the potential to offer a solution for.
The problem has nothing to do with commercializing image gen AI and all to do with Emad/Stability having seemingly 0 sensible business plans.
Seriously this seemed to be the plan:
Step 1: Release SD for free
Step 2: ???
Step 3: Profit
The vast majority of users couldn't be bothered to take the steps necessary to get it running locally so I don't even think the open sourcing philosophy would have been a serious hurdle to wider commercial adoption.
In my opinion, a paid, easy to use, robust UI around Stability's models should have been the number one priority and they waited far too long to even begin.
There's been a lot of amazing augmentations to the stable diffusion models (ControlNet, Dreambooth etc) that have propped up, lots of free research and implementations because the research community has latched onto the stability models and I feel they failed to capitalize on any of it.
The dream
AI startups need not an insignificant amount of startup capital , you cannot just spend weekends to build like you would a saas app . Model training is expensive so only wealthy individuals can even consider this route
Companies like that have no oversight or control mechanisms when management inevitably goes down crazy paths, also without external valuations option vesting structures are hard to ascertain value.
20% investors 70% founders 2-3% employees (1% emp1, 1% emp2, 0.5% emp3, 0.25% emp4) 7% for future employees before next funding round
The more I think about the AI space the more I realize that open sourcing large models is pointless now.
Until you can reasonably buy a rig to run the model there is simply no point in doing this. It's no like you will be edified by setting the weights either.
I think an ethical business model for these business is to release whatever model can fit into a $10,000 machine and keeping the rest closed source until above machine is able to run them.
Also, things like this are in the works:
https://news.ycombinator.com/item?id=39794864
Which will put the system RAM of the new 24-channel PC servers in range of the Nvidia H100 on memory bandwidth, while using commodity DDR5.
The medium sized models like GPT3 and Grok are 185b and 314b respectively.
There is no way for _anyone_ to run these on a sub $50k machine in 2024, and even if you can the token generation speed on CPU is under 0.1 tokens per second.
The withdrawn paper: https://arxiv.org/abs/2310.17680
The wrong source: https://www.forbes.com/sites/forbestechcouncil/2023/02/17/is...
The discussion: https://www.reddit.com/r/LocalLLaMA/comments/17jrj82/new_mic...
I believe there is some research on how to distribute large models across multiple GPUs, which could make the cost less lumpy.
And "depending on the task" is the point. There are systems that would be uselessly slow for real-time interaction but if your concern is to have it process confidential data you don't want to upload to a third party you can just let it run and come back whenever it finishes. And releasing the model allows people to do the latter even if machines necessary to do the former are still prohibitively expensive.
Also, hardware gets cheaper over time and it's useful to have the model out there so it's well-optimized and stable by the time fast hardware becomes affordable instead of waiting for the hardware and only then getting to work on the code.
2 x 32GB: $142
2 x 64GB: $318
8GB: $16
2 x 16GB: $64
2TB of 128GB DDR4 ECC: $9,600 (https://www.amazon.com/NEMIX-RAM-Registered-Compatible-Mothe...)
> Servers that support that much are actually cheap (~$200)
What does this mean? What motherboards support 2TB of RAM at $200? Most of them are pushing $1,000. With no CPU.
It may not hit $50K, but it's definitely not going to be $2K.
https://www.ebay.com/itm/176298520843
Here are 128GB LRDIMMs for $98:
https://www.ebay.com/itm/196305803969
For 2TB and the server you're at $1698. You can get a drive bracket for a few bucks and a 2TB SSD for $100 and have almost $200 left over to put faster CPUs in it if you want to.
That's stinking Optane, would work if you're desperate. Normal 128GB LRDIMMs cost more than other DDR4 DIMMs. You can, however, get DDR4 RDIMMs for ~$1/GB:
https://www.ebay.com/itm/186345903230
With 32GB RDIMMs that machine would max out at 768GB, which could still run a 1T model at q4 or grok at FP16. And then it would cost less than $1000.
Or find a quad-socket system with 48 memory slots and then use 64GB LRDIMMs ($1.12/GB):
https://www.ebay.com/itm/176299295509
The quad socket systems aren't $200, but you can find them for $550 or so:
https://www.newegg.com/hp-proliant-rack-mount/p/2NS-0006-3E5...
Maybe less if you shop around (they're not as common).
As models advance, they will become - not just larger - but also more efficient. Hardware advances. Large models will run just fine on affordable hardware in just a few years.
And while AIs may become more compute-efficient in some respects, the tasks we ask AIs to do will grow larger and more complex.
Sure you might get a good image locally but what about when the market moves to video? Sure chat GPT might give good responses locally, but how long will it take when you want it to refactor an entire codebase?
Not saying that local compute won’t have its use-cases though… and this is just a prediction that may turn out to be spectacularly wrong!
You get all the benefits of academics and open source folks pushing your model forward, and a vastly improved hiring pool.
But it doesn't stop you launching a commercial offering, because 99.99% of the world's population doesn't have 48GB+ of VRAM.
But for the individuals involved, it might also be
Step 2: Leverage fame in AI space for massive VC injection on favorable terms.
I was a heavy user since the beginning but my usage has dropped to almost 0
That's why I asked that question to see if others notice something similar or if that's just me
It’s a shame because they’re literally just using stable diffusion for all their tech but built a nicer front end and incorporated control net. No-where else has done this.
Controlnet / instantID etc are the really killer things about SD and make it way more powerful than Midjourney, but they aren’t even available via the stability API. They just don’t seem to care.
You know, the old tried and true licensed merchandise model. Everybody gets paid.
It's just not there yet. GenAI outputs aren't something audiences wants to hang on a wall. It's something that evoke sense of distress. Otherwise everyone's tracing them at least.
> It's just not there yet. GenAI outputs aren't something audiences wants to hang on a wall.
People have a wide range of standards. Last summer I attended the We Are Developers event in Berlin, and there were huge posters that I could easily tell were from AI due to the eyes not matching; more recently, I've used (a better version) to convert a photo of a friend's dog into a renaissance oil painting, and it was beyond my skill to find the flaws with it… yet my friend noticed instantly.
Also, even with "real art", Der Kuss (by Klimt) is widely regarded as being good art, beautiful, romantic, etc. — yet to me, the man looks like he has a broken neck, while the woman looks like she's been decapitated at the shoulder then had her head rotated 90° and reattached via her ear.
[0] This is also why people look at a Google street view image with a ©2017 Google[1] tiled over on a blue sky and say "LOL, Google's trying to own the sky", or why people even on this very forum ask how some new company can trademark a descriptive term like "GPT"[2], seemingly surprised by this being possible even though there's already a very convenient example of e.g. Hasbro already having "Transformers".
[1] https://www.google.com/maps/@33.7319434,10.8655264,3a,77.2y,...
The point is, generative AI images are not widely regarded as good art. They're often seen as passable for some filler use cases and hard to tell apart from human generations, but not "good".
It's not not-there-yet because AI sometimes generates sixth fingers, it's something another level from Gustav Klimt, Damien Hirst, Kusama Yayoi, or the likes[0]. It could be that genAI is leaving something that human artist would filter out, or because images are too disorganized that they appear to us to be encoding malice or other negative emotions, or maybe I'm just wrong and it's all about anatomy.
But whatever the reason is, IMO, it's way too rarely considered good, gaining too few supportive celebrities and artists and audiences, to work.
0: I admit I'm not well versed with contemporary art, or art in general for that matter
> It's not not-there-yet because AI sometimes generates sixth fingers, it's something another level from Gustav Klimt
My point is: yes AI is different — it's better. (Or, less provocatively: better by my specific standards).
Always? No. But I chose Der Kuss specifically because of the high regard in which it is held, and yet to my eye it messes with anatomy as badly as if he had put 6 fingers on one of the hands (indeed, my first impression when I look closely at the hand of the man behind the head of the woman, is that the fingers art too long and thumb looks like a finger).
wait what? Isn't that missing the point of expressionism? Klimt's Judith I is basically a photo, surely he can draw sh*t if he wanted to?
But myriad predecessors such as Vermeer, Rembrandt, Van Gogh, da Vinci, et al., have done enough in realism, and also photography was becoming more viable and more prevalent, that artists basically started diversifying? Isn't that what lead to various forms of early 20th century arts like surrealism(super-real -ism), cubism, etc?
I don't mean offense but that's just, surely that level of understanding can't be basis of policy decisions when it comes to moral rights and licensing discussions and "artists should just use AI" and such???
I am asserting here that the AI is (at its best) more competent, not any of the other things.
I suspect that the law will follow the economics, just as it often has done for everything else before — you're communicating with me via a device named after the job that the device made redundant ("computer").
But I said "often" not "always", because the business leaders ignoring the workers they were displacing 200 years ago led to riots, and eventually to the Communist Manifesto. I wouldn't discount this repeating.
--
I've just looked up "Judith I" (I recognise the art, just not the name), and I don't even understand why you're holding this up as an example of "basically a photo".
As for the other artists demonstrating realism: photography made realism redundant despite being initially dismissed as "not real art". Artists were forced to diversify, because a small box of chemistry was allowing unskilled people do their old job faster, cheaper, and better. Photography only became an art in its own right when people found ways to make it hard, for example by travelling the world and using it to document their travels, or with increasingly complex motion pictures.
I suspect that art fulfils the same role in humans as tails fulfil in peacocks: an expensive signal to demonstrate power, such that the difficulty is the entire point and anything which makes it easy is seen as worse than not even trying. This is also why forgeries are a big deal, instead of being "that's a nice picture", and why an original painting can retain a high price despite (or perhaps because of) a large number of extremely cheap prints being plastered onto everything from dorm rooms to chocolate wrappers.
"Cool with AI" and "sell my likeness so nobody ever needs to hire me again" are too close for comfort on this one.
I'm betting the list of folks who would sign the AI license are pretty small, and mostly irrelevant.
> Step 1: Release SD for free
> Step 2: ???
> Step 3: Profit
That’s not true. He was pretty open about the business plan. The plan was to have open foundational models and provide services to governments and corporations that wanted custom models trained on private data, tailored to their specific jurisdictions and problem domains.
Automatic1111, ComfyUI, Oobabooga. There's more value within these 3 projects than within at least 1 billion dollars worth of money thrown around on yet another podunk VC backed firm with no product.
It appears that no one is even trying to seriously compete with them on the two primary things that they excel at - 1. Developer/prosumer focus and 2. extension ecosystem.
Also, if you're a VC/Angel reading my comments about this, I would very much love to talk to you.
I see it this way to be honest:
- companies will aggresively try to use AI in the next 2-3 years, downsizing themselves in the meantime
- the 3-5 year launch mark will show that downsizing was an awful idea and took too many hits to really be worth it. I don't know if those hits will be in profits (depends on the company) but it will clearly hit that uncanny valley.
- 6-8 year mark will have studios hiring like crazy to get the talent they bled back. They won't be as big as before, but it will grow to a more sane level of operation.
- 10-12 year mark will have the "Apple" of AI finally nail the happy medium between efficiency and profitability (hopefully without devastating the workers, but who knows?). Competitors will follow throw and properly usher the promises AI is making right now.
- 15 year mark is when AI has proper pipelining, training, college courses, legal lines, etc. established and becomes standard faire, no stranger than using an IDE.
As I see it, companies and AI tech alike are trying to pretend to be the 10 year mark all the while we're currently in legal talks and figuring out what and where to use AI to begin with. In my biased opinion, I hope there's enough red tape on generative art to make it not worth it for large studios to leverage it easily (e.g. generative art loses all copyright/trademarkability, even if using owned IPs. Likely not that extreme, but close).
They will not downsize, they will train their workforce or hire replacements that are willing to pick up these more powerful and efficient tools. In the hands of a skilled professional there will be no uncanny valley.
This will result in surplus funds, that can be invested in more talent, which in turn will keep feeding AI development. The only way is up.
Not allowing copyright on AI generated work is a ridiculous and untenable decision that will be overturned eventually.
Sure, the smart companies will use it as a tool, but most companies aren't smart, or just don't care. It'll vary by industry. There is already talks of sizing down VFX/Animation for a mix of outsourcing and AI reliance, for example. And industry that already underpays its artists.
>Not allowing copyright on AI generated work is a ridiculous and untenable decision that will be overturned eventually.
Maybe, once the dust settles on who and what and how you copyright AI. It'll be a while, though. But I get the logic. No one can (nor wants to) succinctly explain what sources were used in a generative art work right now, and that generative process drives the art a lot more than the artist for most generative art. Even without AI there is a line between "I lightly edited this existing work on photosshop" and "I significantly altered a base template to the point where you can't recognize the template anymore" where copyright will kick in.
Still, my biased hopes involve them being very strict with this line. You can't just give 2 prompts and expect to "own" an artwork.
I disagree completely. But I should note I was referring to medium to large scale companies. Nothing in those companies happens in "months" these days.
Maybe some startups rise from being first to market much faster, but given the huge legal issues I'm not seeing it. Microsoft et al. can afford A lengthy legal battle and make backup plans. A startup can't.
Open-source AI is a race to zero that makes little money and Stability was facing lawsuits (especially with Getty) which are mounting into the millions and the company was already burning tens of millions.
Despite being the actual "Open AI", Stability cannot afford to sustain itself doing so.
They already pivoted away from open-licensed models.
got any links?
Getting fired and making moves to capitalize on the current crypto boom while it lasts
The drama around OpenAI is well documented, there are multiple lawsuits and an SEC investigation at least in embryo, Karpathy bounced and Ilya's harder to spot than Kate Middleton (edit: please see below edit in regards to this tasteless quip). NVIDIA is pushing the Dutch East India Company by some measures of profitability with AMD's full cooperation: George Hotz doesn't knuckle under to the man easily and he's thrown in the towel on ever getting usable drivers on "gaming"-class gear. At least now I guess the Su-Huang Thanksgiving dinners will be less awkward.
Of the now over a dozen FAANG/AI "boomerangs" I know, all of them predate COVID hiring or whatever and all get crammed down on RSU grants they accumulated over years: whether or not phone calls got made it's pretty clearly on everyone's agenda to wash all the ESOP out, neutron bomb the Peninsula, and then hire everyone back with at dramatically lower TC all while blowing EPS out quarter after quarter.
Meanwhile the FOMC is openly talking about looser labor markets via open market operations (that's direct government interference in free labor markets to suppress wages for the pro-capitalism folks, think a little about what capitalism is supposed to mean if you are ok with this), and this against the backdrop of an election between two men having trouble campaigning effectively because one is fighting off dozens of lawsuits including multiple felony charges and the other is flying back and forth between Kiev and Tel Aviv trying to manage two wars he can't seem to manage: IIRC Biden is in Ukraine right now trying to keep Zelenskyy from drone-bombing any more refineries of Urals crude because `CL` or whatever is up like 5% in the last three weeks which is really bad in an election year looking to get nothing but uglier: is anyone really arguing that some meme on /r/crypto is what's pushing e.g. BTC and not a pretty shaky-looking Fed?
Meanwhile over in crypto land, over the same period of time that AI and other marquee Valley tech has been turning into a scandal-plagued orgy of ugly headlines on a nearly daily basis, the regulators have actually been getting serious about sending bad actors to jail or leaning on them with the prospect (SBF, CZ), major ETFs and futures serviced by reputable exchanges (e.g. CME) have entered mainstream portfolios, and a new generation of exchanges (`dy/dx`, Vertex, Apex, Orderly) backed by conventional finance investments in robust bridge infrastructure (LayerZero) are now doing standard Island/ARCA-style efficient matching and then using the blockchain for what it's for: printing a consolidated, Reg NMS/NBBO/SIP-style consolidated tape.
As a freelancer I don't really have a dog in this fight, I judge projects by feasibility, compensation, and minimum ick factor. From the vantage point of my flow the AI projects are sketchier looking on average and below market bids on average contrasted to the blockchain projects, a stark reversal from even six months ago.
Edit: I just saw the news about Kate Middleton, I was unaware of this when I wrote the above which is in extremely poor taste in light of that news. My thoughts and prayers are with her and her family.
NMS: Neuroleptic malignant syndrome, No Man's Sky
NBBO: National Best Bid and Offer
SIP: Session Initiation Protocol, Systematic Investment Plan, Security Infrastructure Program
I love their models and I love how they have changed the entire open source AI ecosystem for the better, but the writing was always on the wall for them given how unprofitable they are.
I don't think much of the AI startup scene or socials groups like e/acc would have existed if it weren't for the tech that they just gave away for free.
Its interesting how Stability AI and their VC funding have done a much better job of acting effectively as a non-profit charity (because they don't have profits. lol) to speed up AI development and open source their results as compared to other companies that were supposed to have been doing that from the beginning.
They really were the true ActuallyOpenAI.
Related to this, if you are an aspiring person who wants to improve the world, tricking a bunch of VC investors to fund your tech and then giving away the results to everyone free of charge is the single best way to do it.
Good riddance.
At least the open source AI people have code that you can use, freely without restriction.
The doomers, on the other hand, don't do anything but try and fail to prevent other people from releasing useful stuff.
But, in some sense I should be thanking the doomers because I rather that people with such incompetence were the enemy as opposed to people who might have a chance of succeeding.
The US is great but the best AI researchers aren’t gonna want to live here if it becomes a hyper libertarian hellscape. They want to raise a family without their kids being exploited by technologies in ways that e/accs tell us we should just accept. It’s not sustainable.
Releasing cool open source AI tech doesn't turn the world into a libertarian hellscape.
You are taking the memes way too seriously.
Mostly people just joke around on twitter while are also building tech startups. e/acc isn't overthrowing the government.
Exactly. There are two groups of people: ones that defend Effective Accelerationism online with a straight face, and ones that take memes too seriously
It means unrestricted technological progress. Unrestricted, including from annoying things like consumer protection, environmental concerns, or even national security.
If e/acc was just about making cool open source stuff and posting memes on the Internet, you wouldn’t need a new term for it, that’s what people have been doing for the past 30 years.
Regardless of whether a couple people who are taking their own jokes too seriously truly believe that they are going to, I don't know, create magic AGI, the fact remains that the actual measurable results of all this is only:
1: funny memes
2: cool AI startups
Anything other than that is made up stuff in either your head, or the heads of people who just want to pump up the valuation of their startups.
> Marc Andreessen’s manifesto
Yes, I'm sure he says a lot of things that will convince people to invest in companies that he also invests in. It is an effective marketing tactic. I'm sure he has convinced other VCs to invest in his companies because of these marketing slogans.
But regardless of what a couple people say to hype up their startup investments, that is unrelated to the actual real world outcomes of all of this.
> you wouldn’t need a new term for it
The fact that me and you are talking about it, actually proves that yes some marketing terms both make a difference and also don't result in, I don't know, the government being overthrown and replaced by libertarian VCs or whatever nonsense that people are worried about.
E/acc has many interpretations. In the most basic sense it means “technology accelerates growth”. One should work on better technology and making it widely distributed. Instead of giving away money, one can have the biggest impact on humanity with e/acc.
we’ve been effectively accelerating for the past 200 years.
Nothing hyper libertarian there.
Yeah uh.
Really? Because stability AI caused a very large amount of good in the world.
It arguably kicked off the entire AI startup industry.
> Tricking people to obtain money? How's this not fraud?
Its not fraud because you don't have to lie to anyone. You can tell VCs exactly what you plan on doing. Which is to open source all of your code... and... uhhh... yeah that will totally make the company valuable.
There are lots of ways of making sales pitches about open source, or similar, that will absolutely pass regulatory scrutiny and are "honest", and yet still have no hope of commercial success and also provide a huge amount of value to the public. Like what stability AI did.
A company that sets out to obtain VC money and blow it all on open source software without turning a profit, is going to leave behind smoking guns. Those will turn up in discovery and make one's life rather difficult.
Sure they did. They have been open from the start that they were releasing everything open source. They have been very up front about that!
> is going to leave behind smoking guns. Those will turn up in discover
No it won't, and it didn't. The VCs all hopped on board onto a very transparent open source giveaway. Good on them! Nobody lied to anyone about their open source plans.
Their platform for easy distribution and management of models has sped up the ecosystem more so.
First the OpenAI rebellion in November, then the Inflection AI acqui-hire from Microsoft not willing to pay the over-valued $4B and deflected that to $600M instead (after making $0 revenue) and now a bit of in-stability at Stability AI with the CEO resigning after many employees leaving.
What does that say about the other AI companies out there who have raised tons of VC cash and aren't making any meaningful amount of revenue? I guess that is contributing to the collapse of this bubble with only a very few companies surviving.
Yes maybe the business model wasn’t perfect but ad revenue was already well and truly a thing by the time Google invented a better search engine. All they had to do was serve the ads and the rest is history.
Amazon nailed high revenue growth from the very beginning, just reinvesting in growth & deferring the margin story. They could have stopped at any time.
Google nailed high traffic from the beginning, so ad sales was always a safe Plan B. The founders hoped to find something more aesthetic to them, failed, and the conservative path worked.
The reason I write this is misleading is b/c this is very different from a ZIRP YC era thinking that seems in line with your suggestion:
- Ex: JustinTV used their VC $ to pivot into Twitch, and if that didn't work, game over.
- Ex: Uber raised bigger & bigger VC rounds until self-driving cars could solve their margins / someone else figured it out. They seem to be figuring out their margins, but it was a growing disaster and unclear if they could with such an unpredictable miracle.
In contrast, both Amazon & Google were in positions of huge cash flows and being able to switch to highly profitable growth at any time. They were designed to control their destinies, vs have bankers/VCs dictate them.
Amazon was famous for not taking profit, and instead putting profit back into the company, but revenue? Amazon was generating gobs and gobs and gobs of revenue from day 0.
This phrasing makes me feel like we're living in some techno-future where corporations are de-facto governance apparatuses, squashing rebellion and dissidents :^)
Infinite sum games...
"infinite games" + "positive sum" => "infinite sum"
has big Sarah Palin energy: "refute" + "repudiate" => "refudiate"
https://www.npr.org/sections/itsallpolitics/2010/11/15/13133...
1. Inflection AI -- ceo out to MSFT
2. Stability AI -- ceo out to ____ (infinite sum games? with EigenLayer?)
What else? Is there a "GenAI is going great" website yet? (ala "web3 is going great": https://www.web3isgoinggreat.com/)
“ In reality, Mostaque has a bachelor’s degree, not a master’s degree from Oxford. The hedge fund’s banner year was followed by one so poor that it shut down months later. The U.N. hasn’t worked with him for years. And while Stable Diffusion was the main reason for his own startup Stability AI’s ascent to prominence, its source code was written by a different group of researchers. “Stability, as far as I know, did not even know about this thing when we created it,” Björn Ommer, the professor who led the research, told Forbes. “They jumped on this wagon only later on.” “
“ “What he is good at is taking other people’s work and putting his name on it, or doing stuff that you can’t check if it’s true.”
https://petapixel.com/2023/06/05/so-many-things-dont-add-up-...
Also all those 'controversies' were mostly the result of an aggrieved co-founder/investor who decided to sell their shares before SD1.4's success. Emad may not have proven to be competent enough to run an large AI lab in the long run, but those complaints are just trivial 'controversies'.
Paid off how?
Value can go up and down though..
I wonder how many ppl among VC/PE circles are also sugar coating their experiences and successes
I have to say, this is a quite common ignorant statement that's said about almost every CEO.
I'm not sure if there's more to it in this particular case, but no, CEOs aren't stealing your work. Similarly, marketers aren't parasites. Designers aren't there to waste your time. Many engineers seem to hold similar belief that others are holding them down or taking advantage of their work. This is just a congnitive bias.
Actually that's wrong, even the idea that engineers are smarter than managers is very prevalent here.
Where is the controversy here? Is the CEO expected to contribute to research? there seems to be some context I'm missing
If Emad had been supporting the research team from the beginning, one might argue that he created the conditions for their success, or whatever. But he wasn't there at that time.
None of this is related to whatever Emad is accused of. I'm just making a point about how we attribute contribution for success.
I have the DMs to prove this, and have not ever said something like this about someone. I wouldn’t make this accusation lightly, for whatever it’s worth. In the HN discussion of that article, I had left a comment, which Emad DMed me on Twitter about, saying no no, he never lied to investors, and tried to convince me that what I was saying wasn’t true. I was wondering why he cared so much. In retrospect it’s probably because it was correct.
I’ve never worked with him, to be clear, and the few colleagues who have worked at Stability have had generally positive things to say. But there was one that was screwed over by them hard (he was doing contract work, and never got paid for it), and I can think of at least four other alarming data points that all point to the same thing.
It’s unsettling not knowing whether to speak up about this. On one hand it doesn’t really matter that much. On the other hand, it’s the fundamental difference between a CEO that tends to IPO vs one that tends to fail. I hate seeing people fail, and I genuinely thought that my feelings about Emad were mistaken since empirically they were doing fine. Turns out, nope, not fine, and the original impression was right. Weird experience.
So there's that
He couldn’t even get his own story straight regarding his education and qualifications which should be a pretty clear disqualifying red flag from the outset.
The Forbes article from last year was dismissed on here as a hit piece but the steady flow of talent out was the clear sign, capped by the original SD authors leaving last week (probably after some vesting event or external funding coming through).
Reading explanations of "decentralized AI" [0] sounds like a sales pitch for investors that missed out on the AI and crypto hype. Technically, it sounds like a distributed file storage which already exists.
[0] https://www.forbes.com/sites/digital-assets/2024/02/24/decen...
I was not feeling confident about them as a company that I wanted to work for before that interview, afterwards I knew that was a company I wouldn't work for.
Today he resigned.
It was bound to happen anyways..
that wikipedia page screams "grifter", wow
Anyone else have information about them?
With his name attached, any crypto AI coin will launch straight to $500m mcap
Startup idea: Unprofitable.ai
That just sounds so simplistic that I don't believe he believes it himself.
[0] https://finance.yahoo.com/news/stability-ai-ceo-no-human-193...
People really oversell fiduciary duty. Yet the whole point of top-level corporate roles is to steer a company predicated upon opinion, which means that you have great latitude to act without malfeasance.
We (I) tend to use the term "programmer" in a generic way, encompassing a bunch of tasks waaay beyond "just typing in new code". Whereas I suspect he used it in the narrowest possible definition (literally, code-typer).
My day job (which I call programming) consists of needs analysis, data-modelling, workflow and UI design, coding, documenting, presenting, iterating, debugging, extending, and cycling through this loop multiple times. All while collaborating with customers, managers, co-workers, check-writers and so on.
AI can do -some- of that. And it can do small bits of it really well. It will improve in some of the other bits.
Plus, a new job description will appear- "prompt engineer".
As an aside I prefer the term "software developer " for what I do, I think it's a better description than "programmer".
Maybe one day there'll be an AI that can do software development. Developers that don't need to eat, sleep, or take a piss. But not today.
(P.S. to companies looking to make money with AI - make them able to replace me in Zoom meetings. I'd pay for that...)
Then we invented another AI to tell the assembler what kind of program you wanted, and called it a "compiler". All you had to do was tell the compiler what kind of program you wanted it to tell the assembler you wanted, and it would do all the not-exactly-programming work for you!
And so on...
My day job is programming in an environment which originated in the mid 90s. A contemporary of the Visual Basic era, but somewhat more powerful, and requiring substantially less code.
While I, and a few thousand others still use it (and it gets updated every couple years or so) it has never been fashionable. Ironically because it's perceived as 'not real programming'.
We routinely build systems with hundreds of thousands of lines of code, much of it founded in the 90s and having been added to for 25 years. Most of it was built, and worked on, by individuals, or very small teams. Much of it today is still active doing the boring business software that keep the lights on.
But its not "main stream" because programmers pick language based on popularity, and enterprises pick programmers based on language. A self-fueling cycle of risk aversion.
A lucky few though hot off the treadmill a long time ago and "followed a path less travelled by". And that has made all the difference.
“Mostaque had embezzled funds from Stability AI to pay the rent for his family's lavish London apartment” and that Hodes learned that he “had a long history of cheating investors in prior ventures in which he was involved”
They have raised 110M in October.
Stability transferred investor money to countless AI users in the world. Emad certainly didn't get a billion dollar payday.
Also not all CEO replacements turn out bad. Uber certainly has turned itself around.
1. Open source even more capable 'adult models' and get sued to oblivion (They still have massive lawsuits)
2. Neuter the model to uselessness and have users abandon it.
Both are bad choices, and require a very, very skilled CEO to thread the needle. Emad failed. that's all.
1. Some people are mad that Stable Diffusion might be trained on CSAM because the original list of internet images they started with turned out to link to some. (LAION doesn't actually contain images, just links to them.)
This one isn't true, because they removed NSFW content before training.
2. Some other people are mad that they removed NSFW content because they think it's censorship.
That actually isn't the legal issue I meant though. It's not that they trained on it, it's that it contains adult material at all (and can be shown to children easily), and that it can be used to generate simulated CSAM, which some but not all countries are just as unhappy about.
Several huge commercial and academic projects scraped billions of images off of the Internet and filtered them best they knew how. SD trained on some set of those and later some researchers managed to identify a small number of images classified as CP were still in there.
So, out of the billions of images, was there greater than zero CP images? Yes. Was it intentional/negligent? No. Does it affect the output in any significant way? No. Does it make for internet rage bait and pitchforking? Definitely.
Also, research isn't the only benefit, code generation, roleplay bots, are pretty good too.
Every single one of them converted flawlessly when I brought them into my markdown note tool.
If you're using chatGPT as nothing more than a glorified fact checker and not taking advantage of the multimodal capabilities such as vision, OCR, Python VM, generative imagery, you're really missing the point.
There are a few options. I work with a REPL, so I usually load the answer from a scratch file and put some representative data into it. When the result is wrong, which often happens, I feed ChatGPT the error, and it corrects the code accordingly. Iterating this results in a working function about 80% of the time. Sometimes it loses the plot and I either give up or start over with more detailed instructions.
You can also ask it to write tests, some of which will pass, some of which will fail. It's pretty easy to eyeball whether or not a test is valid, and they won't always be valid, I just fix those by hand.