what we did at fal is take the model and run it on our inference engine optimized to run these kinds of models really really fast. feel free to give it a shot on the playgrounds. https://fal.ai/models/fal-ai/flux/dev
what we did at fal is take the model and run it on our inference engine optimized to run these kinds of models really really fast. feel free to give it a shot on the playgrounds. https://fal.ai/models/fal-ai/flux/dev
Bummer. After seeing what was generated in the blog post I was excited to try it! Now feeling disappointed.
I was hoping it'd be more like https://play.go.dev.
Good luck.
Remarkably better than the "DrawThings" iPhone app (my only reference point).
Recently Claude began to allow generation of SVG drawings, and asking it to draw a unicorn and later add extra tails or horns worked correctly.
A fork exists in physical space and it's pretty intuitive to understand what it can do. These models exist within digital space and are incredibly opaque by comparison.
That sounds interesting! Were the results somewhat clean and clear SVG or rather a mess that just looked decent?
For what it's worth, I've previously asked in the Stable Diffusion Discord server for help generating a "lamb with seven horns and seven eyes" but the members there were also unsuccessful.
> A Gary Larsen, "Far Side" comic of a racoon disguising itself by wearing a fedora and long trench coat. The raccoon's face is mostly hidden by the fedora. There are extra paws sticking out of the front of the trench coat from between the buttons, suggesting that the racoon is in fact a stack of several raccoons.
Every human I've ever described this to has no problem picturing what I mean. It's a classic comic trope. AIs still struggle.
The intelligence just doesn't "fit" in there.
Personally I'm curious to see what would happen if someone burnt $100M of compute time on training a truly enormous image generator model, something the same-ish size as GPT4...
But AIs learn and therefore create in exactly the same way as humans, ostensibly on the same data. How can this be possible? /s
The first attempt at this based on https://reductress.com/post/my-boyfriends-are-always-two-kid... ... really misunderstood the image. This may also be part of the problem.
The image shows an imaginative, whimsical illustration of a character composed of two parts. The upper part features a man dressed in a long, elegant gray coat, wearing a bowler hat and round sunglasses, with a sophisticated white polka-dot ascot tie. His face has a subtle smile. The lower part of the character transitions seamlessly into a smaller figure of a cat, appearing to wear striped pants, with its tail visible. The entire character combines human and feline elements, creating a surreal, anthropomorphic appearance. The illustration is in black and white, emphasizing a stylized, cartoon-like design.
I then went to the image from https://www.reddit.com/r/DnD/comments/c6fdw4/oc_introducing_...And that provided:
The image captures a whimsical and secretive scene featuring three dwarves stacked in a totem formation, each attempting to conceal their nature under a large brown cloak. The top dwarf has a bright, cheerful expression and blond hair, holding the cloak wide to mimic wings, and is dressed in black armor adorned with teal gems and matching earrings. The middle dwarf displays a fierce expression, sporting a bushy orange beard, and is also clad in similar dark armor with teal embellishments. The bottom dwarf, an older figure with a long white beard, is adorned in a royal dark outfit with gold accents and a small crown, clasping a glowing white orb. This trio of dwarves, each with distinctive fantasy armor, unites in a playful attempt to disguise their stature and nature, adding an element of adventure and mystery to the scene.
Working off of that idea of the totem formation ... "Create an image featuring three children in a totem pole formation that are trying to conceal their nature in a single oversized trench coat."That produced https://imgur.com/a/Of9FsJl
I suspect the orange beard came from the previous part in the session. But that might be an approach to take in trying to describe it in a way that can be used.
It’s more like a giant telescope of many lenses (the latents from the prompts) and you’re adjusting the lenses to bring a possible reality of many into focus.
"a woman lying on her back wearing a blouse and shorts."
But it wouldn't render the image - i instead got a NSFW warning. That's one way to hide the fact that it cannot render it properly i guess...
PS: after a few tries it rendered "a woman lying on her back" correctly.
Also, everybody should remember that these models are not copyrightable and you should never agree to any license for them...
It would be nice here if you give some examples of what you call open source model. Please ;) Because the impression is that these things do not exist, it's just a dream which does not deserve such a nice term..
However, plenty of open source software exists. The fact that open source models don't exist doesn't excuse attempts to falsely claim the prestige of the phrase "open source".
You are wrong about that. It's a file with numbers. Which makes it a database or dataset and very much protected by copyright. That's why licenses are needed. For the phone book, things like open street maps, and indeed AI models.
> The fact that open source models don't exist
The fact that many people (myself included) routinely download and use models distributed under OSI approved licenses (Apache V2, MIT, etc.) makes that statement verifiably wrong. And yes, I do check the license of stuff that I use as I work with companies that care about such matters.
> As far as I know ...
Now you know better.
This is only true in jurisdictions that follow the sweat of the brow doctrine, where effort alone without creativity is considered enough for copyright. In other places, such as the USA, collections of facts are not copyrightable and a minimal amount of creativity is required for something to qualify as copyrightable. The phone book is an example that is often used, actually, to demonstrate the difference.
Not every collection of numbers is a database, and a database is not the same thing as a dataset.
Databases have limited copyright-like protection in some places. Under TRIPS, that extends to only databases that are "creative by virtue of the selection or arrangement of their contents" or something along those lines. In the US they talk specifically about curation.
ML models do not meet either requirement by any reasonable interpretation.
> The fact that many people (myself included) routinely download and use models distributed under OSI approved licenses (Apache V2, MIT, etc.) makes that statement verifiably wrong.
The "source code" of an ML model is most reasonably interpreted as including all of the training data, which are never, ever available.
Now you know better.
[On edit: By the way, the people creating these works had better hope they're outside copyright, because if not, each one of them is a derivative work of (at least some large and almost impossible to identify subset of) its training data, so they need licenses from all the copyright holders of that training material, which few of them have or can get.]
However, transformativeness is a factor in whether or not there is a fair-use exception for the derivative work. And these models are highly transformative, so this is a strong argument for their fair-use.
"Fair use" is pretty much entirely a US concept, and similar concepts in other countries aren't isomorphic to it.
The model does have a radically different form from its inputs. So you could easily imagine that being "transformative enough" for US fair use. A lot of the other fair use elements look pretty easy to apply, too. Although there's still the question of whether all the intermediate copies you made to create the model were fair use...
In fact, I'll even concede that a court could find that a model wasn't a derivative work of its inputs to begin with, and not even have to get to the fair use question. The argument would be that the model doesn't actually reproduce any of the creative elements of any particular training input.
I do think a finding like that would be a much bigger stretch than a finding that the model was copyrightable. I could easily see a world where the model was found derivative but was not found copyrightable. And it's actually not clear to me at all that the model has to be copyrightable to infringe the copyright in something else, so that's another mess.
Somewhat related, even if the model itself isn't infringing, it's definitely possible to have most models create outputs that are very similar to (some specific examples in) their training data... in ways that obviously aren't transformative. Outputs that might compete with the original training data and otherwise fail to be fair use. So even if the model is in the clear, users might still have to watch out.
What criteria for copyright protection are they missing?
I can tell you a secret. What you call 'open source' models are impossible. Because massive randomness is a part of training process. They are not reproducible. Having everything you cannot even tell if the given model was trained on the given dataset. Copyright is a different thing.
And a bad news, what's coming is even worst. Those will be the whole things with self awareness and personal experience. They can be copied, but not reproduced. More over, it's hard or almost impossible to detect if something undeclared was planted in their 'minds'.
All together means 'open source' model in strict interpretation is a myth, great idea which happen to be not. Like Turing test.
> However, plenty of open source software exists.
Attempt to switch topic detected.
PS: as for that massive downvote, I even wasn't rude, don't care. This account will be abandoned soon regardless, like all before and after.
The Llama models aren't. Some of the Mistral models are (the Apache 2 ones). Microsoft Phi-3 is - it's MIT.
I've decided to draw my personal line at Open Source Initiative compliance for the license they release the model itself under.
I respect the opinion that it's not truly open source unless they release the training data as well, but I've decided not to make that part of my own personal litmus test here.
My reasoning is that knowing something is "open source" helps me decide what I legally can or cannot do with it when building my own software. Not having access to the training data downs affect my legal rights, it just affects my ability to recompile myself. And I don't have millions of dollars of GPUs so that isn't so important to me, personally.
Tough beans? There's lots of actual software that can't be open source because it embeds stuff with incompatible restrictions, but nobody tries to redefine "open source" because of that.
... and, on a vaguely similar-flavored note, you'd better hope that the models you're using end up found to be noninfringing or fair use or something with respect to those "unlicensed data", because otherwise you're in a world of hurt. It's actually a lot easier to argue that the models aren't copyrightable than it is to argue that they're not derivative of the input.
> I've decided to draw my personal line at Open Source Initiative compliance for the license they release the model itself under.
You're allowed to draw your personal line about what you'll use anywhere you want, but that doesn't mean that you should try to redefine "open source" or support anybody who does.
Never underestimate the value of getting hordes of unpaid workers to refine your product. (See also React, others)
I'd prefer "false advertising" - it's more direct and without the culture war baggage.
That said, I don't think outputs of the model are derivative works of it, any more than the model is a derivative of its training data, so it's not clear to me they can actually enforce what you do with them.
Are you talking about https://en.wikipedia.org/wiki/Database_right or plain old copyright?
I'm no IP lawyer, but I've always thought that copyright put "requirements" on the artefact (i.e the threshold of originality), not the process.
In my jurisdiction we have database rights, meaning that you get IP protections for the artefact based on the work put into the process. For example a database of distances between adress pairs or something is probably not copyrightable, but can be protected under database rights if enough work was done to compile the data.
EDIT: Saw in another place in thread speaking about the https://en.wikipedia.org/wiki/Sweat_of_the_brow doctrine, relates to Database rights. (Neither of which notably are not applicable in the U.S)
The only thing that's really specified about the model itself is its architecture, which is (1) dictated by function, and (2) usually deeply stereotyped.
Fair enough, but those datasets are also primarily copyrighted material. If the software here merely transforms the input material (which I agree it does), then the output is a derivative work.
If I take a string of data from a true hardware RNG, XOR it with a Taylor Swift song, and throw away the original random stream, is the resulting fundamentally random bit string still a derivative work of the song? As with the ML model, you can't recognize the song in it. And as with at least some training examples in the inputs of most ML models, you can't recover the song from it either.
It feels like the test for whether X is derivative for copyright purposes should include some kind of attention to whether X is a creative work at all. Maybe not, but then what test do you use?
I do recognize the possibility that the models might not themselves be eligible for copyright as independent works, yet still infringe copyright in the training inputs. It seems messy, but not impossible.
... and as I said elsewhere, it's also messy that while you generally can't recover every training input from the model, you can usually recover something very close to some of the training inputs.
It's not a copy of it, and when you distribute it you're not distributing the original. So it's not a derivative for copyright purposes.
It can still be a derivative for other legal purposes. Judges don't appreciate it when you do funny math tricks like that and will see through them.
> It feels like the test for whether X is derivative for copyright purposes should include some kind of attention to whether X is a creative work at all. Maybe not, but then what test do you use?
Yes, that's how US copyright law works. (well sort of…)
Being a transformative work of something makes it less of a copy of it, the more transformed it is, since it falls under fair use exemptions or is clearly a different category of thing.
If a model was a derivative of its training data, then Google snippets/thumbnails would be derivatives of its search results and would be illegal too. Unless you wrote a new law to specifically allow them.
In other countries (Germany, Japan) fair use is weaker, but model training has laws specifically making it legal in certain circumstances, and presumably so do Google snippets.
A compressed (or normally encrypted) version wouldn't be a copy that way, either, but I would still absolutely go down for distributing it. The difference is that the compression can be reversed to recover the original. Even lossy compression would create such a close derivative that nobody would probably even bother to make the distinction.
You're right that "math games" don't work in the law, but that cuts both ways. If you do something that truly makes the original unrecoverable and in fact undetectable, and if nothing salient to the legal issues at hand about the new version derives from the original, then judges are going to "see through" the "math trick" of pretending that it is a derivative.
> then Google snippets/thumbnails would be derivatives of its search results
Thumbnails are legally derivative works, in the US and probably most other places. In the US, they're protected by the fair use defense, and in other places they're protected by whatever carveouts those places have. But that doesn't mean they're not derivative works.
In fact, if I remember the US "taxonomy" correctly, thumbnails are infringing. It's just that certain kinds of ingfringement are accepted because they're fair use.
If thumbnails weren't derivative works at all, then the question of fair use wouldn't arise, because there can be no infringement to begin with if the putatively infringing work isn't either derivative or a direct copy.
Where thumbnails are different from ML models is that they're clearly works of authorship. In a thumbnail, you can directly see many of the elements that the author put into the original image it's derived from.
The questions are (a) whether ML models are works of authorship to begin with (I say they're not), and (b) whether something that's not a work of authorship can still be a derivative work for purposes of copyright infringment (I'm not sure about that).
So far as I know, neither one is the subject of either explicit legislation or definitive precedent in most of the world, including the US.
Since it costs millions to produce one of these models, it's not just taking the software and running it to compile them.
Thanks for pointing that out @Hizonener
I'd suggest re-wording the blog post intro, it reads as if it was created by Fal.
Specific phrases to change:
> Announcing Flux
(from the title)
> We are excited to introduce Flux
> Flux comes in three powerful variations:
This section also comes across as if you created it
> We invite you to try Flux for yourself.
Reads as if you're the creator
This library is quite well known, 3rd most starred project in Julia: https://juliapackages.com/packages?sort=stars.
It has been around since, at least, 2016: https://github.com/FluxML/Flux.jl/graphs/code-frequency.
I hope this one doesn't stir as much discussion. It has 4000 stars, there isnt a large mass of people who view the world through the lens of "Flux is ML library". No one will end up in a "who is on first?" discussion because of it. If this line of argument is held sacrosanct, it ends up in an infinite loop until everyone gives up and starts using UUIDs.
https://en.wikipedia.org/wiki/Go!_(programming_language)
Disclosure: I work at Google but not on the Go team.
also search engines are context aware, if your search history is full of julia questions, it will know what you're searching for
Flux A is the ML library
Flux B is the T2I model
Flux C is the React library
Flux D is the physics concept of power per unit area
Flux E is the goo you put on solder
If it's unlimited or "throttled for abuse," say that. Right now, I don't know if I can try it six times or experiment to my heart's desire.
I’d bet that fine art training would further improve the compositional skills of the model, plus it would open up a range of uses that are (to me at least) a bit more interesting than just illustrations.
Does it respond to any names? I noticed SD3 removed all names to prevent recreating famous people but as a side effect lost the very powerful ability to infer styles from artist names too.