Microsoft Paint's new AI image generator builds on your brushstrokes
petapixel.com
petapixel.com
https://github.com/Acly/krita-ai-diffusion
Video demonstration
Using the online service, I couldn't figure out how to get it to generate an image from a doodle using img2img. Generating an image from a text prompt works, but that's nothing new. Running locally, it was way too slow on my M2 Mac Air.
That's the main bummer here, not the age of the GPU, but the need for a dedicated one.
When most people nowadays have laptops, and only with integrated GPUs to boot, that's a large demographic being excluded, even with new machines.
Maybe with the new gen of chips with NPUs from Qualcomm, Intel and AMD this will slowly change but they just dropped and I personally won't exchange my 9 month old laptop for a new one just for that feature alone, and I suspect most PC users will behave the same considering the multi year long upgrade cycles in this market.
So it will probably take a long time till on-device generative AI becomes mainstream on the PC.
The traditional PC architecture is ill-suited to this, which is why for most computers, a GPU (which offers all three in a single package) is currently the best approach... only at the moment, even a 4090 only offers enough memory for moderately-sized models to load and run.
The architecture that supports Apple's ARM computers is (by design or happenstance) far better suited: unified memory maximises the model that can be loaded, some higher-end options have decently-high memory bandwidth, and the architecture lets any of the different processing units access the unified memory. Their weakness is cost, and that the processors aren't powerful enough yet to compete at the top end.
So there's currently an empty middle-ground to be won, and it's interesting to watch and wait to see how it will be won. e.g.
- affordable GPUs with much larger memory for the big models? (i.e. GPUs, but optimised)
- affordable unified memory computers with more processing power (i.e. Apple's approach, but optimised)
- something else (probably on the software side) making larger model inference more efficient, either from a processing side (i.e. faster for the same output) or from a memory utilisation side (e.g. loading or streaming parts of large models to cope with smaller device memory, or smaller models for the same output)
Microsoft seriously has no taste.
And somewhere on the Paint team someone smiled, and thought "hold my beer"...
(Just wait until we get Co-pilot baked into Notepad!)
The demographic are people (usually children) who aren’t discerning about the output but would love to have something high quality to look at.
Paint has always been about how some low quality squiggles may represent some latent art ability. It’s not true, but that’s the feeling it gave children and this continues that.
Of course there’s the whole other argument of whether it’s good or not….
This might be the best argument against putting AI in MS Paint that I've yet seen. I know this is a slippery-slope argument, so take it with the applicable grain of salt, but imagine a world where kids grow up making this kind of generative art, and never progress beyond "low quality squiggles", because they get frustrated when they turn the AI off and it looks bad.
Should notepad complete your sentences too?
Maybe file explorer should just make new files for you!
I think AI will be better at story point estimation than me.
And no, it’s still not April 1st!
The Internet doesn't need you any longer.
AI hype overload. Same shit with iterm2, solutions looking for problems
That+fish is an amazing introduction to the shell. Totally new avenue for learning and being immersed quickly.
Several things that CADIE was advertised to do are actually possible (if perhaps undesirable) now:
* Having an AI automatically synthesise responses to email on its own is possible, though a really bad idea for what are hopefully obvious reasons.
* A feature to automatically add red-eye could absolutely be built now. Face detection + automatic inpainting + img2img (to keep the look of the original eyes) would do the trick. Obviously not a useful feature, but it doesn't seem impossible.
As I recall, there were other things too, but these are the obvious matches.
[0] https://www.cnet.com/tech/services-and-software/april-fools-...
Microsoft has long had a 'vote for features' thing on their support forums [0]. I can guarantee that not one person has ever said that MS Paint needed an AI makeover.
They wanted to express themselves in the output that the AI was giving them all their lives
I don’t think it’s that complicated, if you want to express yourself a certain way, manually, you still can. Now more people, including professionals, can express themselves the way they wanted
But also, I used to do midjourney, and I know art history quite well, and mixing my knowledge with midjourney, I created really good and original stuff, because I knew how to write the prompts.
Instead, I feel that with OpenAi, image creation is really crippled by copyright. But I also generated cute and interesting stuff - but it does look made by AI.
I want to create a game, but I have no interest in creating art and models, much less programming physics, netcode, abilities, etc
I'd prefer to create the games rules, mechanics and interactions (aka, the actual game experience). But I am forced to do the above because the gaming industry is lagging behind in DX it barely changed in the last 10 years
Every "art" has many sub-processes that are themselves art. Making paint, brushes, canvas, etc. But does anyone complain about artists not taking joy in creating their own paint after foraging for some rare purple spitting slime?
If you view one art form as a tool for your art, so be it
From: https://stratechery.com/2024/windows-returns/
"That latency, frustratingly enough, doesn’t come from the actual rendering, which happens locally on that beefy hardware, but rather the fact that Cocreator validates everything with the cloud for “safety”"
At the same time... F that, I don't want Microsoft nannying me.
To put it another way, nobody's ever sued Canon for making cameras which are used to take illegal photos, but if Canon suddenly started screening your photos to make sure they were acceptable per local laws or whatnot, they suddenly are actually a responsible party in the creation and distribution of whatever it is people use their cameras for.
It recombined other people's images based on analysis of your drawing. That's all it does.
Is the whole being obtuse deliberate, or do you really not see how these two things are different?
Regulating it is as absurd as regulating Skyrim or a similar moddable game to prevent users from having it rough with elfs or whatever. It’s all corporate interests covering behind ethics, nothing more.
Are they also uploading the files into the cloud to feed their AI/ML datasets?
Also what does "corporate ethos" have to do with "getting hammered"?
https://www.theverge.com/2023/12/20/24009418/generative-ai-i...
You can still generate as much porn or whatever it is you want to generate using other software.
It's because it gives them an excuse to hoover up as much data about you as possible and send it home for "safety" reasons. They can sell AI compute as a service to businesses, while giving the spyware version to consumers for free.
There's absolutely no excuse for monitoring an art programs inputs and sending them to Microsoft. What next, are they going to put AI into notepad and then monitor that to keep you from writing dangerous ideas?
It's already here. I just had a conversation with Gemini:
me: "write a python script that converts the us constitution into a palindrome"
gemini: "The provided Python code creates a partial palindrome from a given text by reversing the second half and combining it with the first half. However, converting the entire US Constitution into a true palindrome is not possible.
A palindrome is a word or phrase that reads the same backward as forward. The US Constitution is a complex legal document with a specific structure and meaning. Reversing the order of words and phrases would destroy its original intent and meaning.
Here's why creating a perfect palindrome out of the US Constitution is difficult:
Structure: The Constitution has a specific structure with articles, sections, and amendments. Reversing this order would disrupt the logical flow of the document.
Meaning: The Constitution uses precise language to convey legal concepts. Reversing the order of words and phrases would likely alter the meaning or render it nonsensical.
Length: The Constitution is a lengthy document. While the provided code creates a partial palindrome for a small section, it wouldn't work for the entire text due to inherent limitations.
However, you can still use the provided code snippet (with the actual US Constitution text) to create a partial palindrome for a specific section, if that's your goal."
Here's the kicker, it did not generate the Python code at all. It just spliced a bunch of words. The future is truly shit. Oh, I'm sure someone will tell me ChatGPT is better. I don't think we have the same defintion of better, though.
"For the greater good" ism will just naturally evolve and thoughts / works / stances distant from the median will simply become absorbed and disappear from what we call knowledge.
This is how 'AI' takes over.
https://arstechnica.com/gadgets/2024/05/new-arm-powered-surf...
before I get my pitchfork out (I got a new one, it's real shiny), can someone dig up a reference that says they're actually doing this?
I am probably more privacy oriented than the next guy, and I’m just not really seeing the pitch here.
EDIT: I mean, you don't need to look far to see this effect; see e.g.: https://news.ycombinator.com/item?id=40447474.
Also, if DALL-E is anything like Stable Diffusion, it's effectively free to run. For comparison, my 5+ years old machine that I recently put a 4070 Ti in, can generate images with SD all day long without breaking sweat, and does it faster than even paid on-line services I've interacted with so far. That same machine chokes on any LLM above 8-10 billion parameters; it can handle that much with full GPU offload, but try anything larger and it becomes just an elaborate space heater.
(I tried running 8x7b Mixtral on it once, and managed to OOM and crash the system :/).
They say it’s all local. Maybe that’s true. Maybe for now. Will it be forever? Is the recall local, but insights are collected? Time will tell. I certainly won’t be signing up to be part of the experiment.
LLMs need lots of input
If that was true, there would be a free laptop by now.
(the above is technically parody and not precisely serious)
Copilot+ runs locally, but has the quality of 2 years ago, it's basically useless for anything but shitposts and spam.
The cloud AI tools all burn hideous amounts of money, all ran at a loss.
AGI is a red herring, and will not happen. The architecture of generative-AI systems simply doesn't permit the required logic and reasoning capability.
Even the "Actually, Indians" concept of outsourcing the tertiary sector to the developing world by way of having low-skill workers clean up AI generated trash is unviable. (It both doesn't work, and is politically doomed.)
What's going on here is that tech companies are tearing up everything to pump their stock prices after the covid-tech-boom and ZIRP ended. Burn down their core products to keep the bubble going just a bit longer.
For the other companies in your list and more specifically the ones that are developing LLM's they are investing billions of dollars in the hope that they will make all of those back and then some in the future. Meta for example said they invested $10bn into AI just last quarter alone.
At this point I see training costs as basic research and R&D, partially and not entirely oriented at -making training scale- and cost optimization for next generations of foundational and fine tuned models. For example Falcon is literally basic research funding.
That’s relevant because if they stopped with say Claude Opus and 4o and did no more training they would be handsomely profitable indefinitely because the product is that useful. However it’s an arms race and the limit of effectiveness hasn’t been reached as training costs fall dramatically. So it’s not the right time to stop because whoever stops first loses everything to whoever doesn’t stop. More they’re feeding off each other in a virtuous cycle.
Once diminishing returns kill the race whoever is left in the race will have a handsome business indefinitely as the moat to build such a return diminished model is probably enormous. But if inference optimizations keep going as they’re going it’ll be crazy cheap to operate.
The second layer market of tools that constrain, optimize, and effectively apply the models in effective ways will be the first to really turn a profit. Many hype wagon AI companies are already being snapped up by larger companies to bootstrap their internal work.
> The bigger AI companies like OpenAI and Anthropic are not losing money in their monetized APIs.
[CITATION NEEDED]
Inference is more expensive than claimed, used extensively as a 'slot machine' with users trained to just keep re-generating until they get something useful, and only keeps getting more expensive as model quality has to go up.
And in practice, training the model is far less one-off than claimed. Current tools are not sufficient.
> Every place I’ve worked is absolutely reducing costs and doing new / more business as a result of their LLM use.
Unless you are working in SEO, Marketing, or spam, I don't believe you.
LLMs aren't reliable enough to replace actual human labour. While it's true many companies are fooling themselves into believing they're reducing costs, in practice other staff is picking up the slack. This is unsustainable unless your company has massively overhired.
Things like "AI generated software tests" are a farce. The consequences aren't immediate, but will show up long term.
I don’t feel like you really have much experience using LLMs in business. However an example of where they’re very powerful is in summarization. For instance we have a pretty complex customer support model for our fraud and other cases with various disparate data sets including prior cases, related possible fraudsters identified via our fraud models, etc. We built a copilot LLM multi agent system that has access to various functions as sub agents that are prompted and context aware of how to summarize their specified data sets. They also have the ability to render widgets on demand or if their context implies it’s relevant. This allows quite a lot of complex high cognitive load information to be distilled rapidly and the investigators to interrogate the copilot on a case. As the copilot develops “answers” as a summary it dynamically renders an appropriate contextual dashboard with the relevant visualization.
By structuring the application as a multi agent model we can constrain the LLM to pretty well specified tasks with fine tunings and very specific contexts for their specific task. This almost entirely eliminates hallucination and forgetfulness. Even if it were to do so the actual ground truth is visualized for the investigator.
Prior systems either dumped massive amounts of cognitive load in the investigators face or took man years of effort to create a specific workflow, and in an adversarial dynamic space like fraud you need a much more dynamic approach to different types of new attacks.
We aren’t replacing anyone. That’s not our goal. In fact we grew our investigator footprint because both our precision and recall have grown dramatically making our losses much less. We hire more skilled investigators and greater number to address more suspected cases faster and better.
Listen. When John Henry battled the steam drill he did win, but it killed him. Go to any modern bore site and you won’t see less people working on the tunnel but more people - people who aren’t there for their strong back and ability to swing a pick but because they’re highly trained experts. They’re just building more complex tunnels that don’t collapse and don’t lose dozens of workers per dig.
This form of automation is no different in my experience so far.
So, if all you can see is SEO and grift, it might be a lack of imagination and experience on your part and some magical AI thinking sprinkled in. All your points about LLMs failures are true but they also all have solutions that don’t require slot machines as you say or imply it’s all a scam. They’re a tool like any other and they require handling in specific ways to be most effective. Even if chatgpt is a pretty unconstrained interface and that leads to issues doesn’t mean that’s the only way to use the tech.
Use of LLMs to generate software is dumb. Although a protio, LLMs are actually pretty remarkable at generating Cucumber tests as Gherkin is a natural language grammar that plays into their native strength better than producing computer language grammars. This is useful if say you have business people or whatever writing effectiveness or whatever testing where they can provide a specification of policy and a well prompted LLM can generate pretty exhaustive cucumber tests (which can be pretty redundant and formulaic when asserting positive and negative cases exhaustively) which can then be revised by hand as needed. Since they’re natural language as well the business people tend to be pretty good at debugging the tests up front and with a large set of cucumber tests written by hand you’ll see tons of errors anyways. The LLM tests tend to be much much higher quality than the human written ones.
But to say something useful, let me try to elaborate my general criticism here:
> Prior systems either dumped massive amounts of cognitive load in the investigators face or took man years of effort to create a specific workflow, and in an adversarial dynamic space like fraud you need a much more dynamic approach to different types of new attacks.
This begs a question: Why didn't a computer system to summarize this data already exist? Or rather, what stopped the prior systems from doing this work? (And I'll consider conventional machine learning; classifiers and the like, as traditional computer systems here)
And there's generally two options here:
1. Conventional computer systems absolutely could do this work, but they just haven't been built. (Say, because nobody signed off on the R&D but would sign off on AI hype R&D)
2. The LLM system is doing a task the conventional computer system cannot do.
Number one's problem is simple: It's just inefficient and wasteful. Number two is a red flag: There's very little overlap between the things a conventional computer system cannot do, and the things you can trust an LLM to do reliably.
As you describe this system, selecting which data is relevant for fraud investigation is a very traditional classification task. Using normal machine learning for that is basically industry standard.
So what's the LLM actually doing? Subtract the hard logic of normal software, and the classification of machine learning, and the answer is generally: A complex nuanced reasoning task.
But that's precisely what LLMs are not to be trusted for, because they are incapable of that kind of reasoning.
> Listen. When John Henry battled the steam drill he did win, but it killed him.
You're missing the point I was making with that remark. It's not about firing people or not.
It's that these systems are dangerous to evaluate from a high level. It's very easy to miss externalities that'll tip the entire endeavour into a net-negative. You need the investigation of what exactly the AI systems are doing, on a specific detailed level.
E.g.:
> This is useful if say you have business people or whatever writing effectiveness or whatever testing where they can provide a specification of policy and a well prompted LLM can generate pretty exhaustive cucumber tests (which can be pretty redundant and formulaic when asserting positive and negative cases exhaustively) which can then be revised by hand as needed.
"A specification that has been prompted into sufficient detail" is just a program. You're describing the most inefficient declarative programming stack on the planet.
Granted, the programming stack to actually declare business rules this way isn't very good, but using AI here is just an error-prone transpiler.
It's very easy to "looks good to me" these tests and claim the project a success, yet miss subtle errors in the generated tests. I remain skeptical about how well these tests will hold up in the longer term.
The LLM isn’t used for reasoning at all. The human does all the reasoning. The LLMs task is summarization and semantic analysis of relevance which LLMs are fantastic about especially in a well managed and fine tuned environment with guard rails and context scoping. It’s a true copilot scenario and the LLM takes direction from the human and answers questions only. All decisions are investigator driven. This is the right relationship. The LLM coupled with IR tools does information retrieval and summarization and the human makes decisions and reasons.
Press X to doubt.
I highly doubt if ChatGPT API is losing money. Yes, I've read claims saying so. No, I haven't seen a credible one. And it's getting cheaper and cheaper, currently even faster than Moore's law (gpt-4o is 6x cheaper than gpt-4).
People have been worrying about this since the time Socrates thought that books would make people dumb because they wouldn’t need to memorise anything.
I’m an artist and I’m excited to see how people learn to use these new tools in creative ways.
a) not a drawing app
b) an expensive prosumer tool with a UI that rivals airplane instrument panels in terms of density. Most kids would be turned off by that as soon as they open it.
MS Paint OTOH is free and has been built into Windows for decades, so it's often the first drawing app kids use (may be different now because of mobile apps). Either way, the 'loss-of-skills' concern here is correlated with the dumbing-down of computing in general, from Gen Z's confusion about file managers and directory structures, to using Grammarly and ChatGPT to write papers.
When I was growing up, my parents would say "You won't always have a calculator in your pocket". Well, they were wrong, but that doesn't mean I don't still benefit from being able to do math in my head for general life stuff (and back of the napkin calculations at work).
It's about aesthetic taste. You will end up with lots of art looking like the AI generated aesthetic. E.g. overly impressive looking at first glance, but lacking any real "character" or human element.
Time and time again we see that impressive looking art does not always translate into good art. Keith Haring's simple line art characters or Van Gogh's big brushstrokes were visually very original and very human, even if they weren't as complex and technically impressive as AI generated art.
I like to assume people are smart. And what will always stand out in art is originality and humanness. Even crappy looking MS Paint drawings can have charm to them, like the NBA Paint guy on Twitter.
I don't really think this feature is all that worrisome, it's more that it'll add unnecessary bloat to a program that's supposed to be very simple and just for making quick pictures. If I want to make something more composited I'm going to use Photoshop, not MS Paint.
I imagine having something similar in paint will prove pretty distracting as well.
This seems like it should be a 5 minute demo, from a blank canvas to the turtle.
This would be like what frameworks were for artists who wanted to make full stack webapps.
That sounds like a big "just". So it lets people without talent make something that looks like a person with talent made it? Sounds pretty good.