Photoshop for Text
stephanango.com
stephanango.com
Photoshop doesn’t use AI models to generate new things for the artist, photographer, whoever (it does have generation of course, but not in the same way). A tool like Photoshop for text, in the way described, seems more like a writing aid than a writing tool; something I can give vague commands like “make this paragraph more academic” (whatever that means), and it’ll spit out something approximating the academic style it analyzed. Whereas “auto balance the levels of this image” is much more concrete in what it means, there’s no approximation of “style.”
I feel like Photoshop for text should be an editor that cares more about the structure of stories and who’s in them, with ways to organize big chunks in easy ways, rather than something that generates content for you.
I’m decent with Photoshop but have not had a need for it lately. This makes me want to fire it back up.
Each video is 1-2 minutes long. Mind blown again.
It certainly does have this, such as tools to remove objects and paint over them as if they were not there, automatic sky replacement (replace a cloudy sky with a sunset), super resolution (AI up scaling) and a range of what they call 'Neural Filters'.
You want your low quality boring midday image of some famous bridge to be a high resolution, taken during sunrise without that person riding a bicycle? Photoshop will do it with very little user skill or input.
Some beta features include style transfer, makeup transfer and automatic smile enhancment.
Hope this question is not too dumb!
"Families enjoy this restaurant"
And started typing and changed "Families" to "our family" this "photoshop for text" would change "enjoy" to "enjoys" for you.
Enough little facets like that, I could see being useful.
I can see the effect of a change to an image at a glance. Judging whether it is a much worse outcome takes seconds usually.
With wholesale changes to text, seeing if it inadvertently makes something worse takes minutes. It's orders of magnitude slower. That makes the editing process much slower, whilst still fraught with risk.
I am not talking about not trusting AI to get it right. But about a general change that happens not to work. Maybe your academic writing was better of with one narrative paragraph. Maybe you want to change tense but forgot to mark a quote and you are putting false words in someone's mouth. We proofread changes by human editors. AI editors will require the same.
An image is 2D in space, 0D in time. Because we perceive 3D/1D, i.e. one dimension above in each aspect, we can see a whole image at once.
We can't see an entire video (2D/1D) at once - we have to move in time. We can't see an entire 3D object at once either - we have to move in space.
Technically, text is just as visual as an image. However, meaning is conveyed by the sequence in which we interpret the symbols, and that sequence is 0D in space, 1D in time.
Maybe beings in an 2D time dimension would be able to parse text instantly?
As with anything the shorter the feedback loop the better. The difference being that an image has an information density which is much higher than text, "image worth 1000 words", and you can make sense of it easier.
If you want to edit lots of text and see if it flows well quickly, some other ways would involve mapping the textual representation to something we can perceive quicker. The same way sometimes systems will have different sound tones for different actions, making a skilled operator able to detect mistakes by "listening" to the UI. If you can represent textual changes in an aggregated visual form, you might get what you describe.
For example if you could make nice grammar produce a nice musical sound when "read" by a computer program, you could potentially assess for correct grammar quicker by listening to the whole thing in much less time than you'd take reading it.
Arithmetic: quantity.
Geometry: quantity in space.
Music: quantity in time.
Astronomy: quantity in space and time.
Back in my uni days, the student centre featured a painting which was a representation of a musical piece --- something classical, possibly Beethoven's Fifth. It displayed in space (or more accurately, in a plane), what was usually performed over time. The painting was more conceptually interesting than visually appealing, though I like the idea.
Whilst I think you have a point re: images, there are often those which reward a more measured appreciation. A "where's Waldo" type visual puzzle might be a more trivial example but there are images which reveal themselves over time.
There's also what works leave us with. Text ultimately conveys, well, textual or verbal information. Speech is similar. Both can deliver a mood, though that's not essential. Music on the other hand seems to me to be far more emotional. Images in their simplest form are literally iconographic, more representing something than portraying it, though with detail they tend toward the latter. (Plastic arts such as sculpture seem similarly iconographic, and of course, the original icons were often such portrayals.) Video and drama seem to me more related to music than texts, working on moods and emotions. That though is occurring to me as I write this, as well as the realisation that both often rely heavily on a soundtrack. In the case of opera or musicals, to a dominating extent, less so in a straight stage play. And of course, on television, there's the infamous laugh track to guide us in our emotional response to a sitcom, standing in for the immersive live-audience experience.
Back to the suggestion: the idea of being able to autotune a text, so to speak, seems ... possible, now or soon, with advances in AI, GPT3, and the like. But as rocqua said, far harder to take in at a glance as compared with visual or video arts.
Most people are bad at writing. This will make them "better" at writing, but only in a certain way. I love reading things people write, because it gives me a window into their mind, how they think, who they are at a deep level. That is the joy of reading, and the structure of writing is a big part of that.
Now, a lot of people's writing has become homogenized by the computerization of our world already, but only at a low level: spelling, basic grammar, and so on.
When people have the ability to inpaint whole paragraphs, dreamed from the blob of internet text (which is mostly corporatized, computerized, email-ized, sterile in the way described above) we will lose something essential.
And another problem arises. I send an email asking something, to a coworker or to a friend. They inpainted their response. Did they really understand what I was asking? Did we really communicate at all?
Images, I guess, suffer from the same problem. But images are less interpersonal. They are communication, but not communication like writing is communication. In nearly all circumstances, images are less subtle than words are (the subtlety of visual communication happens irl, where such machines have yet to insert themselves).
Worse yet, not only will the 'in-painting' obscure the writer's meaning, or even whether they even had meaning, it will ALSO render that meaning more generic, and eliminate the most information-dense surprising bits.
All of these tools are merely synthesizing new text/images/code/etc. from millions of existing examples. When there is any doubt about the intent or the output, it fills in the MOST EXPECTED output. It does NOT fill in the possibly intended but unique image, phrase, code that would be the output of a unique insight. Real brilliance will be simply lost in the sauce.
Ugh. I won't be using these, and I'll shun those who do.
Communication only has value if it is surprising. That much has been known since Shannon.
Put another way, their prompt is arguably a pointer to data buried in that model. You're raising the question of whether they should only send the pointer, or just cache the query result, in a sense, and save you the trouble of looking it up.
But, when they are trying to express what WE are saying, it looks like a very lossy solution at best.
I've had a deposition taken with an "AI" stenographer, and it was horrific, frequently reversing the meaning of sentences I said, or replacing an uncommon name with a common name (e.g., "John Kemeny" replaced with "Jack Kennedy"). Of course the transcript LOOKS great, it doesn't have any of the "(unintellegibile)" notations of a human transcript. It also does not go back at break points and ask for proper spellings of names, addresses, etc. like a human transcriber.
This is in the context of a legal trial with consequences, and I'm horrified to see this kind of crap passing for usable products, and here we are looking to foist it off on the general public as writing tools. We're forking doomed by smart idiots looking to make a quick buck with novel "tools".
That is to say, these tools will exist and they will be used. But people will still write and make art - often perhaps using these tools in some way - because they'll have the time and find it rewarding to do so. And others will also assign some value to this. We might briefly get addicted, and then have a society-wide discourse on what healthy use is, similar to social media.
Isn't it compelling to wonder what humanity will decide to do with technology when technology were to be limitless? As in, what essentially human choices will we make in what, when and how to use technology? (Singularity-themed scifi tries to provide some answers since the 90s.)
Easier put: the joy of writing and the joy of reading are strongly linked. I don't write things with the expectation that nobody will read them, and I don't read things with the expectation that nobody wrote them. Or, in this case, a machine.
And I think a "writing photoshop" will be much harder to detect than an image Photoshop. Did I picture the author as a white collar Yale graduate because they are one, or because the machine told them that was best?
* For example: I would sooner give up my ability to walk than give up my ability to communicate myself to others; that's the difference I'm talking about.
True, but most people don’t write. ~Everybody reads, and people are, on the whole, supremely atrocious at it.
The problem of writing getting “easier” is trivial compared to the problem of bad reading comprehension.
I agree with you overall. It will likely have similar impact to grammarly and similar services.
And maybe that's not as bad as it seems right this moment. Once upon a time great painters actually made their own paints. These days we wouldn't think about that skill as in the necessary catalogue of an artist, and in a bunch of ways -- most ways -- we're the better for it. Perhaps something similar will unfold here.
I've never been upset about Stable Diffusion, because as you say, once upon a time, you had to physically paint an image to be visually creative. Never posted a comment about SD with "I cannot stand the idea of this. I know it's coming, I know I can't stop it, but it will be a catastrophic loss." Now, at the suggestion that it might happen for writing, all the sudden it seems wrong! Someone else noticed this in the other thread about Copilot: HN is not upset about Stable Diffusion, but seems to frequently be upset about Copilot! I think you may be right actually -- calm down, it's just a tool.
I guess I want to revise my prior comment. I don't just want to talk to a machine. Just like I don't just want to see Stable Diffused images. The human part is essential, it's the heart of the thing. My intuition is that you're supposed to augment the human part, not replace it.
Wow. Using the term inpaint (from AI art apps) illustrates your point with great clarity, at least for me.
Textshop, in-paint the rest with something nice, will ya.. oh and add a friendly sign off.
The authors of most shared articles and most comments are not even passing a “turing test”. In the vast majority of cases the readers just consume the data.
With GPT-3 we can already make “helpful and constructive” seeming comments that 9 out of 10 times may even be correct and normal. But 1 out of 10 times be kind of crappy. Aby organization with an agenda can start spinning up bots for Twitter channels, Telegram channels, HN usernames and so on, and amass karma, followers, members. In short, we are already past this point: https://xkcd.com/810/
And the scary thing is that, after they have amassed all this social capital, they can start moving the conversation in whatever directions the shadowy organization wants. The bots will be implacable and unconvinced by any arguments to the contrary… instead they can methodically gang up on their opponents and pit them agaisnt each other or get them deplatformed or marginalized, and through repetition these botnet swarms can get “exeedingly good at it”. Literally all human discussion — political, religious, philosophical etc. - could be subverted in this way. Just with bots trained on a corpus of existing text on the web.
In fact, the amount of content on the Internet written by humans could become vanishingly small by 2030, and the social capital — and soon, financial capital — of bots (and bot-owning organizations) will dwarf all the social capital and financial capital of humans. Services will no longer be able to tell the difference between the two, and even close-knit online societies like this one may start to prefer bots to humans, because they are impeccably well-behaved etc.
I am not saying we have to invent AGI or sexbots to do this. Nefarious organizations can already create sleeper bot accounts in all services, using GPT-4.
Imagine being systematically downvoted every time you post something against the bot swarm’s agenda. The bots can recognize if what you wrote is undermining their agenda, even if they do have a few false positives. They can also easily figure out your friends using network analysis and can gradually infiltrate your group and get you ostracized or get the group to disband. Because online, when no one knows if you’re a bot… the botswarms will be able to “beat everyone in the game” of conversation.
https://en.m.wikipedia.org/wiki/On_the_Internet,_nobody_know...
Alternatively you can lie and say 'this is what I saw' about something that is not even close. Once images are used to promote falsehoods, your worry becomes true. But much editing, either for accuracy or for beauty, has no such harmful effects.
I'm not sure when this practice started, and it may have gone by several different terms --- "sample business letters", "model business letters", and "handbook of business letters" are terms I'm coming up with.
I'd consider 1960 to present to be fairly recent but two books from that period are the McGraw-Hill Handbook of Business Letters (<https://www.worldcat.org/title/28181038>) and The Complete Book of Business Letters (<https://www.worldcat.org/title/1975991>).
I'm finding a similarly titled work or collection from 1885 though I'm not sure it's the same concept, Sample business correspondence, 1885:
<https://www.worldcat.org/title/831831066>
What all of these address, however, is the fact that most people aren't good at composing correspondence, and/or that businesses benefit by standardised forms of communications.
Photoshop lets a conscious hand change the medium. Now we have tools that generate the message out of thin air. We're about to outsource human expression to machines.
And why? To automate things we don't want to write about, and automate reading them. I picture a future of machines talking to each other in increasingly clever human language while humans on each end just get the straight talk: "Did you put the cover on the TPS reports?" "Yes, I did. I got the memo."
This also ushers in a terrifying era of communication. "I'm so sorry for your loss", says a machine on behalf of someone, using all the right words but not feeling any of them. "Thank you" replies another, on behalf of someone else.
The effect is the same on the reader. This looks hand-wavy to me.
Spelling and grammar checkers are amazing. It becomes a problem when we outsource expressing ourselves to a machine. If kind, thoughtful words come from an artificial brain, expressing sympathies just becomes a box-ticking exercise.
We mocked "press F to pay respects", but we've already built it in real life. "Congratulate John about his new job" (LinkedIn), "Wish Jane a happy birthday" (Facebook). Having the machine write the message for us is the next step, and I find it terrifying. What value is there to a gesture if it asks nothing of us? No thoughtfulness, hell, no thought.
At this point, you might as well go all-in. Attach a pen to a CNC, connect AI to it, and offer heartfelt handwritten notes as a service. Offer to connect it to social media and auto-detect heartfelt handwritten letter opportunities. Have that letter in the mail while the corpse is still warm.
Grammarly[2] also have a desktop app that appears to offer editing advice, although generally I think they're focused on grammatical correctness over anything else.
The question I'm left asking myself at the end of the article is, to what end do we need to edit text like Photoshop? Part of me sees this "Photoshop for Text" as something that would be akin to "No Code" tech stacks. Good No-Code/Low-Code solutions usually allow to build specific classes of products (websites, 3D assets) in ways that are faster than the status quo. But anyone who spends enough time in a No Code stack eventually hits the wall where the people who designed the tool had to sacrifice the flexibility of text for the convenience of a GUI.
I yearn for the day that we can set a language model loose on something like the NCBI database or arXiv and have it point out open problems in the field to new PhD students. Or have it figure out whether my ablation studies make sense. Or an AI that can generate math proofs for me. A lot of this linked to model interpretability and understanding, but I think the work that DeepMind is doing is showing that there might be a way to utilize this stuff in expert domains sooner than we think.
If I didn't have this experience, then giving the machine any input on what I write would seem crazy to me. I would think that language is too personal, too contextual, that I need control over every word and every letter.
But now I love writing with the help of the machine. It still feels like me speaking, the machine doesn't add any extra context that I don't approve of. It really feels like the messages are still mine, and the autocomplete just helps me extract my thoughts from my head in a better and more effective way.
1. (of body tissue or an organ) waste away, especially as a result of the degeneration of cells, or become vestigial during evolution.
"without exercise, the muscles will atrophy"I'm not a native speaker. It's nice to have training wheels sometimes, even for a language I'm familiar with.
It's sound.
Long before GPT, image synthesis, video deep-fakes and these imagined "Photoshop for words", we had sound synthesis.
That's a very useful marker. Because we can read the things people were saying about the future of sound, their hopes, fears and predictions as far back as the 1960s when Robert Moog and Wendy Carlos were patching modular synths.
Most of the fears and predictions turned out to be rubbish. Musicians, orchestras and live events didn't get replaced. Instead we invented synth-pop bands.
And many of the things technologists imagined people would want to do, turned out to be way off the mark. To my knowledge Isao Tomita was the only talented artist to "replace an orchestra" with synthesisers. Most people who used the tools "as intended" were artless, and forgettable. Everyone else ran riot in the parameter space - messing and subverting the technology to get the weirdest punk-ass squelches and wobbles possible.
So I always have to look on these "How synthetic X is going to make the real X obsolete" with a pinch of salt.
From your comment, it seems that the linked article is far more fear-mongering than it is; I gathered a mostly optimistic tone from it.
The final paragraph -
> While some of these capabilities sound a bit scary at first, they will eventually become as mundane as “desaturate”, “Gaussian blur” or any regular image filter, and unlock new creative potential.
I will also put it forward that for reasons I'm ignorant of, eye seems to be more readily fooled than ear. 20 years ago with crappy tools all I had to do was smudge and clone a hydrant in a photo and it would effectively be gone for 99% of observers. But similarly primitive ways of trying to change or alter a sound file were immediatelly noticed by all listeners.
Photoshop conceptually sits somewhere between these two extremes of electronically-aided creation, but much closer to sound synthesis, than what the author is hypothesizing. I can’t even think of a text analogy for sound synthesis as you’ve described. The least nonsensical imagined example I can think of is “this word does not exist” (as in word synthesis), which would be more valuable as a game or a gag than as a tool.
For core text processing that would be similar to how a bitmap editor processes text (fitering, replacement, conversion, etc.), there are some tools, like the aptly named:
# Photoshop has since added some AI stuff, like for object removal and such, but its main functionality wasn't and is still not about that).
Can’t say that what it generated was particularly insightful, however it was helpful to reach an obituary word count.
I'm pretty sure you meant "obligatory" word count. But "obituary" is an awesome accident. Like Copilot is waiting, happy to sum up your life in a tidy paragraph, when the time comes.
I rarely need more than the builtin macro system though, as macros can basically do anything including regex-based search it can do any formatting or change as long as the steps stay the same each iteration.
I find it necessary to heavily edit when writing - up to, and including, this comment. I don’t mind it, and I don’t mind doing it to other people’s writing either, so this new way of doing things appeals to me.
I’d be interested to hear what anyone who’s able one-shot their writing thinks of this. I feel like that type of person may have less of a desire for this kind of stuff?
Yep. I don't use spell-checking and I don't use auto-complete bars on phone keyboards, largely because I feel it keeps my skills sharp. I would use tools like the article describes when I deem them to have become necessary to stay competitive at what I do, but at the moment I don't feel it's clear-cut whether their use would promote or harm my faculties to write.
I also wonder what it means if everyone is farming out parts of their intellect to similar models, which might be limited pathologically, by training data or enforcement.
AI text generations will be more like games than like photoshop. Photoshop is 2 step: do and undo, but editing text is more like a continuous stream of refinements.
If by "machine generated fluff" thou dost refer to the banal, trite and insipid content that oft pollutes the Interwebs, then nay, I believe not that people desire to read such.
The "thought-processing" angle is cool though.
https://github.com/Hellisotherpeople/Constrained-Text-Genera...
A blurry image is less likely to get you into trouble than saying the wrong words, so it is critical to validate manually what it says.
That said, such tools could be useful for recommendations on how to rephrase things.
Text? I'm not sure what kind of model we have here. Only now that we have word embedding and other nlp models can we hope to do the same kind of thing we do with text as we do with images?
The surprise is not that it is coming, but that it hasn’t been a thing for years given how flammable the internet is.
"Photoshop for Text" is, arguably, called "Typesetting". Aldus Pagemaker and Quark XPress were both quite popular at this task.
Always the promise
Vim can do this, not in the sence you putting into the article but at least it does it without requiring using anything else except 3 rows of keyboard.
> Text filters will allow you to paraphrase text, so that you can switch easily between styles of prose: literary, technical, journalistic, legal, and more.
Pfff. Styles of prose are: trolling, documenting, cat-talking, legal and just making a list of something. Trolling cannot be augmented, documenting feature begs of some connections to reality, proper cat-talking requires throwing a lot of synonyms really fast, legal is kind of Java programming when you type one line and get 40, and augmenting lists might be done with either md-style but without requiring to draw every symbol of that ascii tables or excel-style but without gui.
So in what sense? Some other irrelevant sense?