Using Deep Learning to Create Professional-Level Photographs
research.googleblog.com
research.googleblog.com
The method proposed in the paper(https://arxiv.org/abs/1707.03491) is mimicing a photographer's work: From taking the picture(image composition) to post-processing(traditional filter like HDR, Saturation. But also GAN powered local brightness editing).In the end it also picks the best photos(Aesthetic ranking)
Selected comments from professional photographers at the end of paper is very informative. There's also a showcase of model created photos in http://google.github.io/creatism
[Disclaimer: I'm the second author of the paper]
I think you're trying to say professional photographers don't apply tone mapping to single exposures, but that is false. There are times when it's not 'technically proper' to apply tone mapping to a single exposure, but it makes for a better picture.
Really it sounds like you're trying to make a snarky comment referencing how some professional photographers complain about novices over using HDR on flickr. But who cares. You can have a professional quality photograph that uses what some pros think is incorrect tone mapping. Just like how you can have a novice hack together a working professional-level product that many professional engineers think is trash. If it works and meets the criteria, who cares.
14 stops of dynamic range in a single shot (RAW on a modern DSLR) counts in my book.
My point is that with 2005-era DSLRs, I had to use HDR techniques (or strobes) to, say, correctly expose a room and the environment out a window. Shooting into the sun, I'd have to use HDR, polarizers, and/or ND grads to expose both the sky and the foreground. With my D750? I've got 4 stops more dynamic range and the situations where I need HDR have almost all gone away. I can get the same thing from a single RAW exposure.
These days, HDR only seems necessary if you want to have ridiculous tone mapping where you pull in your global highlights to be within a stop or two of your global shadows and still keep the noise levels down.
But to your point... it seems that "HDR" means two different things these days.
1. A tool to expand the dynamic range.
2. A style of tone mapping.
These two things frequently coincide, but there's no reason for #1 to have crazy surreal tones, and #2 frequently comes from a single exposure.
On the con side: only an AI photo-bot can find and capture the most photogenic scenes in the world.
I'm curious: Is this creatism link what the algorithm considers the best images, or did you manually filter out some of the less good ones?
Here's the thing. If she had say, 2 shoots in a week, maybe a total of say 8 hours of shooting time, the selection, editing, and styling of those photos could take 3x or 4x as long as the time spent shooting. She would have to be editing constantly.
I've always thought a perfect application for this sort of technology would be a model trained on that particular photographer's style, which goes through a batch of photos and selects the best candidates and then presents the user with a few choices for each photo it selected and styled.
I think the ultimate would be a NAS with this capability embedded. The photographers I have met through her over the years seem savvy about photo tech, but don't generally use online storage solutions or know much more than how to admin a WP site at best. SOHO solution would be ideal imo.
Great work! I'm excited to see where this goes.
It's more likely this will hit consumers hands via an automatically applied 'enhancement' in Google Photos first.
Storage is always a problem with every photographer I have met. They keep stacks of external USB drives or burn archives to DVD. They don't trust or want to use online storage.
A package that provided these models, trainable on that photographer's particular style, with autocropping, watermarking, etc. would sell very well among this group.
At these small numbers, they are more likely to just give away access to the software for free.
If you had a system to:
- batch ingest new photos
- stylize them to your own style
- categorize them
- autoselect candidate photos for final tweaking
- back up final album photos to local storage and S3
- archive to Glacier the rest of the photos from the shoot that will likely never be touched again
You would have a product that I think would be very attractive to professional and semi-professional photographers.
Was there any control for, e.g., the dynamic lighting filter being the factor that carried most of the water, rather than the other manipulations?
FWIW, I did think your tool did a good job finding good compositions within mundane scenes. Looking at the originals, I have to admit that in many or even most of the cases, I wouldn't have spotted the opportunity - which points to an area I can improve my craft.
Photography, on the other hand, is a very common hobby in the tech community. And the comments here seem to reflect that this effort strikes a little close to home: “Those pictures are lousy, if you find them appealing you have no taste! Just because they're 'professional' doesn't mean they're good! Machines can’t replace human judgment, they have no soul! I bet that machine had a lot of human help!”
Tech people may tell you great stories about meritocracy and reason, but in the end we are just emotional monkeys. Like the rest of humanity.
Those of us who can accept this may at least aspire to be wise monkeys.
But you know what? I love this shit. I like photography for the good photos, not because I'm building my self esteem on top of it. I want to capture the scene or the moment in a way that, IMO, does it justice. If this technology makes that easier, and gives it to more people - great. Progress.
And what I really love is how it subtracts a huge amount from the price of entry. It's like when software synths came out and made it possible to make music without $10k+ worth of MIDI hardware. What followed? An explosion of creativity which I've been luxuriating in ever since. Get the ability to create beautiful images in as many hands as possible, says I. For me, at least, "more beautiful images in the world" was the whole idea.
There are still many situations where more professional expertise is required (portrait photography is probably one of the best examples, knowledge of proper lighting is still required for best results). But I've gotten the impression that overall demand for photographers -- and pay -- is slower these days, largely due to technological advances. (See articles like this: http://www.nytimes.com/2010/03/30/business/media/30photogs.h...)
In a way, this is fine though -- as you say, it gives power to more people, which is great.
Per the parent post's point, to be honest, I perceived much of the criticism to be more that they really couldn't think of a good way to use it in their own hobby level work. I personally can't think of a good use case. On the other hand, the most useful application I can think of for this is more in reverse: social media companies with huge image libraries their users have graciously "donated" to them, could use this AI algorithm or something similar to pick, or construct, the most "professional" looking image of anything they can identify from their vast library, and sell it or license it to media / content publishers for a relatively low cost. Such could be but one more competitor to any professional landscape / scenery photographers out there.
This doesn't mean that someone armed with their phone will take good photos. There's still the matter of composition, as well as capturing the "the decisive moment"[1].
The tool described in the OP provides an automated means of finding some worthwhile composition within a library of images, thus providing an aid to someone less skilled at composition - although it's not going to produce something from nothing. If the photographer failed to point the camera in the direction of the best photo, then they've failed.
But I can't see any way that automated tools can, after the fact, help us to capture that decisive moment.
[1] http://truecenterpublishing.com/photopsy/decisive_moment.htm
Truly good photographers don't just produce beautiful photos; they produce meaningful photos.
The tech is great but I'm not a fan of the title of the article ;)
Photojournalism is almost always about the story and what the picture evokes. It could be a simple pic of a boy covered in dust, but made evocative because the dust is from a bomb explosion that killed his father and not from playing in the park.
But in case of a lot of landscape photography, the content is intrinsic to the picture itself—the photo does not get its value from anything external to it. Likewise for portraits.
Great landscape photography takes you to an interesting location or interesting angle. A post-processing algorithm cannot take you somewhere interesting.
Good landscape photography makes you say, "Wow, that's beautiful!"
Great landscape photography makes you say, "Holy crap, what is that? Where is that?"
https://500px.com/photo/4505244/rainbow-glass-by-junya-haseg...
https://1x.com/photo/41156/portfolio/60972
https://500px.com/photo/103746295/the-white-silence-by-danie...
http://ngm.nationalgeographic.com/2008/08/photo-contest/stei...
http://hongwrong.com/verical-horizon-hk/
https://500px.com/photo/103960697/one-thousand-and-one-night...
Buildings are boring as is snow, photo of person or animal / shadows not landscape, and over edited. Now, I suspect as someone that's looking at a lot of photos those pinged some novelty feeling in you, but that does not mean they are objectively good.
By comparison the article's photos (and those linked: http://google.github.io/creatism) generally evoked that I wish I was there feeling.
I predict a similar divide will crop up in computer generated music: very soon normal people will prefer their pleasant sounds, even while more discerning (semi-) pros will deride their lack of artistic merit for a lot longer.
Computer generated art will shine as Gebrauchskunst first---eg get a 'professional level' soundtrack and video editing for your youtube video shot on a phone.
This isn't just a post-processing algorithm. It also picks the place from Google Street View.
I struggled for years to make good photographs until I learned to post process well enough. For an average viewer, the difference is one of "side glance and ignore" and "wow, great photo".
It's instructive to look at photographers who post their originals. Even ones that appear to show an "interesting angle" can be quite distorted to show just that angle (as in if you travel there, you will never see the nice angles you saw in the photo).
I didn't say that you can't have great landscape photography or that there can't be mediocre photography.
I was talking about how there typically isn't an 'external' story to landscape photography the way there is for photojournalism.
Any photo is a tiny square cut out of an infinite number of 360 degree spheres taken at a single point in time. For every photograph there will always be much more outside of the frame than in it.
That doesn't even begin to touch on things like lighting for portraits - but to think that even a landscape photo is anything but subjective expression is naive at best, dangerously so at worst.
I suppose one could get into deep philosophical discussions about if a machine intelligence can make art, if it does so, does it have to be regarded as a higher level of intelligence, or does the art have to be regarded as inferior art, etc. Interesting stuff, but not my point at this time.
I personally enjoy learning how to use tools so I like messing with processing techniques. Since I'm very much an amateur, it's an effective way to improve the objective quality of an image. Still, it has made me very aware of the limitations when it comes to improving the subjective quality of my photos.
Essentially, I can save an underexposed RAW photo or use levels, curves, and masks to improve the quality and balance of lighting and color...but it will never fix a poorly composed shot or help me to better understand how to get a point across.
As someone who's also spent a fair amount on gear, the post-processing part is the thing I've had most issues (read: spent least time with) to get right, and while I haven't tried all software, I haven't found a one-click solution of sufficient quality.
I've been hoping for a while that ML might lead to something like this. I would like if I could spend all that time taking photos, and no time fiddling with them afterwards.
Could you recommend any "enhance" feature that you've encountered that you think is okay?
* Choose good scenes AND make them pretty AND paint well.
With photography, to make a good picture, you have to be able to:
* Choose good scenes AND make them pretty.
With this technology, to make a good picture, you might have to be able to:
* Choose good scenes.
Reducing the number of skills required massively increases the number of people who can do them all.
First, semi-pro means you've made money through photographs - not spent money.
Second, I heavily dispute the utility of those "Enhance" filters. They often do make a photo nicer, but they never make great photos.
It's interesting to see algorithms catching up to being able to replicate this. However when you mention these kind of abilities to photographers, they get defensive, almost like you are threatening their identity by saying a computer can do it.
It's a fairly common reaction from most people. Hayao Miyazaki, director of Spirited Away, got very upset after being shown a demo of AI doing animation:
https://qz.com/859454/the-director-of-spirited-away-says-ani...
His reaction of disgust isn't to the general idea of computer-generated animation, but to the specific animation he's being shown, which features a corpse thrashing around on the ground, and specifically in relation to having a severely disabled friend.
It's entirely possible he's not a fan of computer-generated animation in general, but that clip doesn't indicate much one way or another.
On the other hand, imaged thoughtfully and without enhancement, the mundane can be truly beautiful.
Subjectively speaking, the world always seems more saturated through my eyes. Objectively speaking, the eye has something like 20 stops of dynamic range while paper has 9 at best...
1. The diagonal lines in the clouds and the bright tree trunk at the extreme right of the first image are distractions that don't support the general aesthetic.
2. The bright linear object impinging on the right edge of the cow image and the bright patch of the partial face of the mountain on the extreme left. Probably the gravel at the left too since it does not really support the central theme.
3. The big black lump that obscures the 'corner' where the midground mountain meets the ground plane in the house image.
4. The minimal snow on the peaks in the snow capped mountain image is more documenting a crime scene than creating interest. I mean technically, yes there is snow and the claim that there was snow would probably stand up in a court of law, but it's not very interesting snow.
For me, it's the attention to detail that separates better than average snapshots from professional art. Or to put it another way, these are not the grade of images that a professional photographer would put in their portfolio. Even if they would get lots of likes on Facebook.
Again, it's an interesting project and a significant accomplishment. I just don't think the criteria by which images are being judged professional are adequate.
Edit: I guess it depends if there were other images in their corpus that didn't have any flaws, but who's to know?
Or to put it another way, compare the images in the blog post to Carter Gowl's http://www.gowlphoto.com. Gowl produces a couple of images a year and those are good enough to charge a couple of hundred dollars per print. Even if the images in the blog post were appropriately high resolution, I don't see them commanding a similar price. YMMV.
To be fair, there's plenty of technically good but uninteresting professional photography, too.
I hope that one day our driverless cars will alert us when there is a pretty view (or a rainbow) so we take a moment to look up from our phones. Every route can be a scenic route if you have an artistic eye.
Interesting book that centers around some of these themes. Worth a read if you haven't yet.
https://2.bp.blogspot.com/-6bVWUgA8NEI/WWe1uoW8ayI/AAAAAAAAB...
The model has the reverse situation, of course: it cannot perfectly guess the emotional response for any one person, but it has access to a larger assortment of data.
In addition, in different contexts it may be easier/cheaper to place a machine vs. a human in a certain locale to get a picture.
If my theorizing makes any sense, it suggests that this technology would be useful in contexts where: the locale is hard to reach and the topic is likely to evoke a wide variety of emotional responses.
So what? Maybe I missed it, but what are some potentially meaningful applications of this technology? What motivated this to begin with? Or are these questions that we even bother asking anymore?
I remember the first time someone showed me the Snapchat app -- it would make them look like a cartoon dog, or all these other real-time overlays. I thought, 'jesus, so glad we're all getting advanced computer science degrees so we can work on utterly useless shit like this...'
I think this research is spot on, and can't wait to have it on my phone. And I love taking photos the old fashioned way, too.
Saw a few people talking about retouching and studio work - I do a lot of studio shoots and retouching on my own, and would be happy to help or participate in projects. Feel free to reach out.
Oh really.
Unless someone can put a huge battery in a small mobile, forget about running big (and good) networks in mobile.