PySceneDetect – A tool for detecting scenes in movies
pyscenedetect.readthedocs.io
pyscenedetect.readthedocs.io
I remember that I used two techniques for extracting scene cuts:
* Difference in the brightness (Y) histogram of the YUV video between frames; when that difference is more than a threshold there's a scene cut
* Counting the number of Intra macroblocks per frame on an H.264 encoded video; if that number was more than a threshold then there's a scene cut
There are some more info in the wikipedia: https://en.wikipedia.org/wiki/Macroblock and also there's a nice article explaining some of the magics of H264 compression: https://sidbala.com/h-264-is-magic/
Other techniques being considered for future work include use of optical flow, background subtraction, and analyzing histograms.
It seems like the only way to do it is click edit, get an error message and then hit backspace in the url to get to the root of the repository.
Now that you mention it, I might ask the Readthedocs folks to see if they have any idea why this is happening.
Of course, the project is still very basic in it's current form, but does offer a good platform for testing more advanced scene detection methodologies (and has also been used as a baseline in various research/academic contexts). There's been plenty of suggestions for other detection methods (e.g. using histograms) which I'm very interested in looking into adding for future releases.
Lastly, one thing I do want to address going forwards is the performance, to make the algorithms more suitable for real-time systems (have decided to rewrite the core algorithms in C++ when I have some time, and interface that with the Python library/CLI). That being said, I'm always open to feedback or different ways of approaching things. If you or anyone else has any feedback or suggestions, feel free to create a new issue on the Github issue tracker [2] and I would be happy to discuss the possibility of including it in a future release of PySceneDetect.
[1]See e.g. docs of requests: https://requests.readthedocs.io/en/master/
I'm most open to any feedback or feature requests/ideas/suggestions; feel free to checkout the issue tracker on Github, or create a new entry:
https://github.com/Breakthrough/PySceneDetect/issues
Some ideas being considered/researched for future releases:
- looking at changes to image histograms
- using edge detection to improve robustness
- camera flash / foreground object suppression
- automatic threshold detection using statistical methods (currently is just a heuristic)
Eg. Let's say that you wanted to watch a prerecorded football game or baseball game without all the commercials, timeouts, commentators talking about the fans, etc.
Or... Let's say that you wanted to re-cut a movie in a certain way, by re-ordering the scenes, you could just generate a new scene data file and let the encoder/player use that.
[1] https://en.wikipedia.org/wiki/Edit_decision_list
[2] https://github.com/Breakthrough/PySceneDetect/issues/101
Pornhub et al often show scene markers, but I have assumed they're manual or extracted in the same way as DVD chapters.
Would you be able to share a small sub-set of the episode, in particular the area where you're unable to detect the starting segment? (If not, no worries!)
There's a few issues with PySceneDetect currently that may lead to false or missed detections, but these are things that I would like to solve in the long run:
- threshold is heuristic/fixed right now, but I would like to change it to an adaptive/statistical method which can dynamically change
- single-frame events can trigger false scene changes
Thanks for your feedback, and feel free to share any other suggestions you might have.
I'm trying to help a friend in media industry. His requirement is ability to identify different voice in a movie and have the output time stamped - Ex. Voice A: 00.00.00sec - 00.05.30sec, Voice B: 00.05.31 - 00.06.30, etc.
It would be very helpful if anyone can point to any tools that exists that can do that (open source or otherwise).
Transcribe seems to be more for speech-to-text irrespective to who is making the speech.
Here the requirement is to identify the unique voices. Ex. if "Mary had a little lamb" is voiced by two different voices then the engine should identify Voice A said "Mary" at 00.00.00sec-00.00.01sec and then Voice B said "had a little" at 0.00.02sec-00.00.03sec, then Voice A again said "lamb".. etc.
Recognize Multiple Speakers
Amazon Transcribe is able to recognize when the speaker changes and attribute the transcribed text appropriately. This can significantly reduce the amount of work needed to transcribe audio with multiple speakers like telephone calls, meetings, and television shows.
So they claim to be able to do it. However, having to do speaker diarization doesn't exactly make speech recognition easier, so you should adjust your expectations regarding the error rate accordingly.
But thanks; on doing little bit of search on 'speech diarization' saw that Google Cloud has that service. https://cloud.google.com/speech-to-text/docs/multiple-voices
And then came across this comparison between various cloud transcription service, which should be helpful for evaluating on diarization aspect as well.
AI-POWERED TRANSCRIPTION SERVICES SHOWDOWN: AWS VS. GOOGLE VS. IBM WATSON VS. NUANCE https://www.armedia.com/blog/transcription-services-aws-goog...
I would also recommend comparing it to Google’s and maybe MS/Azure’s services.
Aside from the fidelity of the transcription itself, and the accuracy of disambiguating the different speakers, I’m also not convinced that all of these services will give you an end timestamp (start timestamps are sometimes there, but not necessarily for every word or sentence).
You could try a multi prong approach by doing the transcript to get text and speakers only (using the services mentioned above), and then using an aligner such as “gentle” [0] to find the start / end times.
You will have some gaps and wrongly transcribed words, but it may be a start..!
Most of what you'll find would require downsampling your audio to 16khz and you'll find a combination of NN based diarizer and hmm based models pre-ML.
One thing to note a lot of the systems will work well for interviews, broadcast media and footage taken from a camera because the audio will tend to be clean.
Film and movies will be a challenge because of the background music being identified as a separate voice. It tends to to confuse it.
Haven't used the cloud based systems with audio the background sounds you tend to find in film movies.
I ended up using optical flow and key frame information instead.
Just curious, what didn't you like about the output quality? What version of PySceneDetect were you using?
The latest version (v0.5.x) uses ffmpeg instead of mkvmerge for output by default now, which produces significantly more accurate and higher quality output than before.
That being said, you are correct in that optical flow and keyframe information is currently not being used during the detection phase. There are several proposals to incorporate this into a future release, however, along with several other techniques: - histogram analysis - edge detection - background subtraction - automatic or dynamic threshold detection
Thanks for your feedback!
Optical flow for scene detection? How does this work?
This approach promises to detect scenes. I’m not enough of a coder to know if it is valid, but I can certainly imagine that it is a solvable problem. Most scenes are defined by location, and movies makers like each locatation to be distinct from the preceding one.
To my mind, sequences are where the action is at. Sequences are defined by narrative development... These can usually be defined by a trained human, but are probably tough to computationly define.
There is also a huge difference between film genres. The scenes and sequences of Finding Nemo are pretty easy to define. But you try the same approach on an art house film like The Scent of Green Papaya, and see how far you get.
https://www.helpingwritersbecomeauthors.com/movie-storystruc...
I've also seen people combine the cuts detected by PySceneDetect with subtitle (.srt) information to generate a more comprehensive list of scenes, which are sets of shots joined together by some context (to allow for jumping forward to a particular scene by it's description).
It would be kind of awesome if we could “hash tag” scenes such that at some point we could search fot “remember that #action scene in movie X where the [something happens]” and it can pull up that clip of the scene...
Disclaimer: I'm the author of PySceneDetect.
Or you could make some kind of generic feature vector from each frame that includes contrast, etc. The interiors are the same, but the camera focus/fov will likely be different.
Nowadays just take an annotated movie and throw it into an encoder/decoder network to assess if a frame is from the same scene?
The one area where PySceneDetect does struggle with currently, however, is dealing with sudden/rapid flashes, or momentary obstructions to the viewing scene.
However, there are scenes in the narrative sense, but also scene main the way an AD making a schedule will see them, that is, scene headings (sluglines), which would be easier to differentiate and detect.
I built it for auto-summarisation of video and categorisation as a few day hack and it was OK (better than I expected, not good enough to use). Totally failed at sports though, as the commentary usually has more varied speech when there's little happening.
(Disclaimer: I'm the author of PySceneDetect.)
For detecting fades-to-black, you want to make sure that you're using the detect-threshold command (not detect-content). For example:
scenedetect -i somevideo.mp4 detect-threshold -t 12 list-scenes
Where -t specifies the threshold to use (the default being a value of 12). Full documentation for the detect-threshold command:https://pyscenedetect-manual.readthedocs.io/en/latest/cli/de...