Simply by erasing stop words from a headline and then scanning the remaining terms for likely hits. The media has quite a few terms it likes to repetitively use in headlines for certain topics. I currently have about 100 terms to check a headline against. Think words like "environment", "bribe", "scandal", "accuse", "working conditions", etc...
It's... not a sophisticated system. It misses some major stories, gives me a few false positives, and of course lots of duplicates from differing sources.
But compared to what I had before (basically, nothing) it's a million times better.