Show HN: TidalWaves API – live, tokenized news metadata from around the world
tidalwaves.io
tidalwaves.io
So, I built an API around it and released it on Rapid API today.
I'm hoping my pricing isn't too high.. It's currently set so that a Tier 1 sub can create the marketing site I linked, and Tier 2 subs can access historical data as well. But, I also hope to create a steady income so I can work on other projects in the future.
Hope you guys like it, and please send your comments and crits.
What are you using instead of GDELT now?
Our domain was pretty well defined, and we weren't building a real-time system, so we developed an adjudication tool and hired a few part time employees to evaluate and clean up our output.
More info: https://blog.gdeltproject.org/gdelt-translingual-translating...
(unrelated: scrolling on that page is horrible)
I think GDELT is the best source, nothing else comes close to coverage in my opinion, it's just that all the articles it gathers revolve around world events, as opposed to all of the millions of topics of discussion and niche interests people post about.
Asking on behalf of Rhys Darby. https://www.youtube.com/watch?v=HynsTvRVLiI
But then, Amazon's awash with tat too these days, and I wouldn't be surprised at anyone selling their product there.
I see it as a good way to test the waters. If people are interested in it, I'll eventually move it off of Rapid API for direct access and use Stripe for subscriptions.
https://autocode.com/lib/url/temporary/
We're ostensibly an automation platform but our core technology around Connector APIs automatically generates docs for you and a whole ton of other neat things - like your API can show up in autocompletion dropdowns for Autocode users. We're a Stripe-backed company, if you haven't heard of us it's because we rebranded / relaunched on July of this year. Disclaimer; am founder. :)
- Missing keywords. Huge topics like "brexit" won't be found, so you need to extract those yourself.
- Dirty titles. You need to extract the article titles from the website metadata, and more often than not they'll pre/append their site name like "[title] - DailyStormer" and links like "[title]: Business, Stocks, News" etc. That requires NLP to fix.
- Non-standardized locations. Location tags are all over the place, so "united_states" might have many tags like "us", "usa", "america", "the_states", etc. I'm still working on combining these tags in the API.
https://www.splcenter.org/fighting-hate/extremist-files/indi...
https://dataverse.harvard.edu/dataset.xhtml?persistentId=doi...
https://dataverse.harvard.edu/dataset.xhtml?persistentId=doi...
I understand your data comes from GDELT. I'm new to that and all I could glean from visiting the GDELT site at first go was that they have access to a lot of historical data sources - but I couldn't see a list of 1000+ current news publications. Is there a list somewhere? Are all the sources typographical, or might some be radio/tv? Thank you.
There is no official list currently, but you could query either GDELT or TidalWaves to see which are available in the data sets.
I could add a page to the marketing website which displays all of them, instead of the top sources for the latest news - would that be something useful for people?
I would like to see the list of publications if that is possible.
I am wondering what "trending stories" means in TidalWaves? For example, if the NYT at any point contains 500 stories, and is refreshed say 3 times every 24 hours, with say 100 stories aged off and 100 new stories added each addition, what is a trending story in the NYT? Does the front page headline count more than a page 27 minor traffic report?
How is the top 5 trending stories for USA derived?
As you can see, there are already almost 8000 sources. I'm honestly not sure how GDELT scrapes its data though, but this is what they do for over 5 years now and it's backed by Google Cloud processing, so I'm sure it's ever expanding and very thorough.
The kinds of articles which are discovered revolve around world events, however. So you won't see random blog posts, or traffic reports. Instead you'll see geopolitics, big tech news, and cultural discourse.
Trending stories are found for the entire set of articles in a batch, with no separation for source, category, or location. If the top stories all happen to be in the USA in a given moment, then that's all you'll find. If NYT happens to publish an article that fits an overall story being talked about in the rest of the batch, it'll get added to that story.
Trending popularity is decided by how many individual sources talk about the given topic. This is good enough to identify the Zeitgeist of a moment, but it likely won't catch nuances in evolving stories.
Some GDELT methodology is extolled here:
https://blog.gdeltproject.org/gdelt-2-0-our-global-world-in-...
Which ISO standards is this statement referring to?
I've never heard of GDELT before, but I've found news websites to have incredibly tight ToS's which prevent you from keeping a database of their articles offline, regardless of whether or not you distribute the articles.
Your map looks fantastic, btw.
The problem with GDELT is that they don't have that much coverage. I discovered it when we began to build our own News API [1].
I failed to sell this data over an API. I think you should see much more interest for an application. API like that is a bit complex. Almost all potential clients who really needed GDELT already parsing this data from Big Query, Redshift, or GDELT's dump files.
Anyway, check our solution, probably we could collaborate somehow. Map is cool! Feel free to reach me over artem [at] newscatcherapi [dot] com
Oh nice, I actually saw Newscatcher on Rapid API but didn't see the marketing page. It looks like a great product, and I think suited for large-scale solutions. TidalWaves will likely find a niche for people making visualizations and personal apps.
What I want to do is make a mature pipeline that handles at least 90% of what amateur and solo users would want to use GDELT for, so hopefully people see the value in that. After all, I could make this visualization using the API - so I'm sure people can think of many other cool uses for it.
Also, thanks for reaching out, I'd definitely like to collaborate in the coming months.
Not too interested in the most popular topics, but the obscure stuff like why some dept in govt has changed something related to topic and there are a couple journos who have been monitoring/reporting on that dept and so over time have a high count related to that topic.
They also could be press releases, considering the same title was found for multiple articles over the past 30-45 minutes. The data pipeline currently removes duplicate titles, but only inside single batches. So, it could be a long sustained press release.. like propaganda.
As of writing it shows three different stories though:
- 12:15 Will there be a coronavirus relief bill with $1,200 checks?
- 12:30 Nigeria protesters break curfew amid gunfire, chaos in Lagos
- 12:30 Google antitrust case to turn on how search engine grew dominant
Maybe what you're seeing is a bug in the website client, I'll investigate it.
Finding duplicate stories that have different texts and titles sounds like something you should use AI for
The locations are ordered by salience (proximity to start of the article, number of times mentioned), but not completely filtered, so on this specific visualization you will see a lot of matched locations which you wouldn't consider part of the article, but were mentioned in passing.
Trying to understand the content of an article is still really, really difficult, and error prone.
These scores are averaged for each spot on the map, showing the associated color.