Show HN: Extract main ideas in your texts with SummarizeBot
summarizebot.com
summarizebot.com
I assume the expected reaction is "Cool! They use blockchain." My reaction is moreso, "What on earth does this have to do with blockchain?"
> We apply decentralized architecture to train and test our AI models. Using blockchain technology helps us not only to get more training data but also to improve the trustworthiness of our algorithms.
I still don't understand what this has to do with blockchain.
Funding from VCs?
1. Summarizing entire bodies of text down into "bite-sized" chunks isn't inherently a good thing. It seems the main use case (and at least the one suggested in the demo) is to be used for news articles. Now, I'm totally understanding of the fact that not everyone has the time to read every news article, but as it is, only reading part of the article (or more commonly, only reading the headline) is a huge issue with current consumption of content. This attempt to further summarize articles into small, context-less bites seems to be going in the wrong direction.
2. On the demo page, there is a "Fake News Detection" feature. I threw a couple of articles at it and it left me with so many questions I don't even know where to begin. For a few articles, it just gave me a binary "Real:1 , Fake:0" output. For others, it spit out a couple of numbers for stats like "conspiracy", "irony", "bias", "pseudoscience". Why are these the attributes chosen to measure? How are they calculated? Is something like "irony" even meaningful when trying to detect fake news?
Viewing the documentation section of the site, there is a small blurb claiming that it uses "custom AI classifiers", "custom machine learning models trained on fake and biased articles", and "database of trusted and biased websites created by our experts" to calculate these numbers. AKA, there is absolutely zero meaningful explanation as to how these numbers are calculated and why they should be trusted. This entire feature is a complete black box, and for all we know, the "database of trusted websites" could be created by Russian spies trying to sow misinformation.
1) I did try on https://www.reuters.com/article/us-southkorea-prisonstay-idU... and from 20% and 40% summary I have no clue of the actual meaning of the article. It seems that it just gets some sentences out, but they don't really combine in summary as a whole.
2) And seems like it has some fake news element as well "fake: 0.343". Confusing.
To sum up my experience: I confused, real: 0.9; fake: 0.1.
Anyhow it seems that has some great potential and may be useful in some general knowledge fact summarization in the future.
--
[0] - as long as it wasn't a cloud SaaS where I have to share my data with vendor's machines.
What exactly does this have to do with blockchain?
The landing page is more marketing than technical and may not really be a good fit for this site.
In theory that makes the language models auditable and tamper-proof. I'm not so sure about the supposed benefit of that, though. Yes, it means that the model itself cannot be tampered with (in order to introduce bias to the summaries, for instance) but as long as the algorithm itself remains closed source you could still alter the results by for example boosting some values while attributing less significance to others.
Simply publishing both the algorithm and the model as open source alongside with an SHA-2 hash to make sure neither has been tampered with would achieve a lot more in terms of reproducibility and trustworthiness.
Then again, they would've had one buzzword less in that case ...
When people say blockchain the meaning that there is a distributed consensus comes into the picture. In this case, there is no reason for a distributed consensus on ordering or anything.
But if you are suggesting there are many text parsers that train the model, and there is a central modal that's held by the network state, sure. But I don't know what's the benefit to that as I don't think simply training on more text will allow this bot to produce better summaries.
https://gizmodo.com/yahoo-shutters-that-30-million-app-it-bo...
https://www.summarizebot.com/api/378d4eec8d0e4ddeb8142c84433...
Say, someone scanning through a list of legal docs to identify the most relevant ones. A short blurb would be pretty helpful.
on edit: fixed misspelling, just woke up from nap.
https://github.com/contentinnovation/NeurIPS-2018-papers
It would be really cool to be able to translate via ML high level AI progress into standard American journalistic english. To some extent Bloomberg TicToc, Jinri Toutiao are already generating short form video for breaking news stories.
For instance, Ocean Protocol plans to use TCRs for data quality, meaning that data providers and data consumers evaluate the quality of a dataset in a continuous way, so that certain assets can be moved up/ down the ranks in near real time.