Python API for Zero Day Phishing Detection Based on Computer Vision
github.com
github.com
That said, your browser plugin is extremely aggressive at uploading screenshots of potentially sensitive information to your servers (for instance, someone's stealth mode startup's unindexed Heroku staging URL), doing so by default if it hasn't been whitelisted [0]. And your privacy policy [1] permits you to resell "non-personally identifiable visitor information... to other parties for marketing, advertising, or other uses."
I don't blame you for needing to make money or for launching an MVP, but there must be better zero-knowledge ways of accomplishing this. It's an important goal and I wish you all the best... but don't be the source of the data leaks you intend to prevent.
[0] https://github.com/phishai/phish-protect/blob/master/js/back...
Per the browser extension, we don't upload every screenshot in the free extension only if you like to scan a particular webpage as we can't afford to scan every webpage for free. We don't sell any information as this is not our business model - I got your point with the privacy policy and we will make it more accurate.
The business model is mainly with API product where you will see more integration coming up in the next couple of months and as I said the use case for the API is more for incident response teams and hosting providers that need to go manually through a lot of URLs and tag them.
And then with the extension it takes a screenshot of every page you visit and uploads to their server that does something with CV fast enough that it can block you from submitting data to a phishing attempt.
Anybody have any details of how this actually works?
Seems a little magic.
1) we crawl websites of known brands and take screenshots of them. then we extract AI & computer vision features from them and create signatures for every website.
Now there are two products 1) API: you can use this for any type of use case: incident response, automated phishing classification for hosting providers. or any other use-case that you can think of. 2) chrome extension: The free version is not uploading screenshots in real-time as it's very expensive and resource consuming to process a lot of screenshots so it's available only for enterprise version. The free chrome extension has two features: unicode detection and link to scan the current website actively with Phish.AI - so really nothing to call home about.
Hope it made some things a bit clearer
The use case is more appropriate for Incident-response teams that go manually through tons of URLs to classify them or Hosting providers that go through tons of websites to check if phishing websites are hosted or some sites were hacked. Essentially any use case where you have a feed of urls that you can access from the web and tag/find phishing websites
'electronic communication' is generally considered a synonym for 'e-mail'. If that's not what your service is targeting, you should probably explain it.
We specialize in comparing visually website to our own database of legit websites and then comparing the domain to detect fake/phishing websites that hosted on different domains. For example, we detect that website looks like paypal and then we check if it's hosted on one of paypal domain (paypal.com,paypal.uk etc...)
New and hot: Breach at phish.ai‘s screenshot database leaks all my company‘s emails.
https://github.com/phishai/phish-ai-api/blob/master/phish_ai...
So, yes, you are likely both in agreement.
Instead of what content often shows up on HN with lots of technical detail, this amounts to a service announcement.
This service sounds interesting but a much better HN submission would be one that talks about how it works, how it was improved, etc. Browsing the phish.ai website uncovers several (very brief) articles that are more interesting than this repo.
BTW footnote wrt the API's design -- {"verdict": "clean"} -- this is not ideal for a non-interactive service. It should either return a numeral (like a confidence interval) or a boolean value.