Using Google Cloud Vision OCR to extract text from photos and scanned documents
gist.github.com
gist.github.com
https://chrome.google.com/webstore/detail/cloud-vision/nblmo...
Try it out, let me know what you think. File issues at github.com/GoogleCloudPlatform/cloud-vision/
Btw, here is a Ruby script that will take an API key and image URL and return the text:
I read about these limitations in the Cloud Vision OCR API docs, but could not believe that they would indeed not provide data at the word or region level. Anyone has any idea why?
I mean, they must have this data internally and it is key for useful OCR.
Currently I am using the free ocr api at https://ocr.space/OCRAPI for my projects. It also has a corresponding chrome extension called "Copyfish", https://github.com/A9T9/Copyfish
It's been used with different kinds of financial docs such as municipal bonds. Implemented in pure python, it has a web interface, simple API and does nifty type inference (dates, interest rate, dollar ammounts...).
- http://tabula.technology/ (Java)
- https://github.com/jsvine/pdfplumber (pure Python as well)
Similar to plumber and opposed to Tabula, the goal was to extract tables from a swath of documents without user intervention. Additionally, no knowledge about the location tables in the document is required. A fully automated workflow would curl -X POST localhost/analyze/... and filter down the json to the type or types of tables needed (via context lines, data types, column headers).
I hope it will be useful for you who want to try Vision API without being bothered to get the token API from Google Cloud.
We used to run this on videos.
https://gist.github.com/dannguyen/a0b69c84ebc00c54c94d#perfo...
Basically, about 2 seconds for the road signs photo. 6+ seconds for the spreadsheet image (with occasional timeouts). So, probably not optimized/ideal for reading large amounts of text
Compiles fine on Windows.
var request = require('request')
var file = require('fs').readFileSync('./testimage.png').toString('base64')
var body = {
requests: {
image: {
content: file
},
features: [
{
type: 'TEXT_DETECTION',
maxResults: 10
}
]
}
}
var url = 'https://vision.googleapis.com/v1/images:annotate\?key\=your_api_key'
request({
url: url,
method: 'POST',
body: JSON.stringify(body)
}, (err, res, body) => {
console.log(err)
console.log(body)
console.log(JSON.parse(body).responses[0].textAnnotations[0].description)
})
Basically you want to convert image data into base64, put it in the requests.image.content field and make a POST request and you'll get back the text.a simple service that has a free plan on top of it can be found here - https://scanr.xyz/
https://gist.github.com/dannguyen/a0b69c84ebc00c54c94d#1a-go...
(better than I thought, actually)