Welcome to the New AWS AI Blog
aws.amazon.com
aws.amazon.com
I feel like Amazon could do some work in this area to support users to use their own engines and not be bound to AWS AI Platforms and Services.
This could be as simple as publicly documenting the time lambda's stay 'warm' for and retain data in /tmp persisting through multiple invocations or some other examples on how an AI workflow could be implemented with popular custom engines such as TensorFlow.
Does anybody else have any experience in this regard?
- Run strip all .so libraries -- many aren't stripped fully
- In Python I manually deleted sub packages of numpy/scipy I didn't need
- If you're loading large models at initialize, numpy load routines are _much_ faster than cPickle. Have it load at module initialization, not during each invocation.
I should really write a blog post about my experience with it.
At a certain point I decided I was doing something that lambda really wasn't designed for -- I'm looking at migrating off, but the current implementation makes capacity planning super easy. Provisioning 1000 machines with 1GB of RAM for 15 minutes every day to read off a queue isn't a trivial problem.
(Also, if anyone from AWS is reading, being able to limit the max concurrency of a single function vs account level limits would be super useful).
Would love to see a blog post to compare experiences..
(Speaking of blogposts.. that was somewhere on the todo list..)
Have you considered extracting the assessment functionality into an SWF activity, invoked from your Lambda via StepFunctions?
We considered running it on EC2, but the economics just didn't work out for our needs (hundreds of parallel processed jobs with irregular spikes, <5s runtimes per invocation and some others)
I'm curious what sort of things you did to shrink your model?
Have you considered pulling the model data from S3 outside of the main lambda handler and seeing if that negatively impacts performance -- with a cloud watch event running every 5 minutes or so to keep the function warm?
Hadn't thought about a CloudWatch even to keep the function warm, I might suggest that to the team.
We haven't shrunk the model, we've deleted superfluous files from the TensorFlow python library and dependencies (we don't need tensorboard, for example).
It would be nice if you could package TensorFlow up into a minimal component just for assessment and not have any of the 'learning' stuff or other added-on libraries, but we couldn't find a simple way of doing that - we're not pro C++ engineers and even our python is not the greatest. We're managing for now, but if our model grows any bigger we will run into issues, but there have been some good suggestions in here.
We had considered the S3 store, but we ruled that out quite quickly for cost & performance at the number of invocations a months we're looking at - but that was before we knew more about the 'keeping warm', so that may be revisited, too.
For example in one Lambda function, we use Guice to get resources from S3.
Disclaimer: work @ algorithmia
https://aws.amazon.com/blogs/compute/seamlessly-scale-predic...
Both providers offer you raw VMs with GPUs and such so you can run popular machine learning frameworks yourself by hand. After that the three providers diverge a bit, and I've not seen a good writeup myself. Roughly:
- Google has both a hosted TensorFlow (Cloud ML) as well as specific, pre-trained models you can simply use (Cloud Vision, Cloud Speech, etc.). For an easy to use interface, we have direct TensorFlow (and more) integration in Datalab.
- AWS also has some pre-trained services (Rekognition, Polly, Lex) but for "obvious" reasons doesn't do hosted TensorFlow. Instead Amazon Machine Learning is a bit more like Azure's offering: "Put data in, wire up stuff in the console and hit go".
If you're really interested in ML, my biased opinion is that you'll be using TensorFlow. And as you surmised, we're committed to making TensorFlow the "best" ML framework and making sure it runs well on Google Cloud. Like Kubernetes, we're not going to handicap it elsewhere, but having it managed and accelerated for you, is extremely convenient.
[Edit for formatting. I also should have mentioned there will be lots of ML-related talks at NEXT in San Francisco in two weeks!].
We provide a machine image with TF, MXNet and others pre-installed, along with Keras, CPU and NVIDIA divers, and other libraries for deep learning. We just added Ubuntu support too:
https://aws.amazon.com/blogs/ai/the-aws-deep-learning-ami-no...
http://www.allthingsdistributed.com/2016/11/mxnet-default-fr...
but to each his own!
Are there (artificial) datasets that can be used that showcase particular fit for deep learning?
I think there is way to little research in building artificial datasets (using domain knowledge of course).
It might even be possible to run these generative models and have this type of data very soon.
I agree though, the kind of artificial / play-against-yourself datasets that the folks at DeepMind created for say Alpha Go are an entirely different beast.
[1] https://cloud.google.com/blog/big-data/2017/02/google-cloud-...
[2] https://research.google.com/research-outreach.html#/research...
I guess it's more comparable with Google Cloud Functions: https://cloud.google.com/functions/
For more info, see our new blogpost (discuss here: https://news.ycombinator.com/item?id=13697666).
Again, important disclosure: I work on Google Cloud.
Polly is a stand alone component but the reverse is closely bound up into Lex which is a conversational interface API.
Amazon has internally built an engine I could ask to convert an audio file in S3 into a text content representative of the audio file... yet I can only use Lex to drive a conversation via text and audio.
If AWS really want to give me the power of their AI tools. How about unbundling them?
A big win for Amazon is that so many companies already have huge data sets in S3. Having AI APIs 'close to' existing data makes it easier getting started.
I have a million scanned images of court documents. Some are briefs, some are motions, some are court orders, etc... Given that I have images and their types, could I "train" the AI with these million documents to recognize a new image that might come in?
> Given that I have images and their types, could I "train" the AI with these million documents to recognize a new image that might come in?
Do I understand this right:
1. You have lots of documents as images, and their type (brief, motion, court order)
2. You get a new document, as an image.
3. You want to assign a type to this new document.
Is that right?
What's your goal in terms of quality? A key way of thinking about this is:
1. What's the risk/cost if you mis-classify a document?
2. What's the risk/cost if you fail to classify a document?
Are the documents typically very structurally different? Could I probably tell them apart without wearing glasses? Or are they largely the same, but with nuanced differences in the text?
That is exactly right.
As far as quality, if I classify the document wrong today (or fail to classify it), it's not that big of a deal to the system as a whole - but users will be "very" annoyed and have to correct it.
> Are the documents typically very structurally different? Could I probably tell them apart without wearing glasses? Or are they largely the same, but with nuanced differences in the text?
The documents are structurally reasonably different, but not "very" different. If a trained human looks at it, they would be able to tell them apart. There are exceptions, of course. For instance, if the attorney doesn't follow accepted convention, but those are rare.
While there are many benefits of end-to-end training, I don't think it might be best suited for this case. This is because we already know that the only useful feature in the document is the text, and not the contours or textures. Theefore wasting neurons in your neural network to learn the wastefulness of this features is just waste of resources. Furthermore, you benefit from even a larger corpus of learned data, which the char recognition engine has been trained on.
Thus, I was looking for a more advanced approach.
OCR is a better understood problem than a general neural net, so I think it's likely easier to improve its quality that to superseded the quality with image-based recognition.
With OCR, you get bits and pieces of information, but because I don't know what the type of the document is, it is difficult to determine where, structurally speaking, this information resides on a page.
If I could use AI to determine the type of the doc, I would know the structure of the document and I could then use OCR to pinpoint specific information on the page.
Most court documents are created from templates.
Leaving the quality part aside -- this job itself is easy to parallelize in that you can split it up by document or by page.
Open option is to run each job in Lambda asynchronously, with the input being a URL to the page or the full document, and have the job call back to you with the text of the page (or put it on S3 as a text file, or add it to a message queue, or whatever works). Regarding splitting: we've been using a python wrapper + pdfium for splitting PDFs into page images on Lambda, with excellent results.
To make the Lambda function, you'll either have to build e.g. Tesseract such that it fits into a 50MB zip, or download it while the Lambda function executes. LambCI has a set of docker containers that they've made for simulating lambda, and the "lambda:build" container makes building things easy and repeatable: https://github.com/lambci/docker-lambda. In a pinch, you can build on an Amazon Linux EC2 instance and it should work on Lambda, but you will have to be more careful about dynamic linking.
As another option: I'm not sure if it's been mentioned, but you can also try a ready-made OCR service before packaging up Tesseract, like this one: https://algorithmia.com/algorithms/ocr/SmartOCR.
So anyway, the performance part has good solutions, at least.
For fixing the accuracy: I know next to nothing about approximate string matching, but perhaps it would then be possible to do a fuzzy search over the text using something similar a Levenshtein automaton: https://en.wikipedia.org/wiki/Levenshtein_automaton.
You may also want to take a look at this: https://en.wikipedia.org/wiki/Bag-of-words_model
More broadly, I'm sure that there are text-based document classification methods that are robust against sloppy OCR. It may just take some research on the main approaches people take to document classification -- it's not my area, but my understanding is that this is typically approached with statistical methods. Otherwise your spam filter would get defeated by typos.
Disclaimer: pretty new to ML
But I know next to nothing about AI and ML - that's why I was asking this question.
Often ML problems way more experimental than "normal" coding. I would first try modelling it as a classification-problem and just do some cross-entropy validation to check the performance of the model. If it's useful, go with it, if not back to the drawing board. You will need some serious computing power, so either buy some GPUs or use the cloud.
You could train a random forest based on the inputs of OCR and NN, if you want to get total ML. You would gain some interpretability (i don't know whether thats important, but i would guess it might)
I am sorry that I can't give you a more concrete answer, these are just ideas. They are probably wrong. Like i said, i am a beginner and also don't really know the problem.
Edit other idea: If you know the location of the relevant OCR-text, you could use the following approach: Use the NN for classification. It will return probability-like values for every category. Take the top 2 (or 3, or every top until they add up to 70 percent...idk). Then do some OCR for every category you have to check. If one is positive you have your result, if not run the others.
I've had success with Clarifai's [0] custom CV model API in the past. You basically upload batches of labeled images to train a model, and then you can submit new images for classification.
Of course, I have no idea how effective it would be for your documents. Obviously it depends on how visually distinct the different types are.
From https://developer.clarifai.com/quick-start/
Seems simple to train
// add inputs with concepts
app.inputs.create([{
"url": "https://samples.clarifai.com/dog1.jpeg",
"concepts": [
{ "id": "cat", "value": false },
{ "id": "dog", "value": true }
]
}, {
"url": "https://samples.clarifai.com/dog2.jpeg",
"concepts": [
{ "id": "cat", "value": false },
{ "id": "dog", "value": true }
]
Then you predict what the another image is: // predict the contents of an image by passing in a url
app.models.predict(Clarifai.GENERAL_MODEL,
'https://samples.clarifai.com/metro-north.jpg').then(
function(response) {
console.log(response);
},
function(err) {
console.error(err);
}
);Is that possible with AWS AI?
> Finally, we provide AI engines, a collection of open-source, deep learning frameworks for academics and data scientists who want to build cutting edge, sophisticated intelligent systems, pre-installed configured on a convenient machine image.
Which is to say anything that TensorFlow can do, AWS AI can do. So yes. But is it possible without implementing it yourself? It looks like Amazon Rekognition may be able to do what you're looking for, but I'm not certain. You'll have to research that one.
If not AWS, how could I get started elsewhere? Every tutorial I've looked at deals with a CVS data file, not an image.
Could you point me to a resource that has an example of how to train AI with image documents?
You can make it more performant by downsizing the documents, but obviously this is incredibly GPU-intensive to train. I second the suggestions to just use OCR.
Look at the text API which parses text from an image and gives you it's pixel location.
More than happy to chat with you
1) Training ML models does not require network access, which is one of the biggest competitive advantages of the cloud.
2) Training ML models is typically a batch process, which benefits minimally from the scale-on-demand model of the cloud.
Since the cloud premium is a significant exchange for the value that it adds, I don't see this being a big win for cloud providers. I can't help but think that if I were making use of extensive machine learning with continuous training, I'd have it training models on a local bare metal cluster statically scaled to my application's demand with minimal network connectivity needs. And then ship the serialized trained models to the cloud. The potential cost difference is huge.
> benefits minimally from scale-on-demand
Neither of these is true when dealing with terabytes of data (or more, if you're working with image/video corpora). Many AI/ML problems have stages that are trivially parallelizable - if you can divide your problem into iterations where a subgraph of nodes communicates internally, then sends/receives updates to other subgraphs, it's very similar to an iterative map-reduce algorithm, perfect for networked cloud systems.
And as you're tuning your hyperparameters, you don't know what the performance characteristics are, and you will absolutely want to run experiments in parallel, until you find the right settings that you'll use in production. You'd need to invest in a LOT of redundant bare metal to have that capability. As Netflix puts it in this presentation, the key to effective machine learning is iterating often, and that means having a lot of parallelism to bring to bear. https://www.infoq.com/presentations/machine-learning-netflix...
Of course it depends on how much data you're training on, and how up-to-date you need your model to be.
https://aws.amazon.com/blogs/ai/the-aws-deep-learning-ami-no...
Quick Heads up in case anyone wants to do the same. The only European region that supports p2 at the moment is Ireland (I tried Frankfurt at first).
You cannot currently count objects other than faces and bounding boxes for non face objects are not currently provided.
For the time being existing solutions in the opencv/tensorflow/MXNet realm might provide more of the result you're looking for.
I would go with a DIY solution, since Rekognition will get expensive quick ($1/1000 images). If you did 1fps running continuously, you'd pay $87/day.
Nearly any ML framework with an ImageNet pre-trained model will work. Accuracy will be better if augmented with more labeled deer images (ImageNet probably only has frontal poses) but should be alright straight out of the box.
You could run the whole thing locally on a cheap android phone for free instead of AWS cloud pricing, 1fps will be fine with TensorFlow or MXNet on Android. Then just setup email alerts.
I wonder if this impresses this blog's audience, or does exactly the opposite...
His wealth.
The obsession with cyber-malthusianism among the owners of tech informs me more about their valuation of humanity than about the future of tech. Fortunately there's a more level headed analysis that was published by the NYT recently for those of us that aren't consciously or subconsciiously all-in about the idea of making Snow Crash non-fiction https://www.nytimes.com/2017/02/20/opinion/no-robots-arent-k...
edit: I did see that they have a podcast, nice, something to listen to while I walk 2 hours in the middle of the night.
I highly recommend Inoreader as a great RSS reader.
Regarding Blogger back in like 2011 or so, I saw RSS feed all the time at the bottom (usually still do).
What I want is offline daily download of specific stuff. Being lazy myself, to build scrapers and download them to my phone but not an Android dev only web dev at this stage.