Andrew Ng is raising a $150M AI Fund
techcrunch.com
techcrunch.com
That was one helluva course, challenging and interesting, and fun all at the same time (and so much "concretely" - lol).
From what I understand, that course is still available thru Coursera (which Ng booted up after the ML Class experiment; Udacity was Thrun's contribution after his and Norvig's AI Class, which ran at the same time in 2011).
After learning more about him, he's probably the only person in the world that I envy. I resonate with his ideas a lot, but I'm like 1% of what he is. It makes me a bit sad. I'll be taking his new course as well and hopefully one day I will be able to work in the same field as him.
I wrote down my thoughts after taking his new DL course. Hope it helps you all :) https://medium.com/towards-data-science/thoughts-after-takin...
When I took the ML Class (I also took the AI Class at the same time, but had to drop out due to personal reasons - but I stayed in on the ML Class and finished it), I hadn't really touched linear algebra since high school.
I graduated high school in 1991; Ng's course was 20 years later.
I also didn't have any stats or probability experience under my belt. Nor anything about derivatives or integrals.
I basically had to pick all of this up on-the-fly (fortunately there are internet resources), and even to this day, I barely understand them (I understand matrix operations mostly, but I struggle with probabilities, and I have little-to-no idea on derivatives or integrals).
After high school I went on to get a 1 year, virtually worthless today associates degree from a now-defunct voc-tech school here in Phoenix. Since then, I've been steadily employed as a software engineer here in the valley, and well compensated (I believe) for it. I own my own house, and I have zero debt except for a mortgage.
Given all of that, one should be able to see how such a course would be a challenge. There were a ton of people who signed up, but from what I understand, the majority dropped out after the first couple of weeks. This actually seems "par for the course" though for MOOCs.
I know it was a simplified intro to ML, but for me, it and what I took of the AI Class taught me more than what I ever was able to figure out on my own, especially on neural networks. The light really clicked on for me there. But I was really disappointed to have to drop out of the AI Class.
Later, in the Spring 2012, after Udacity had been established, they weren't able to offer the AI Class as one of their courses. So Thrun came up with another course, which was originally titled "CS373 - How to Build Your Own Self-Driving Vehicle" - and I jumped on that one, and completed it as well. I found it fairly challenging too (but not as challenging as the AI Class was). This course has since been renamed to "AI for Robotics" - which is more apt, I think.
It took a while - but eventually the AI Class was made into a course (I think there was some kind of licensing issue, but I don't know for sure, that was preventing it from being part of Udacity's offerings). I have yet to retake it, but it is on my list (plus a ton of others).
Today, I'm in the home stretch of the 3rd term of Udacity's Self-Driving Car Engineer nanodegree. I'm struggling mightily to get my path planner project to work properly, but I almost have it done (it can make it around the track, but for some reason my behavior planner isn't costing things properly). Got an elective, and the integration project to do, all by mid-October or so.
I don't know if any of this will lead anywhere for me career-wise. I'm happy with my current employer, so I expect to stick around here for a while. I have hopes, dreams, ambitions to perhaps get a degree of some sort in CompSci. I want to really learn more mathematics. I've always been a lifelong learner, but this kind of stuff is really fascinating to me, even if it is (what seems to me at least) complex and not always intuitive. But if it were easy, it probably wouldn't be as fun (but I will say Keras and Tensorflow really make things much easier than when we had to implement a neural net in Octave and Python).
Having not completed an AI course, I tread lightly.. However, I would guess that this project involves re-implementing an established solution. -- It is work like this that drives me away from such courses; as I can't imagine how creative practices are promoted, instead "correct" techniques are repeatedly hammered in.
> I don't know if any of this will lead anywhere for me career-wise. I'm happy with my current employer, so I expect to stick around here for a while.
Professionals learning to program late in their career typically have a misconception that they're only eligible for entry-level positions in the field of software engineering. Many fail to realize that the 10+ years of experience in their own field can be coupled with their newfound-skill, giving them a background unlike that of many existing professional developers.
The Ng class was a good introduction, but it was mostly applications, not the mathematics and theory behind them.
[1] https://www.edx.org/course/learning-data-introductory-machin...
Doesn't die -- HTML being a specification in name only, there are a lot of really crazy web pages out there that render on browsers but are pathological edge cases.
Does a good job of distinguishing 'good' links from 'bad' links on a page. -- Lots of pages have links that should not be followed, some are easy they are rendered in the same color as the background (SEO black hat link juice) and others refer to crawler traps.
Crawler traps come in many forms -- Rich Skrenta created a great example one where the page generated a random number and said "%d is an interesting number" here are two more interesting numbers "%d and %d" the each link went to a new URL that ended in the number. So if you tried to crawl that site exhaustively you would fill your entire crawler cache with random number pages.
Dynamic importance scaling -- you want to crawl the 'best' pages for a topic so you need to figure out a way to measure which pages are important and which aren't. This was the secret sauce of the PageRank patent Google had but it's been gamed to death by SEO types. So now you need better heuristics to understand which are the more important links to follow.
Effective crawl frontier management - for every billion pages you decide to crawl there are probably 20 to 50 billion pages you "know about". These URIs that are known but not yet crawled are referred to as the 'crawl frontier'. Picking where to go looking in the crawl frontier to find useful new pages is half art and half good machine learning.
Good algorithmic de-packing -- many many pages today are generated algorithmicly from a set of rules, whether it is the product pages on Amazon or posts in a PHP forum, if you can recognize the algorithm early, you can effectively avoid crawling pages that are duplicates or not useful.
Good page de-duping -- There is a lot of repetition on the web. Whether it is the 'how to sign up' page of every PHPBBB site ever or the same product with 10 different keywords in the URI.
Selective JS interpretation -- sometimes the page exists in the JS code, not in the HTML code, so unless you want to store 'this page needs Javascript enabled to run' into your crawler cache you need to recognize this situation and get the page out of the Javascript.
That's just off the top of my head.
When you say google and microsoft have this advantage in creating data sets, is it just the massive size of the web indices they are able to compile or do they use their crawlers in specific ways for compiling structured data that would be more useful for certain ML projects than a general web index?
Are there any tweaks you'd make to a crawler if you sent it out with the purpose of creating a dataset for a specific AI / ML project, rather than a general purpose web index?
> ... do they use their crawlers in specific ways for
> compiling structured data that would be more
> useful for certain ML projects than a general
> web index?
There are many uses for a large index. For example, they decode into structured data for many of the 'one box' results, a small box that shows up on the search results which has the answer to your query, even though that answer came from a web page. This is good for the consumer, they get their answer right away without clicking through to a web page, and its good for Google as it keeps the customer on the search results page with its advertising rather than having go to some page on the web potentially with someone else's advertising on it.Google also post processed crawl data to indicate the spread of flu in their experiment of extracting health data from query logs.
> Are there any tweaks you'd make to a crawler if you
> sent it out with the purpose of creating a dataset
> for a specific AI / ML project, rather than a
> general purpose web index?
Yes there are many. Some of them made it into the Watson crawler. One of Blekko's claims to fame was their notion of 'slashtags' which were curated lists of known 'good' pages on a topic. Using such pre-validated URI lists can help you improve the fidelity of the datasets you collect. There are also clever ways to use existing data to validate the new data you are looking at. I'm on a couple of patent applications around that space which, if they ever issue, will make things a bit more obvious than they are today :-).They release the frameworks so people learn them and then it is easier and cheaper to find employees.
That is what I think.
If you are just aggregating data, that's not a moat.
A lot of folks are buying data from a bunch of sources to give complete coverage - e.g. a dataset of all plane ticket prices, where before you could only get separate datasets from each of amadeus and rivals.
Someone, usually someone big, just goes round you to build their own infrastructure for some internal function, e.g. Google flights api, and then realises they can replace you as a secondary revenue stream.
Instead, I think you gotta somehow add value to your data, which is best done as a side effect of another buisness.
Reuters have a huge news dataset with amazing annotation because they got thier editors to curate it as it was produced. That's an unassailable free text training set that no one else is gonna match.
So build a dating app that causes users to create a curated dataset. Or a game. Or a buisness tool. Or an api. Or a really good AI secret sauce built in an expensive privately curated training set that ate most of your funds.
What I find fascinating is what AR might do for training AI models. To execute on AR we'll need to digitize a model of our physical surroundings so the software can interact. At that point we'll have a compelling pipeline of actionable data in regards to machine learning - especially for robotics.
More and more varied datasets are needed for this (but understandably they can be seen as valuable on their own, so reluctance to share is understandable - at least from a business perspective).
I spend 90% of my time manipulating data to try to build bigger, better datasets, and only 5% modelling.
If, on the other hand, they are talking about making progress on the more traditional dream of AI, the focus on building these data sets seems to be a sad way to lock us into a local maximum for a long time.
Massive curated data sets are crutches that lead us into narrow-minded hyper-specialized systems. Are any of these funds investing in people working on systems that try to make sense of raw sensory data streams? I don't think we're ever going to move away from data sets by creating more and better data sets.
Yes. That's also my total focus right now. Happy to discuss, email on profile.
Or is it just the time correlation that interests you? Because some of these data sets are very likely to indeed be time correlated. Like a video/audio data set for example.
The problem with the dataset-first approach is that humans are still providing a lage part of the intelligence: defining the domain, defining what good performance on it looks like, carefully designing model architectures, collecting and labeling large datasets, etc. This is fine for narrow task-specific problems, but is not really the end-all of AI, and does not even seem to work well on all well-defined tasks. As an example of another kind of inference, how about mathematical reasoning? I purposely pick one here that is seemingly very formal; should be possible for a computer to do it. Mathematicians are somehow able to invent conjectures and prove theorems without first being exposed to terabytes of labeled mathematical facts. The scientific method is kind of an even messier version of this. Or to take something laypeople do, people can usually learn games to at least a passable level from just a handful of playthroughs, not AlphaGo-style millions of plays (imagine if you had to play even 1,000 games of MT:G before you got the basic hang of it...).
All this kind of stuff is quite well-represented in the literature though, if you mean the scientific literature. Pop-press AI writing tends to cover a pretty specific subset of what's going on at AI conferences.
You can feed a million images that are labeled "cat" or "no cat" to one of these systems and it can achieve a human-level proficiency at identifying cats in images. But, it won't be able to do anything other than identify cats, it's far too narrow to be considered intelligent in any way.
If you can feed a series of timestamped photos to a system, basically an unlabeled arbitrary video stream, and you can demonstrate that it formed some notion of what a cat is, that would be very interesting indeed.
With modern data sets and labelling, so much of the problem domain is deeply hard-coded into the system. It doesn't have to learn what letters and words mean and how to identify them in a totally arbitrary* visual or audio stream. It just gets a relatively minute amount of structured data that it has an embedded understanding of what to do with.
Sure there are things like OCR and speech-to-text, but I don't think you could just run your streams through those, there's just so much subtle information loss. In order to make AI that really has a chance of reaching what humans would call intelligence, I firmly believe it has to make meaning out of some kind of raw sensory experience analogous to ours.
*Ok, not totally arbitrary, a human would not learn language from a video of a forest, and that's where the "labelling" comes in, from observing other people using language, but a child can still learn any language from just that, and the labels are often vague, inaccurate, contradictory, complex, abstract, etc. There's no master training set with the right answers, you have to decide for yourself. And humans were also capable of bootstrapping language from nothing. I just don't see modern supervised learning systems ever doing things like that.
> In order to make AI that really has a chance of reaching what humans would call intelligence, I firmly believe it has to make meaning out of some kind of raw sensory experience analogous to ours.
And why must it be able to do this from scratch? Why must it be hampered with the same limitations we have?
Self-Taught Learning: Transfer Learning from Unlabeled Data
http://www.andrewng.org/portfolio/self-taught-learning-trans...
Google Brain/DeepMind are also pushing some of those ideas. They must be, since they aggressively poach all the top researchers from those labs...
Ng approach is different: he wants a world powered by Deep Learning, so his goal is to make applied deep learning thrive. His strategy to do that: give those data-hungry models even more data, which is completely reasonable.
Those two approaches - fundamental research and applied deep learning - are often referred to as AI, causing much confusion.
Are they talking about basically getting the most out of the current ML type systems?
Since Deep learning is a subset of machine learning, the sentence retains correctness but there are now two equally valid interpretations.
Okay, I do have something relevant to say. I don't see what advantage there is in raw sensory data streams, models can be trained off-line to operate on sensor streams just fine. What we actually want are systems that learn adaptively and on-line. To do that well, they'd need to also be data efficient.
I mentioned this in another comment, but I don't know that we can train these models just fine, a human can get a lot more information out of an audiovisual stream than a rudimentary transcription of recognized speech, objects and/or text.
If we make these models more sophisticated to capture more information (e.g. body language, tone, context), we have to decide how that information is structured and communicated to the "higher level meaning interpretation" stage. No matter what, their output is going to be more rigidly structured and will contain less information than the raw stream. The extent to which the sensory processing model captures the human-recognizable information in the stream is the extent to which you have created an intelligent system.
There is some form of this structuring and reducing happening inside our brains, but we will never get machines to do that if we continue curating structured data sets. We, the researchers, are using our human intelligence to process raw sensory data and put it into a nice format for the AI. They need to be able to do that themselves.
Do we? I'm not convinced. I've joined different sensory inputs before and just glommed together the top of two nets, I didn't have to create any of my own representations.
> but we will never get machines to do that if we continue curating structured data sets
Curated datasets are important though. Want your machine to understand more about the world around it? Then you need high quality inputs in formats you can load in. These data also need to be licensed correctly.
> They need to be able to do that themselves.
We don't chuck kids out into the wilderness and expect them to come back as a useful member of society. We have a huge range of inputs specifically curated to help (from toys and shows to school curricula). Later on with specialisation we pay for extremely carefully selected data, presented in a specific order!
Creating high quality datasets is vital for really anything from niche specialist systems to large general ones.
Marge: Do I have to be dead before you’ll help me?
Wiggum: Well, not dead – dying. [Marge gets up to leave] No, no, no, no. Don’t walk away. How about this: just show me the knife in your back. Not too deep, but it should be able to stand by itself.
On the other hand, confidential information of corporations is pretty well guarded, and confidential contract terms, costs, pricing, and other information that doesn't get shared among peer companies or competitors is what will propel AI into the economic stratosphere. Finding ways to get confidential data into learning systems and provide actionable feedback is the killer app.
Or is it more hedge funds - like satellite images of Walmart car parks to estimate the revenue figures?
Either way these don't seem like things we need AI to pick out patterns ?
Or am I missing something?
If an AI has access to a broad spectrum of confidential information, it can reliably answer questions about the state of the market on an anonymized basis. This has the effect of normalizing sourcing behavior, which reduces purchasing and sales friction and improves overall efficiencies.
A tollgate recording vehicle pass-through with a timestamp, a train station gate recording individual pass-through with a timestamp, etc, with huge volume could be useful data.
Anyone have links to interviews or information on Ng's vision? I'd love to hear the details.
Apparently they don't know also. The website is empty.
I thought the AI hype is over again but apparantly not.
Start with the word "urn". Now don't drag it out, make it short. Make the "n" and "ng" sound at the end (urng). Now take out the "r" sound (uhng).
It seems like Americans do tend to pronounce it "ehng" instead.
I see this question asked very often in the last year or two. I am not an expert on deep learning, nor have I taken the deeplearning.ai track, but am currently learning about the topic.
Some of the resources out there have nice primers. I think you need to be somewhat comfortable with understanding the usual log/exp functions for the very base, understand calculus, with partial derivatives, and be used to linear algebra and matrix operations. Some good understanding of statistics could be useful as well when learning about ML. I don't think that a good background in CS is necessary for this stuff. This has not much to do with programming languages, operating systems, Turing completeness. Maybe having a good base in algorithms could be useful for _implementing_ the libraries to make sure they are optimal.
I was wondering why do people ask this question (or ask about resources on learning about this topic in general), when answers on this are so easily findable online.
The coding exercises are setup very nicely for you with a lot of code comments and hints. A lot of the time, if not all the time, you're simply completing the right hand side of an assignment.
What are "advanced simulation tools" ? something like https://github.com/marcotcr/lime ?
I believe most regulations that block this kind of investment are done in the name of protecting the little guy.
Otherwise, you will get people that have negative net worth, maybe $20k in credit card debt + student loans that they pay minimums on, and a few kids dumping a year of savings into Snapchat and losing it all.
While it sucks for small players that are ok with high risk investing (and would be ok losing it), there just isn't a way to stop the flood of stupid that would come with it.
Probably the closest thing I've seen to being able to invest in something like this is cryptocurrecy,
That's why they limit investment size. It's not to be purposefully exclusionary, they just have to be mindful of efficient use of time.
So is this Elon Musk's arch-nemesis?
doesn't quite jump off the powerpoint slide as well though
I would love to just use the amp version for all TechCrunch pages. Anyone in the mood to make a chrome extension? (I'm on desktop, and the results are still clean without adblocker)