Ex-Googlers raise $40M to democratize natural-language AI
fastcompany.com
fastcompany.com
1/ Open source companies like huggingface (https://github.com/huggingface), explosion (https://github.com/explosion), Fast.ai that democratized access to ML and provided a set of tools & language models for engineers. These companies didn't only manage to build a company and open source projects, they built a welcoming community and it's impressive to see all these students use these tools to tackle bigger problems. That's what"democratized" mean to an engineer[1].
2/ Media companies that built ML APIs to ease the use of such services like OpenAI. I think this company falls under this category.
[1]: https://marksaroufim.substack.com/p/machine-learning-the-gre...
You can refer to this discussion for more details: https://twitter.com/migueldeicaza/status/1285204129225281536
I guess I'm now "old" in the sense that it never occurred to me to search for this kind of info on Twitter.
I respectfully disagree with this statement, there is a large body of work focusing on making this easier. The HW/SW stack vary from a a company to another.
I guess 'democratize' rolls off the tongue more easily than 'heteroligocratize'.
"Democracy" and especially "the people" can mean lots of different and completely contradictory things to lots of different people. I'm always very cautious of anyone who invokes "democracy" and "the people" directly, I don't automatically consider them untrustworthy but I prefer a direct argument for a particular idea. If someone's truly a democrat, they'd be unafraid to make their point without such potentially dishonest tactics as people would democratically choose that idea anyway.
* I'm massively oversimplifying here - there's such things as Blairism and One-Nation Conservatism which blur these lines enormously.
"Data dumps" are privacy-friendly. For example, a user can download Wikipedias dumps and search through them to her hearts content, and never use the network. Zero data collection by third parties. All those observing the network can see is that she downloaded some data dumps.
"Web APIs" are "tech" company and surveillance-friendly. Network access is required and all activity is observed and recorded. Web APIs are also used as a means of controlling access to what is often publicly available data/info. The company does not own the data/info, its a middleman. "Too many" requests, API user gets cut off.
Too often its publicly available data/info that is being served by APIs. Hard to sell data dumps of public data as a "product". I sometimes see entities that provide "data dumps", e.g., a corpus, for free who then try to restrict usage of it through a license, yet they do not themsleves own the data. Whether they even have the rights to "license" it is debatable. They are never legally challenged so we cannot say for sure. The more interesting issue is whether they had the rights to collect the data thats in it. The so-called "web scraping" issue.
"They say New Yorkers are selfish and unfriendly, but its all untrue. When I visited, a guy overheard I was a tourist and came right up and offered me a great deal on Staten Island Ferry tickets. Just $7.50 for a round trip!"
(In case you don't know, the Staten Island Ferry is free).
If the former doesn't exist (at a comparable quality), then the latter is better than the nothing.
There is no such thing as a startup with altruistic intent.
To your point though: “democratize x” makes my eyes roll. It’s overused hip marketing speak.
I do understand the point you're trying to make: local autonomy is superior to cloud access mediated by a private commercial entity.
However at this time, we may have a counterintuitive situation where API access is more "democratic" than downloading a huge model.
Based on various reports[1], the GPT-3 model was trained on ~45 terabytes of text corpus (Wikipedia + Web Common Crawl + book texts, etc) and the final runtime model (175 billion parameters) requires ~350 gigabytes of RAM. In that case, the model size is ~1% of the training set.
So "democractize" depends on how ambitious the user is. If you want to use a very large 350GB RAM model, the cloud model with API will be more accessible by the masses than running on local hardware. Last time I looked, an Intel Xeon motherboard has max ram of 128GB so scaling up to 350GB RAM is not going to be cheap or trivial to build.
Let's further extrapolate to a future hypothetical GPT-4 using ~10x multiplier: train on 450 terabytes of text with a model requiring 3.5 terabytes of RAM. How do we make that future huge model accessible to the masses? Probably via a cloud API. Unfortunately, there's an unavoidable hardware capital expense barrier there.
You're misinterpreting my comment. I'm directly addressing this fragment by the gp: >which can be downloaded, and accessed as a library,
I'm not making any moral ideology statements about the model's "openness", "transparency", or "intellectual property".
As a person very interested in playing with something like GPT-3, I'm talking about practical concerns of even running the model. Some type of cloud API access let's me run experiments today. Hopefully the API cost is reasonable or free with limits. I believe that's true of most researchers because they can't afford the hardware in the near future to run a GPT-3-size model as local library.
If the model was open source you could have an API market where providers competed to build the most economic service just like virtual machine companies do with Linux. If there is only one API then everyone is stuck with them and just has to hope they don't change the prices.
There is practical concern for researchers within the possibility that what costs you $0.05 to run today will cost you $500 tomorrow (see Google Maps API for when this actually happened).
Capex and opex of these two are quite different.
Even a low-end tower server can be configured with multiple terabytes of RAM.
E.G. https://www.dell.com/en-us/work/shop/povw/poweredge-t640#tec...
I think the issue people have with centralised cloud APIs for this sort of thing is that there's still a single gatekeeper with their finger on the off switch. In my opinion, instead of throwing out cloud APIs altogether a better scenario would be many gatekeepers with a deliberate diversity of socio-political backgrounds.
The best part of the story for me is the work-life-balance he's able o achieve.
Can't recommend the course enough and it's incredible that such a quality resource is available for essentially zero cost
https://github.com/fastai/course-nlp
https://www.youtube.com/playlist?list=PLtmWHNX-gukKocXQOkQju...
Huggingface are a better example closer to what Cohere do.
>Cohere, he says, offers a platform containing a “full stack” of NLP functions, including sentiment classification, question answering, and text classification.
Google (Cloud Natural Language, Dialogflow), Microsoft (Azure LUIS, Bot Service), Amazon (Lex), and IBM (Watson Assistant, Watson Discovery) all offer Conversational AI and NLP APIs that do exactly what these guys are trying to do here. What is new or unique about them? I work in this space, and the Chatbot/Conversational AI market has been flooded with startups all doing the exact same thing for the last 3-4 years.
https://huggingface.co/transformers/index.html
I recommend it since they already are democratizing AI.
We need better lingo.
For a language processing SDK to be useful, it needs to work inside an offline capable app on my phone. But doesn't that make it way too easy for someone to extract the network and use it to train their "own" AI with distillation learning?
And if they instead only run the AI on their servers so users need to connect through an API, how is that any different from the language AI APIs that Google, Amazon, Microsoft already offer?
And let's say they foot the bill to train a new type of language AI, what stops the big cloud providers from just training something similar? A 200mio non-recurring expense won't stop Google if you have proven that it's a viable business for you.
Asking because I'm pondering similar issues for my AI project.
If you're targeting medium-to-large businesses, I suspect lots of them will buy a licence just to make sure they can't be sued (e.g. https://majadhondt.wordpress.com/2012/05/16/googles-9-lines/), even if it would be realistically near-impossible to detect if they 'stole' it.
lol. It's now a weasel word used by start-ups. Merits an eye roll before moving along.
Like, I don't know how to write an A* on top of my head, but back when I was looking at Google a while ago, they always asked that question, and the prep PDF said so. Kind of hard to fail there...
Not everyone will be able to work there, but it's certainly not the status symbol it was a long time ago. Plenty of "Ex Googlers" are code monkeys like anyone else from any other company.
"democratize X" -> "capitalize on X hype"
NB: the challenge is that X can be any important sounding combination of words
I think during a gold rush, it's good to sell shovels even if can't be the leading shovel sales outlet in town.
I think we're at a point where new NLP techniques will create a lot of value, but it's still hard to tell in advance which cases the current techniques can do well enough to be worthwhile, and in which cases they'll be kinda cool but not effective and so we'll continue to require humans to be involved catching errors. While we sort that out, a lot of companies will need to take a few stabs at trying to get some new model to work for their problems.
Oh, and the fact that Cohere is an expensive solution for a problem that's already been solved for free by open source groups.
It's not about particular models, though, but the ease of use in numerous cloud or self hosted scenarios. What's the value add for Cohere?
When you only have $1m you might as well poke around various ideas and see if you can get another million or ten. Now, if you're sitting on tens of billions of dollars your priorities change - it's far more profitable to grow a $10b pie by 10% than a $1m by 1000%. I'm pretty sure that's what happened to Oracle and IBM, hence we rarely hear about them anymore.
Can someone a successful (not necessarily profitable) concrete application of these things, other than "gpt-3 wrote an article in Guardian and said it wouldn't kill us".
Transformers in general have lots of applications (machine translation, information retrieval/reranking, ner, etc).
It was very successful.
You keep using that word. I do not think it means what you think it means.
The word you are looking for is "mass-market" or perhaps "sell to small business".
The 6 billion parameter model was released just last week with the huggingface api hooks. It's just a few percentage points less performant than Gpt-3 DaVinci in most metrics. You can run it on a laptop with 40gb ram, albeit slowly.
All that to say, the horse has left the barn. What openai is doing is just marketing and protecting an investment. They don't have an ethical or moral high ground.
More about the etymology here: https://www.etymonline.com/word/democratize
Examples of things that are democratized (accessible to virtually anyone) but not commoditized (not the same as what you can get from other providers):
- Coca cola
- iPhones
- Teslas
- Google search
- HN
If you're going to make money by creating an NLP service for wide adoption, you'd hope to create some competitive moat. Otherwise there's no way for you to earn profits to amortize your R&D.
(I mean, like coke, they could say it, but it wouldn’t make sense.)
Possibly one could claim something is democratizing access to some specific aspect?
I do think it probably best to tend away from using the term unless it is a particularly good fit though. (“Best” as in “what I would prefer”, not as in “most profitable”)
Words don't have to mean what their Ancient Greek component parts mean. Meaning = use.
It has at least one other meaning ('introduce democracy as a system of rule').
verb
introduce a democratic system or democratic principles to. "public institutions need to be democratized"
make (something) accessible to everyone. "mass production has not democratized fashion"
https://www.google.com/search?q=democratize
I suppose that would be the second meaning.
Now how it's used here? It's more just a buzzword that suggests they are trying to make it more accessible by selling it as a SaaS. They seem to have a focus on ethically providing access to these more powerful ML systems but what this'll actually mean beyond "cover our ass" is up in the air.
So it's censored. No thanks, not interested.
At what point do we 'allow' AI to determine actions based on perceived (programmed) *BIAS* How can one prevent any bias on an AI's ability to be deterministic.
We need a "product recall" method that doesnt involve Blade Runners and campy one-liners...
SERIOUSLY
So the glib answer is "train it on unbiased data". Depending on your philosophy, this translates either to "manually 'fix' anything you see as 'bias' in the training data", or "use a sufficient amount of entirely unmodified raw data along with an algorithm sufficiently insightful to cancel out all of the sources of inaccuracy introduced by the various sources of data and extract ground truth."
But you know if anyone ever does manage the latter, its results will still be decried as 'biased' by anyone who disagrees with them.
that is the 4th industrial revolution, the "post-correct" world where abundance of information (and its consumers) and speed of its production allows for and results in co-existence of multiple truths (kind of like hyperbolic geometry where one can draw through a point multiple different lines parallel to the given line) with the information space splitting into multiple feudal times like dukedoms.
Anyway, in general for biases i think we have D.Rumsfeld situation - the known/expected biases are known while the AI driven world would most probably bring new biases that we don't even expect.
https://en.wikipedia.org/wiki/Ouroboros
but instead its the warning against letting AI iterate upon itself without external intervention...