GPT-3 is no longer the only game in town
lastweekin.ai
lastweekin.ai
You can use reduced versions of language models with extremely good results.
I was involved in training the first-ever GPT2 for Bengali language, but with 117 million parameters.
It took a month's effort (training + writing code + setup) and about $6k in TPU cost, but Google Cloud covered it.
Anyway, it is surprisingly good. We fine-tuned the model for several downstream tasks and we were shocked when we saw the quality of generated text.
I fine-tuned this model to write Bengali poems with a dataset of just about 2k poems and ran the training for 20 minutes in GPU instance of Colab Pro.
I was really blown away by the quality.
The main training was done in JAX, and it is much faster and seamless than PyTorch XLA, much better than TensorFlow in every way.
So, my point is, although everyone is talking hundreds of billions of parameters and millions in training cost, you can still derive practical value from language models, and that too, at a low cost.
I don't see how I can be of help.
But I can talk. Leave me something through which I can reach you. And I will reach you within a week.
Just wait until NVidia comes with a "Neural AppStore" and corresponding restrictions. Then wait until the other GPU manufacturers follow suit.
DeepMind and OpenAI have a bad rep in this regard.
But a lot is available for free (as in beer and speech).
And most of the research papers are released in arXiv. It's very refreshing.
The bottleneck is not the knowledge or code, but the compute. People are fighting this in innovative ways.
I have been an inactive part of Neuropark that first demoed collaborative training. A bunch of folks (some of them close to laypeople) ran their free Colab instances and trained a huge model. You can even utilize a swarm of GT1030s or something like that.
Also, if you have shown signs of success, you are very likely to have people willing to sponsor your compute needs, case in point- Eluether AI.
The situation is far from ideal with this megacorps rat race [0], and NLP research being more and more inaccessible, but it is not completely dark.
[0]: I, along with many respected figures tend to think that this scaling up stuff approach is not even useful. We can write good prose with GPT-3 nowadays, that are, for all intents and purposes, indistinguishable from text written by humans. But we are far, far away from true understanding. These models don't really understand anything and are not even "AI", so to speak.
The Transformer architecture, the backbone of all these approaches- is too brute-force-y for my taste to be considered something that can mimic or, further- be intelligent.
I'm glad this all worked out for you. This is unrelated, but I just want to say that I hate how many people Google managed to convert to TPU with their research program and that their managed TPU/GPU offerings are absolutely horrible and infuriating to work with unless you somehow get on their radar.
I have not converted to TPU, because it is literally offered by only one company. It will be the height of "vendor lock-in".
But, I must say that JAX on TPU is faster than anything and everything that I have ever seen.
Umm, not really.
You are talking about single and multi-label classification tasks, maybe?
Bengali poetry is just like poetry in any other rich languages like English, French, etc.
What these models do, from a high level, is that they learn the distribution of the data. In this case, they learn the style of the poets.
Writing poetry in specific styles has been done long before Transformer architectures came.
See-
- https://www.tensorflow.org/text/tutorials/text_generation
- https://machinelearningmastery.com/text-generation-lstm-recu...
The goal of poetry generation is to generate something that is unique, is in the poetic style, and is coherrent, grammaticallly correct, ideally indistinguishable in the eyes of a human.
It's an interesting subject, I think, since poetry has so many forms it can take and you need the output to capture the idiosyncratic aesthetics and "inner world" of a piece of verse.
I've actually tried using GPT-3 to generate poetry and the naive approach of just sending prompts and text snippets through the API had wildly varying results. Some pretty good, some that were basically word salad and nonsense.
But I also run a poetry journal on the side, so maybe that skews my understanding of it!
There are a lot of pretrained language models (preferably at HuggingFace Hub). Take one, and fine-tune it on a much smaller dataset.
You can then pass prompts to let this model write poetry.
I would also suggest learning the fundamentals of Deep Learning, RNNs, LSTM, Transformer, etc. Because even if you know all these, this kind of taks is not trivial.
I would use HuggingFace + PyTorch for these kind of tasks.
Besides Deep Learning knowledge, you will also need to know your way around tokenizers, controlling parameters, evaluation metrics, etc.
Having some basic understanding of poetry is helpful, but to be able to truly apply your knowledge of running your poetry journal, you will really need to be on the proverbial edge. I hope you get there, and maybe get some papers out!
Although I worked on a language model because none existed in my language, my expertise (and employment) lies in Computer Vision. So, take what I said with a grain of salt.
Many language models do predict.
In this case, they either try to predict what the next word (or character, or sub-character in case of Chinese, Japanese, etc.- this is totally the decision of the DS) is , or what some "masked" word are.
w_i becomes w_(i-1) in the sequence
where w_i is the last word generated
The ones that are trained to be able to predict the next word are the ones that are good generators.We thought Haiku was cool, but none of us knew it, and learning resources were scarce. But we are very happy with Flax.
We used the Bengali subset of the mC4 dataset [0] for training the GPT-2 model.
The most exciting part about it is showing us there is a path forward by scaling and prompting, but you can still do much better with a smaller model and a bit of training data (which can come from the expensive GPT-3 as well).
What I expected from the next generation: multi-modality, larger context, using retrieval to augment input data with fresh information, tuned to solve thousands of tasks with supervised data so it can generalize on new tasks better, and some efficient way to keep it up to date and fine-tune it. On the data part - more data, more languages - a lot of work.
That said, what I'm really curious is how those other models stack up against GPT-3 in terms of performance -- does anyone know of any comparisons?
The answer is at https://github.com/kingoflolz/mesh-transformer-jax
It has detailed comparisons and a full breakdown of the performance, courtesy of Eleuther.
Get it, cause it's a generative transformer? Hah
No it won't. Capital was (and still is) cheap when they made that purchase. A billion used to really mean something, but it's chump change for companies with valuations over $1 trillion.
Seems like aside from Eleuther.ai you can’t use the models freely either, correct me if I’m wrong.
OpenAI has strict requirements for the usage of GPT-3. For instance, you cannot automate posting to social media without a human in the middle.
OpenAI is pivoting to corporate evil, and to do that properly they need proprietary assets to rent out.
None of these systems could work if only facts were used as input. Maybe some "factual" score could be presented with the data, but wow, that's a philosophical problem there.
I think other comments are more valid. Sale potential. Buyouts.
IBM managed to beat Garry Kasperov using symbolic AI did they not? So in what way does it not work?
Ok, I should be clearer. ML approaches are way way better than symbolic approaches. Given almost any problem, it is much much easier to make an ML approach work than any symbolic approach.
Yes, chess was first solved symbolically, but it's since been solved by ML better and more easily, to the point that stockfish now incorporates neural nets [1]. ML has also given extremely high levels of performance on Go, Starcraft, DoTA, and on protein folding, image recognition, text processing, speech recognition, and pretty much everything else.
I would challenge you to name any (non-simple) problem where traditional AI methods are still state of the art.
Theorem proving, classical planning, SAT solving, robotics, search, in particular adversarial search, program induction, knowledge representation.
Plus all the stuff that used to be considered "AI" but aren't anymore, like rule-based systems (e.g. for fraud detection) etc.
Sorry, I know you asked for only one.
Robotics is already on its way there. Is this program induction https://paperswithcode.com/task/program-induction ? Looks like it's headed towards ML too. I suspect you could stick ML into SAT solving and get yourself a system that worked pretty decently.
When you say "ML" you probably mean the deep neural networks approaches that are currently state of the art for machine vision etc. Deep neural network approaches have been proposed for the task of program induction but they generally lag well behind symbolic machine learning approaches.
The most coherent efforts to tackle program induction by neural networks that I am aware of is the work of Dawn Song's group at Berkeley [1] and of Joshua Tenenbaum's group at MIT. I can't find a handy link to a compilation of the latter group's work but the Dreamcoder paper in the paperswithcode search you linked to was an interesting milestone [2].
There is a lot of work on neuro-symbolic approaches to program induction, for example see the recent (two weeks ago) NeSy workshop [3], part of the first International Joint Conference on Learning and Reasoning for some new work in that burgeoning field. Statistical Relational AI combines symbolic with probabilistic learning; see the STAR-AI workshop [4] also at IJCLR.
If you're interested in recent developments on the front of program induction (again, learning programs from examples) then IJCRL is the conference to keep an eye on.
_______
[1] https://sunblaze-ucb.github.io/program-synthesis/index.html
[2] https://arxiv.org/abs/2006.08381
These search and sampling algorithms still play key roles in game playing AI (chess, poker, Go) and natural language generation. It is the human knowledge, specification and heuristics, part that tends to be more readily replaceable. A lot of control flow and data-structures that powered old AI approaches can be found in databases, compilers, type inference, computer algebra and even the autodiff libraries neural nets are written in.
Video game AI, constraint solving and business rules engines are probably closest to still using the full symbolic approach rather than merely extracting the control flow and structures portion.
We can therefore make a compact prediction: learned approaches replace human written computer programs (specifications, rules systems or heuristics) whenever human contribution is not valuable or is somehow harmful to robustness/generalization.
Around November 2023? I'll try to remember.
Did ML methods best classic AI in dialog comprehension, say to the level of SHRDLU? I'm curious, can ML system do that - https://en.wikipedia.org/wiki/SHRDLU ?
A simulated robot hand controlled by natural language to move blocks inside a virtual world. It's not terribly useful but nothing that was created since can do any better and the state-of-the-art NLP approach of large language models is completely incapable of anything like it.
Which is a bit sad, really, if you think that SHRDLU was created by one graduate student, fifty years ago.
Lossless file compression. As far as I know none of the algorithms in widespread use are neural-based, despite the fact that compression is clearly a rich statistical modeling problem, at least on par with GPT-3-style language understanding in difficulty. There are published attempts to solve the problem with neural networks, but they simply don’t work well enough to date. Modern solutions also still use old-fashioned AI ingredients like compiled dictionaries of common natural-language words — any other domain where nat-lang dictionaries are useful has been conquered by neural solutions, e.g. spelling and grammar checkers.
http://mattmahoney.net/dc/text.html
I don't see examples of high performing symbolic AI based compression algorithms anywhere, but again I am very ignorant, do you have examples?
I’m also not an expert in symbolic AI — my comment above is more about neural vs. pre-neural NLP methods, rather than symbolic AI, which I admit drifts a bit from the parent. A compressor replacing word tokens with dictionary indices is definitely symbolic but it’s not especially “AI”.
2. For your example of chess, for some time now ML engines are pretty much untouchable by engines based on pre-ML methods.
Perception tasks were traditionally attempted with statistical machine learning approaches rather than symbolic AI, for example the Perceptron was a very early neural network that was used in machine vision, created by Frank Rosenblatt in 1958.
A lot of that research was carried out under the rubrik of "pattern recognition" rather than machine learning. In any case, no, "we" did not try "to solve [those problms] for decades with symbolic AI". Symbolic AI has traditionally focused on reasoning, which is generally considered to be on some kind of separate level to perception.
As to chess engines, they're still symbolic-statistical hybrids. E.g. the Alpha-x family combines Monte Carlo Tree Search with neural nets that learn an evaluation function etc.
In any case, it seems to me that while real progress has been achieved in language modelling, the same cannot be said for language understanding. That's a bigger conversation but anyway, modelling is still what statistical learning techniques do best, whereas anything to do with semantics, you still need some kind of symbolic approach.
I never thought of HOG and SIFT as "symbolic". If I remember correctly, they were just sets of hand-crafted features? But, features for classifiers, like SVMs and so on.
much like deep learning was invented decades ago but didn't become feasible until technology caught up, could the same be true for symbolic AI?
i.e., is the ceiling for symbolic AI technical and transient or fundamental and permanent?
From that perspective, I don't see how symbolic AI would be competitive but there would be a role for symbolic AI in designing systems that can be comprehensible for humans, but perhaps just as a distillation/compression output from a non-symbolic system. I.e. have a strong "black box" ML system that learns to solve a task, and then have it construct a symbolic system that solves that task worse, but in an explainable way.
I will cling to these goal posts every time. Search was and still is AI, unless you think Russell and Norvig should have named the field's foundational textbook something other than "Artificial Intelligence: A Modern Approach"
There was a point when it was the other way around, this is not static but the result of resources being poured. The data heavy, computational heavy, black box style of ML gives power to large business over small business. So it's seen as a safer bet than symbolic A.I. This in turn makes it work better, which makes it an even safer bet. Notice that startups dream of being big business so they still pick ML.
Also notice that in some domains ML is still behind symbolic A.I., for instance a lot of robotics and autonomous vehicles.
Classical AI has failed in robotics. There's practically no field where it does worse than in robotics. It's getting cut out of every part of every system one piece at a time, and being replaced by ML methods that really work. Even Kalman Filters aren't safe.
IIRC and Waymo vehicles used ML for some of its perception but also depended heavily on the lidar and a rules based approach. Sadly I can't find a link at the moment.
What examples are you thinking of?
[1] https://www.quora.com/How-does-Boston-Dynamics-use-AI-machin...
Boston Dynamics was definitely an exception for a long time. Certainly the perception systems that they are adding use ML, but you are right that they use expressly use classical methods for their control system.
"I'm not a real programmer. I throw together things until it works then I move on. The real programmers will say Yeah it works but you're leaking memory everywhere. Perhaps we should fix that. I'll just restart Apache every 10 requests." -Rasmus Lerdorf
"I was really, really bad at writing parsers. I still am really bad at writing parsers." -Rasmus Lerdorf
"We have things like protected properties. We have abstract methods. We have all this stuff that your computer science teacher told you you should be using. I don't care about this crap at all." -Rasmus Lerdorf
The incumbents do discredit PHP though. For instance, facebook was built on PHP, and still runs on it. They used the language of personal home pages to give every person on the planet a personal home page. Nevertheless, once they suceeded they forked PHP with a new name and isolated devs culturally.
Terrible code in PHP is possible, therefore likely. I say this having spent over a decade writing it and half of that time fixing OWASP bugs created in it.
The incumbents hate it because their vendors use it and everyone is worse off for having their vendors use it. And Facebook did use PHP in the first few years, but they quickly started compiling it (HipHop) and later converted their code based to use a different strongly-typed language based on PHP (Hack). They stopped using PHP because it is a starter language.
If you're just there to paint, but painting with mashed potatoes, you SHOULD have been more worried about your paint chemistry.
"Using these toolkits is like trying to make a bookshelf out of mashed potatoes." -JWZ
Excluding participation in having the time and resources available to overcome the initial inertia required to become productive is a form of opportunity and earning segregation.
Despite having a background in your tech, there is little more if satisfying than people experiencing putting tech to work for them, rather than the other way around or being dependent on others.
This is akin to teaching an adult human about a specific domain. Better to just do that than make a whole new human from scratch!
Only by analyzing the page title (from the bookmark, not by re-fetching the url) and eventually also the domain name.
You could give it a couple dozen bookmarks with example classifications and then feed it a new bookmark and ask GPT-3 what category the page belongs in. Repeat for the entire data set.
For data extraction you could ask questions about the titles. Maybe have it list all machine-learning model names that appear in the set of bookmark titles.
>>> You are spending way to much time browsing.
They can also all be adapted and fine tuned for other tasks in content classification, search, discovery etc. Think facnial recognition for topics. Want to mine a whole social network for anywhere people are talking about _______ even indirectly with very low false negative rate? You want to fine tune a transformer model.
Bert tends to get used for this more because it is freely available, established and not too expensive to fine tune but i suspect this is what microsoft licensing gpt-3 is all about.
> The fact that GPT3 works at all for English indicates that English is too low level and forces a lot of redundancy (low entropy).
I don't think the goal is to compress information/language and maximize "surprise".
Why? 99.9% of programming being done is composition of trivial logical propositions, in some semantic context. The things we implement are trivial, unless you're thinking about symbolic proofs etc
I was a Jetbrains fan but this thing takes productivity to a whole new level. I really don't think I can go back to my normal programming without it anymore.
I'd love to use it but not when there's such a risk of compromising the code base.
Simply leading the horse to water is enough in something like PHP:
// instantiate cURL event from API URL, POST vars to it using key as variable name, store output in JSON array and pretty print to screen
Usually results in code that is 95-100% of the way done.
Nice trick: write a comment describing quickly what your code will do (“// order an item on click”) and enjoy the complete suggested implementation !
Other nice trick: write the code yourself, and then just before your code, start a comment saying “// this code” and let copilot finishe the sentence with a judgement about your code like “// this code does not work in case x is negative”. Pretty fun !
https://sdtimes.com/monitor/using-gpt-3-for-root-cause-incid...
Of course this will be another blow for journalists, who rely on this skill for their income.
How would you tell the difference from the real thing these days?
You can use this instead of having to train a custom model for every specific task.
A good example is DALL-E[0]. Now, what's interesting to me is the emerging idea of "prompt engineering" where once you spend long enough with a model, you're able to ask it for some pretty specific results.
This gives us a foothold in creating interfaces whereby you can query things using natural language. It's not going to replace things like SQL tomorrow (or maybe ever?) but it certainly is promising.
Life online is about to become very annoying.
One example of GPT-3 powered chat bot: https://www.quickchat.ai/emerson
If you think about business uses you can actually get advices from Jerome Powell, simulated by GPT-3.
If someone use GPT-3 to simulate Warren Buffet, he can extract even more value.
https://www.institutionalinvestor.com/article/b1tktmhcfdyqsk...
> Mr. Powell: I don’t want to get into the details of taxes.
> Interviewer: Are you in favor of a cap and trade system?
> Mr. Powell: I don’t want to get into the details of a cap and trade system.
> Interviewer: How do you think a global pandemic would impact the financial system?
> Mr. Powell: A pandemic will have impact on the economy. It will have an impact on the financial system.
> Interviewer: What would the impact be?
> Mr. Powell: If a major pandemic occurred, the economy would be affected, and the financial system would be affected.
Maybe I'm a bit harsh on GPT-3, but I'm not nearly as fascinated by this kind of output as the author.
> Interviewer: How do you think a global pandemic would impact the financial system?
> Mr. Powell: A pandemic will have impact on the economy. It will have an impact on the financial system.
> Interviewer: What would the impact be?
> Mr. Powell: If a major pandemic occurred, the economy would be affected, and the financial system would be affected.
Yes, GPT-3 can provide a convincing chatbot, but it shouldn't be confused for domain expertise.
> > Interviewer: What would the impact be? > Mr. Powell: If a major pandemic occurred)
Output( it could be very disruptive. > > Interviewer: How disruptive? > Mr. Powell: Well, it could be disruptive in a number of ways. > > Interviewer: How so? > Mr. Powell: Well, one of the first things that would be disrupted would be the supply chain. )
Using prompts well makes a huge difference.
If you parse the generated output, classify it, then develop a decision tree that uses further prompts to refine the response, you can get more sophisticated, valuable responses.
The output in the parent is comparable to an off-the-cuff interview response. If you emulate a deeper thought process, you can get more meaningful output, and if you use the right prompts, you can access the semantic networks in the model related to your domain of interest.
I wonder what if somebody designs an electronic currency rewarded as payment for general GPU computations instead of just computing hashes? You pay some $, to train your model and the miner gets some coins.
Every one is happy, electricity is not wasted and the GPUs gets used for a reasonable purpose.
Making the mining process produce useful output that can be resold nullifies the purpose as it means an attacker can now mine "for free" as a byproduct of doing general-purpose computations (as opposed to tying up dedicated hardware), lowering the barrier for a 51% attack dramatically.
This is very much not true. A central class in complexity is NP whose problems are hard to answer but easy to verify if the answer is yes.
E.g. is there a path visiting all nodes in this graph of length less than 243000? Hard to answer but easy to check any proposed answer.
Training NN ML is much slower than inference (1000x at least) because it has to calculate all of the gradients.
And second, the principles of electronic currency are different from gold/money. That's why crypto uses GPU ;)
If their name was ProfitableAI, there'd probably be fewer complaints.
I believe this will displace the workers of places like the Internet Research Agency and some Marketing specialists. While that accounts for a large percentage of propaganda jobs, that is easy enough to adapt from (nor do I really care given who is generally doing it), and marketing people who I generally don't have respect for either (I mean their founding document is literally titled "Propaganda"). The real marketing pros are good at managing hard to handle talent or bad PR disasters, which are actually meaningful skills that this won't affect.
So basically no moral or real consequential impact on society whatsoever.
Maybe that is a good thing, maybe a bad thing, but unless there is a breakthrough in methods this is a dead end. Impressive though.
This will just lead people into the equivalent of using the autocorrect answers too often as if you used the suggest feature of a phone keyboard on a sentence. Eventually it just breaks down and stops making any sense. The code equivalent is going to be bugs, security holes, and programmer headaches.