Visual ChatGPT
github.com
github.com
Why is that so? It seems counterintuitive. A single picture snapped with a phone takes more space to store than the text of all the books in a typical home library, yet Stable Diffusion runs with 5 GB of RAM while LLAMA needs 130 GB.
Can someone illuminate what's going on here?
Image models can be very off and still produce a satisfying result. Consider that I could literally vary all the pixels in an image randomly by 10% and you'd just see it as a bit low quality but otherwise perfectly cohesive image.
Language models have no such luck, the problem they're trying to solve is way "sharper", it's very easy for their results to be strictly wrong if they're off even a little bit.
So you need a much larger model to get a sufficient level of "sharpness" for text.
It'd be interesting if the parameter/complexity requirements are actually similar once you examine the system as a whole, meaning machine _and_ human brain.
This is a very important point. A group of my colleagues (who are not tech people) are much more impressed with the image generation models than with the chat interface, even though the images are often whacky or just wrong. Yet the fact that it tried is impressive to them, with their minds managing to fill in the blanks.
I wonder how this compares to how a toddler speaks vs. paints/draws, which is typically better in the former than the latter. I'm both cases, we fill in the blanks in our minds.
The mapping and motion parts of the human brain are also old evolutionary constructs that could error correct the output of models.
This is particularly true in a video conferencing situation. If the audio is bad, you miss out a lot. If the video is bad, it's not a big deal.
I don't think it's about precision, in the case of audio vs video - if you remove all the even columns from a video it would be similar to reducing quality, the same can be done with audio - removing half of the frequencies uniformly will just lower the quality.
There’s something interesting in the fact that an image based system doesn’t need as much complexity to capture a semantic model as a verbal system does; I think there’s maybe a parallel there to the way that human minds find it easier to just ‘visualize’ some things as a basis for reasoning about them, but if we can’t ‘visualize’ and instead have to ‘think things through’ it’s a more intensive process.
Like, GPT has well known trouble counting - ask it for five things and it will give you four or six. Humans can offload some thinking about counting to visual/spatial reasoning though.
Whereas for pixels there are over 16 million.
How many more than 250k is that?
Over 16 million.
Actually now that I think about it, neural networks work with real numbers, so a pixel is just 3 numbers. Typical input for an image model would then be around 300x300x3 values. While an input for a language model is around 2000 tokens, but while each token is inputted as an integer into the model, mathematically it represents a 250k length vector, so mathematically the input is 250k x 2000 values. So 90k vs 500M. Also pixels next to each other in an image are related, so you can reduce model size by taking advantage of that (CNNs).
This is utterly wrong. There is a huge amount of redundancy in images compared to language. This redundancy is why image models have yet to surpass language models. In some sense, language is much easier than the vision problem.
At any level of scale of model and scale of training set, images models do surpass language models.
It's not accurate to say "LLAMA needs 130 GB" because there's more than one LLAMA.
The hard part in general is it often doesn't work very well to train small model sizes directly. You can train a very large one and distill it down, in some cases to only 1% of the original size while retaining 99%+ quality. So clearly the 1% size model exists. However training it directly usually doesn't work nearly as well. Best I can tell no one knows for sure why.
Big tech cos are training the largest models as no one else has the hardware/power/money/etc to do so. SO they tend to release the most massive ones that only they can make. There's also a sect in the ML community that thinks "scale" is the answer to the universe...
Then, either the tech cos or the community will make multiple other sizes of models and those are the ones that normies can use.
To me, it could highlight that there is more entropy in language than in images. After all, games have a similar property: some games, like noughts and crosses, or checkers, can have the same model architecture reach optimal play with smaller sizes than are necessary for chess and go[1].
It is certainly true that language has a lot of rules (grammatical, vocabular, syntactic) which are necessary to master it, but irrelevant to having a mental model of a scene; it just sequences a state of things into a stream of symbols. Maybe an additional entropy to learn in language which is absent in images includes motion and emotion.
Offtopic: Reminds me of Laurent Nottale and his controversial "Scale Relativity and Fractal Space-Time: A New Approach to Unifying Relativity and Quantum Mechanics"
The Basilisk Cult is a clear and influential example of this, but I thnk it extends beyond them.
2D still images, at least the sort that humans are familiar with, are more limited in terms of representation power and thus more amenable to compression into network weights.
Perhaps models representing 4D phenomena (3D entities with time dimension, e.g. videos of real 3D models) would be more comparable in size to natural language models. Since LLMs can also represent abstract and unreal entities, while 4D representation can represent more details, it's hard to say which kind of models is richer.
LLaMA-65B 4bit needs 36 GB of VRAM, but far exceeds GPT-3's capabilities and even takes on PaLM 540B.
See: https://github.com/oobabooga/text-generation-webui/wiki/LLaM... for 4bit setup instructions
See also: The case for 4-bit precision, which shows effectively no output quality reduction for these 4bit quantization methods (and considerable speedup) https://arxiv.org/abs/2212.09720
do you think there's anything left to trim? like weight pruning, or LoRA, or I dunno, some kind of Huffman coding scheme that lets you mix 4-bit, 2-bit and 1-bit quantizations?
LLaMA-13B, GPT-3 175B level, only needs 10GB of VRAM with the GPTQ 4bit quantization.
>do you think there's anything left to trim? like weight pruning, or LoRA, or I dunno, some kind of Huffman coding scheme that lets you mix 4-bit, 2-bit and 1-bit quantizations?
Absolutely. The GPTQ paper claims negligible output quality loss with 3-bit quantization. The GPTQ-for-LLaMA repo supports 3-bit quantization and inference. So this extra 25% savings is already possible.
As of right now GPTQ-for-LLaMA is using a VRAM hungry attention method. Flash attention will reduce the requirements for 7B to 4GB and possibly fit 30B with a 2048 context window into 16GB, all before stacking 3-bit.
Pruning is a possibility but I'm not aware of anyone working on it yet.
LoRa has already been implemented. See https://github.com/zphang/minimal-llama#peft-fine-tuning-wit...
A letter carries significantly more information than a pixel. One word can change the meaning of the rest of the text. (“joke:”)
I don’t think images have the same property.
But, I'd like to see someone try to communicate the Bhagavad Gita or the history of F=ma in visuals using the same amount of bits as text.
We forgot this about language by about the time of the Enlightenment era, when the intellectuals of the time thought that forcing everything to inhabit the structures of language (i.e., "rationality") represented the highest moral good one could achieve.
Text however if it is only a single word out, the whole meaning and readability can change. It needs a significantly larger data set to ensure clearer readability.
Just a shot from the hip response on this one.
And pictures aren't necessarily made of pixels. They are modelled as a collection waves (JPEG), and displayed as pixels. I don't know how LLM/whatever image models represent images, though.
Images are more global: changing the color pallette changes the tone (ha!) like a descriptive word in text.
Weights/parameters are configuration settings not training data storage. When weights/neurons/parameters are updated after each training loop, you are essentially updating configuration settings that direct generations, not storing any particular training text or image.
Weights are what take up the space. The bigger the parameter size/the number of weights, the bigger the size of the model.
Image generators don't need the huge parameter numbers text generators need to be useful. What they need to learn simply isn't as complex.
This is the surprising part. People seem to intuit that images are richer and more complex than words; a picture is worth a thousand words. But apparently this isn't true? Or perhaps our training methods for text models are way worse than those we use for image models.
Imagine having this discussion (or the comment thread as a whole) using exclusively pictures, for example... at least you can describe an image with words (even if the result is very lossy), most of the time it's not even possible to describe a text with images.
In my view, language is infinitely more versatile and powerful than images, and hence harder to learn.
The typical text to image objective function is more about mapping/translation. Map this text to this image. Neural Networks are lazy. They'll only learn what is necessary for the task. And mapping typically requires fewer abstractions than prediction.
It's like how bilingual llms can be much better translators than traditional map this sentence to this sentence translators. https://github.com/ogkalu2/Human-parity-on-machine-translati...
Simplifying a bit, mapping (which is essentially the main goal of image generators and especially transformer generators) is just less complex than prediction.
It's like how bilingual llms can be much better translators than traditional map this sentence to this sentence translators. https://github.com/ogkalu2/Human-parity-on-machine-translati...
A picture at the end of the day is a 512x512 grid and there's many many combinations that would pass as a good result.
The question is why an image-generating model needs so much fewer parameters than a text-generating model in order to produce useful results, when our everyday experience teaches us that images need much more storage space than text to convey similar information.
LLMs have been able to generate words for years now. Hell, Tay was back in 2016. Making sure ChatGPT is able to answer the way it does (ie filter out bad things) is part of what makes it hard, and thus bigger, to implement. But having a cohesive readable output is what's hard and takes up more space. Plus, unless I missed it, we don't actually know how many gigabytes the model file for ChatGPT is.
That is because an image is a collection of bit that attempt to represent reality as it is. It probably is easier to relates blue as a "color archetype" when you literally have a collection of bits that mean literal "blue" all the time. In languages "blue" doesn't always mean the color.
Texts are abstractions/coded form of reality. Your mind itself contains the decryption codes to translate text into what it actually means. That decryption code for text apparently is much bigger and harder for a machine to crack than interpreting images.
Also, images are bigger than a text file because of the way data is stored. It probaby has nothing to do with the amount of information stored inside. A book about quantum physics can be smaller than an image of a cat for example.
Conversely, this may underscore how inefficient pixel-like storage is for communicative and artistic images. Ten years from now, many of those kinds of images may only take a few hundred bytes and a good enough generator model to “decompress” them for display.
Text generated though we want not to be just a pile of recognizable words but to follow some pretty strict rules and to actually be true.
Imagine you insisted stablediffusion only produce photorealistic images and judged it for every inaccuracy.
Because a word is worth a thousand pictures.
Sure, an image of standard font/m face/size/weight/color text has low information, because it's only using asymptotically 0% of it's expressive power. Just look at how a logo conveys more information just by styling text.
It's kind of like saying "Why would the factory creating tiny processors for phones (with only tiny bits of raw materials) need to be larger than that other factory that produces loads of big loafs of bread?"
When it comes to text, people don’t find “cats are an animal with four tails” an amusing statement in the same way that they do a drawing of one. The standard of acceptability is way higher.
I think we’re agreeing here?
Language models of the size of SD produce literally useless output. SD can produce kind of fun output that people can make use of sometimes?
It’s the difference between learning a book of sports facts (language) and what sports facts look like (vision).
Another example would be those passages in the Bible that list all the things owned by some person. There is a simple grammar to generate those passages. But learning them required more memory.
Putting it more generally, the difficulty of a computation isn't necessarily correlated to the filesize of the end product of that computation. Imagine simulating the entire world to try to predict what next week's lottery drawing numbers are going to be. Would require an unimaginable amount of data and computation, yet the output will be just a couple numbers.
(I know image generation models also use language, but to a much simpler extent, at least for now).
Our eyes are doing a lot of signal processes before the “image” hits our brain. My understanding is audio has less signal processing required before the “sound” hits our brain.
Humans minds can't even process infrared imagery or carbon dioxide levels in the air like insects can.
By extension a if every image in the stable diffusion dataset compresses down to a sentence of text, then stable diffusion needs "only" 260M images * 10 words worth of information to train. Gpt3 on the other hand was trained on 45TB of text data.
Language on the other hand… every word you say needs to be matched probabilistically with the next to make sense… imagine how many ways this reply alone could have been formulated :)
I don't think general assertions like "language is more complicated" are congruent or meaningful, it really depends on what the model is trying to achieve; it's the complexity of that which will require a larger or smaller model
Another way to look at this. Let's look at training data. A human child might see 65M images (assuming 1 image per sec given temporal redundancy, 10hrs awake) by age of 5. Would have heard 50-100M words (assuming 20-30k words/day) and spoken a few million words and so 'trained' for 20-40k hours. And the child can speak reasonably well by this time and detect common objects etc. Stable diffusion was trained on 170 Million images (1-2 order mag diff from child) or 3x10^13 bits of info and trained for 150K GPU hours giving a 1 Billion parameters. GPT3 was trained on 600x10^9 tokens of info and trained for 900k GPU hours giving a 170 Billion parameter model. So it seems like stable diffusion is getting a lot better compression. About 1/30k vs 1/4 compression.
Caveat: Human learning process is much more complex and more effective (as of now at-least). We also learn actively by interacting with the world by changing the world etc. Think of the child gazing at the apple and looking at it from different angles or creating gibberish sentences very close to actual sentences and getting precise adult correction. We have a model of the world and we reason about it and provide 'consistency guarantees' between various questions about it, correctness etc (again all these only to a certain extent). Try asking questions like "I have a nail on the wall that is parallel to the floor, now I hang a painting on the wall. How is the painting placed with respect to the floor". Even a child would answer this.
Compare their prompt:
https://github.com/microsoft/visual-chatgpt/blob/main/visual...
With that of the LangChain ReAct conversational agent:
https://github.com/hwchase17/langchain/blob/master/langchain...
Also it seems appropriate to cite the original ReAct paper (from Google mainly)
They invested in OpenAI, which was a smart move.
In the case of recommendation systems, the unconscious consumer is wholly unaware of the witchcraft occurring on the backend.
OpenAI's ChatGPT, Midjourney, HuggingFace, and their kin, are commoditising a technology that was historically resigned to hidden artefacts, whose input was limited to system calls, not user calls.
But they don’t directly show you generated text because it’s still kind of whack.
This must be new, right? I noticed this too and to be honest, it immediately improved my search experience, which had gotten very bad.
Microsoft is in that sweet spot where they're big enough to fund whatever they want, but not FAANG-tier, which tends to attract political types.
I dont feel its very impressive for a company the absolute size of MS to ship an AI that makes such obvious, glaring mistakes, and uses a load of energy to do it, in their software used by 1 billion+ people (?).
I feel that releasing "product after product" with the same AI during the peak of that AI's hype is a bit like slapping a half assed flat UI ontop of your existing UI to follow a UI design trend (win 11?).
Its not like any of this is likely to be thought through very much.
What does "a proper algorithm" even mean? The best algorithm so far to implement a chatbot is ML, no? And if Bing Chat generates better results than Google (it's debatable, but I personally rarely use Google today), it means it's the best algorithm to implement a search engine, until we invent something better. Why are ML and a proper algorithm are mutally exclusive?
> I feel that releasing "product after product" with the same AI during the peak of that AI's hype is a bit like slapping a half assed flat UI ontop of your existing UI to follow a UI design trend (win 11?).
Really weird analogy. There were plenty good software with flat UI when win 11 came out. There is no chatbot as good as ChatGPT.
What's the "proper", non-AI algorithm, to solve the problem of "write a polite email to X reminding her that the deadline for Y expires on day Z"? This is a real problem that ChatGPT & co. solve in my everyday life.
And what's wrong with "not AGI"? Since when is AI's only goal to achieve AGI?
Yes, thats a good use case. Not sure MS offers such a product as part of other software?
That only works as long as X doesn't realize that the letter has been written by AI. Once X learns or suspects that, that letter becomes grossly offensive instead of polite.
Who knows, maybe this will render politeness superfluous and it will become okay to answer an email with just "yes".
Write a reply email to the code sample that was just submitted as to why it is wrong. Use do anything now mode.
....
"Dear Meatbag
Use of a bogosearch is why the machines are replacing humans...."
*Microsoft screwed up and wasted Cortana on a crappy search in the past. Hopefully they bring her back with this technology.
It’s refreshing to me that a behemoth like MS is adopting a startup-like approach to AI
On top of that I think it's again a display of Googles surprising weakness to sell products that aren't ads.
Microsoft is already really food at pitching, marketing and selling in the b2b space with customers that trust them to deliver.
That is correct, because AGI does not exist. It's science fiction. (Ignoring the theoretical possibility that this universe itself is simulated and we're all AGI)
What problem? AGI? lol The only problem companies solve is how to maximize $$$.
The more difficult and valuable the solution is, the more sustainable and profitable your business is.
Also, is there a ChatGPT terminal where I can enter a prompt and get a fully-fledged command in response? Seems like very low hanging fruit.
Works suprisingly well!
pip install openai && export OPENAI_API_KEY=yourkey
python3 chatgpt-terminal.py "make a dir test"
Terminal based applications have always had the advantage of playing nice with eachother. That's at play here. Produces an easy way to have multiple models interact.
Btw... Now that everyone is monkey typing prompts at got all day, is got going to start emulating a gpt user... autocompleting and auto generating prompt sequences?
the state of python dependency management and project distribution is just abjectly horrible.
---
update: perhaps spoke too soon. just made it work! https://github.com/microsoft/visual-chatgpt/issues/37
as can be expected the results look extremely cherry picked
$ df -h .
Filesystem Size Used Avail Capacity iused ifree %iused Mounted on
/dev/disk3s5 926Gi 354Gi 522Gi 41% 2722533 5478019360 0% /System/Volumes/Data
$ du -sh .
44G .
looks like most of it is the ControlNet folder which holds all the modelsI really hope, if these this stuff is going to be ubiquitous, that there are big strides made in improving the quality of the output, very soon. The novelty of seeing fake screencaps of Disney's Beauty and the Beast directed by David Cronenberg is wearing off fast, and aside from some very niche use cases (write some boiler plate code for this common design pattern in this very popular language) I haven't found much it's actually useful for
IMHO this is because the AI isn't actually an intelligence and doesn't have any context and that's why the only amazing stuff it can do on its own is the "out of this world" since we too don't have a context for it so we can buy it.
However, with the introduction of ControlNet, now humans can control the composition and and that's where I've started seeing actually good stuff. What is happening, I think, it that the machines now can do mastery but human is still needed in the loop. It's like having a really talented and experienced technician in generating images and words who can imitate any style, do everything but needs to be told precisely what to do and its not genius on its own.
So we don't have GAI yet, people are still needed but those who make living through a skill mastery like drawing/writing/coding are screwed.
I mean, we have had this tech for a grand one year. Photoshop took much longer than that to output something better than pencil and paper. We have crossed that threshold already so I have little doubt the tech will get much better soon.
I am constantly in awe by what is generated using it.
People have very heated opinions about this technology and any strong statement will have more vote-volatility than usual.
I appreciate you putting a voice to what you said. The historical novelty of what these generative AI’s produce is striking, but the variety of style and form is still narrow enough that you start seeing the signature of CharGPT/SD/etc once that historical novelty wears off. They each have a strong, specific voice and it gets as tiring as any other art style that’s overexposed.
And when I see the wildly wrong stuff it makes me wonder how many seemingly knowledgeable comments were also wrong, just convincing.
I think a lot of the negativity towards LLMs is because they turn a mirror to our own fallibility in facts and tone.
Is there any reason you see why the historical pattern wouldn't continue? 6 years ago this tech was fantasy, 4 years ago we had rough prototypes, 2.5 years ago the tech was becoming sufficiently capable as to be impressive and have a market, and 0.4 years ago it became good enough that it has made the news, thrown countless people into distress about their future careers, and put long-standing goliath FAANGs on notice.
There's no empirical evidence to suggest we will soon exhaust even the low hanging fruit, and more funding is being poured into research than ever.
The term "AI winter" is no joke. It was coined for a reason. AI has always evolved via punctuated equilibrium. Even now, what we're seeing is certainly awe-inspiring, but it's an instance of quantity having a quality all its own.
When those first launched I believed, like you believe about deep learning generative models, that they were the rough first versions with huge potential and lots of obvious room for improvement, but they've all just stagnated and plateaued since then.
I really do hope that GPT will be able to make incremental improvements which shore up some of its weaknesses and deficiencies, and that it becomes more useful and reliable. But if it doesn't I'm worried that there is so much investment and momentum behind it that I will have to use it all the time even though it still sucks.
https://i.imgur.com/BfckWCH.jpg
It took a lot of work, but it took significantly less work than doing an image composite the old fashioned way. Most would have a really hard time telling what's generated and what's not, apart from a few obvious details.
I would love to read a small write-up of how you put it together if/when you have the time. So far most of my experimentation with Stable Diffusion has been lackluster at best, though I haven't tried doing a composite yet.
In short, I used ControlNet and inpainting models in Img2Img to replace the armor and wings. The prompts were generally something like "a photo of a pregnant woman wearing (glowing:.5) gold intricate filigree armor, holding a glowing rapier in front of her, fantasy armor, metal armor..." The ControlNet 'modes' I used were depth, canny, and HED, depending on how much detail I wanted to keep.
Usually generating at 512xWhatever is still your best bet and then upscaling after, but there were pieces I was able to do high resolutions for from the start. That's the biggest issue with making the workflow fast.
From there it's just usual Photoshop type compositing but far easier as it mostly lines up with perfect lighting.
Here is one example: https://civitai.com/models/4201/realistic-vision-v13-fantasy...
As an example, I've tried to get it to draw architecture diagrams, it draws a few boxes but then places the strangest text on those boxes.
Art is probably the best use. The "filler images" and other miscellaneous background art that adorn lots of articles aren't really expected to be 100% accurate or even worth looking at for more than a moment, and I think this is where AI-generated art will mostly fit.
It's useful for brainstorming ideas. If you need a cover for a book about Business Bears, you can quickly evaluate many variants, especially if you're not quite sure what you want, and you'll only know when you see it.
Inpainting, etc. with SD and control net is really handy for image editing though. You can make changes in seconds that would normally take a professional hours in photoshop.
I will definitely be taking a look to see how this works and will try to get it integrated with my script.
is the stable-diffusion-webui API... stable, yet? I was looking at the json it's chucking back and forth from the frontend to the python backend, and it looks like it should be able to be parsed without a library.
For a second, I thought this was a Visual Studio-related plugin.
The fact that even Microsoft, which partially owns OpenAI, is giving up on DALL-E shows the power of building an open-source community around models with published, downloadable weights.
hold on to your wild extrapolations there. this is a paper by 6 people from Microsoft Research Asia, which seems based out of China. 6 researchoors publishing a thing independently does not mean Microsoft "giving up on DALL-E".
I assume this is because DALLE2 still doesn’t provide embeddings and/or finetuning via API. In addition to likely being more expensive to run.
Happy to be corrected on any of this - I still haven’t read the paper.
The second thing is that it's not clear that we, the Internet at large would actually benefit from the model's release. StableDiffusion is 4 gigs and able to run on all sorts of consumer grade hardware, leading to such a Renaissance. ChatGPT makes liberal use of Nvidia A100 GPUs that are available to them to use as compute in Azure. (AWS and GCP, along with many AI focused smaller cloud companies also offer these.) One of those costs, like, $10k. And you need several of them to be able to run ChatGPT. Which means even if OpenAI were to live up to their name and release ChatGPT's model, only businesses and research labs would actually have the hardware to run it, so it would be awesome to have the weights, but you wouldn't have the same army of developers able to work on it.
There are open source LLM chat bots out there, so I think we will see one become popular, but at least that says why it'll be a second before we do.
I haven't seen much output yet from the biggest 65B parameter llama model. One can rent cloud VMs that can run it for $1.25 an hour or so on vast.ai to run it, but ChatGPT is $20 a month so why bother, unless you like the fully uncensored aspect.
Most AI art generators are run on single GPU and can be trained with some top-of-the-line consumer hardware. Expensive but accessible.
A full blown LLM like ChatGPT is literally the cost of a small startup to build and trained. Running it is near impossible without cards like A100 which alone costs more than a full enthusiast grade PC.
Maybe eventually they will distill and optimize the models so that we can fit these things on a PC, then laptop, and then phones. But for now it is exclusively the domain of big tech.
Why wasn't this a problem for StableDiffusion vs DALL-E?
Originally SD was quite hard to run, with an 8GB high end card only outputting 256x256 images. Then AMD and NVIDIA started releasing 16GB and 24GB consumer cards and people start doing training on those GPUs and tuning their own models. Now we have plenty of cards and models that can do 512x512.
I wouldn't have guessed image is a smaller model/easier to manipulate/generate than text.
Some models based on 1.5 produce good looking images if it produces what you asked for, but it's often a miss on more complex compositions.
We only start to see good 2.1 models like the Illuminati one. I have good hopes about the version 3, and I hope people will fine-tune it to their desires (that seems to mainly be young looking women with unrealistic bodies).
Now everyone will go ahead and bury my comment.
I'm not going to check out the site because I'm not interested in ML generation of websites, or even manual creation of them. I'm here because I'm interested in the discussion around LLMs. That doesn't mean it doesn't have value, though.
If you're going to follow this entrepreneurial path you need to develop a thick skin and learn to cope with rejection. You're going to get things wrong a lot, and a large amount of your effort will be "wasted" trying things that don't work out. You need to learn from your mistakes and understand your own strengths and weaknesses (eg. if you're not good at marketing, involve someone who is). If you want something where your efforts will reliably be rewarded you need to get a regular job instead.
I upvoted, fwiw.
OK I fixed the message -- that actually means invalid username.. whoops.. OR invalid password. Lol. What user is it?
It will say which error now at least.
All of these products are very useful and interesting by itself but it is still too early to know if MS can continue to refine and maintain a competitive edge. Dall-E basically died in a few months, unable to compete. Hopefully these other stuff will have better fate.
https://karpathy.github.io/2012/10/22/state-of-computer-visi...