Oh yes. Simple! Jesus, this ML stuff makes a humble web dev like myself feel like a dog trying to read Tolstoy.
Oh yes. Simple! Jesus, this ML stuff makes a humble web dev like myself feel like a dog trying to read Tolstoy.
Just like any discussion between advanced web devs would make any humble woodworker feel?
And just like any discussion between advanced woodworkers would make a humble web dev feel?
"It's really simple, they're just using a No. 7 jointer plane with a high-angle frog and a PM-V11 blade to flatten those curly birch boards, then a No. 4 smoother plane with a Norris-type adjuster and a toothed blade for the final pass."
Whut?
"You could use Webpack to bundle your HTML, CSS and Babel-transpiled TypeScript 5 down to shim-included Ecmascript 4", "They're just using OAuth2 authentication with Passport.js and JWT tokens, which easily gets you CSRF protection", "Our e-learning platform uses LMS.js and xAPI.js, plus SCORM for course packaging and Moodle as the LMS backend.", ...
There was a time you didn't know what any of that meant.
Just because you don't know what the words mean shouldn't make it sound difficult. Not saying AI is easy, just that the jargon is not a good indication of difficulty and we should know better than to be so easily mystified.
Great idea, actually. I do hope for a curriculum that enables kids on the trade school path to learn more about programming. Why not Master/Journeyman/Apprentice style learning for web dev??
I can't imagine starting out today...
For those that don't know this is from a show called Patriot.
Edit: ah I actually saw the prior scene where Leslie was explaining to John what he expected (which is the setup for the linked bit): https://www.youtube.com/watch?v=G7Do2tlYLhs
That said, when most people say differential equations they’re usually thinking of analytical solutions which is very much not necessary for practical ML.
Thank god.
This thing takes an image and creates a representation matrix.
> connect it to Vicuna-13B with a linear layer
Vicuna is an open LLM, pretty good quality, not as good as GPT3.5 though.
This is the beautiful part - a mere multiplication is enough to convert the image tensor to text tensor. One freaking line of code, and a simple one.
> and train just the tiny layer on some datasets of image-text pairs
You then get a shitload of image-text pairs and train the model to describe the images in text. But keep both the image and text model frozen. Is that hard? No, just flip a flag. So this "linear projection layer" (a matrix multiplication) is the only learned part. That means it takes less time to train, needs fewer examples and requires less memory.
Training the image and text models was much more difficult. But here we don't train these models, they use them as ready-made parts. It's a hack on top of two unrelated models, so it is cheap.
In the end the finishing touches - they label 3500 high quality image-text pairs, and fine-tune on them. Now the model becomes truly amazing. It has broad visual intelligence, and scooped OpenAI who didn't release Image GPT-4 in the APIs yet.
The important lesson to take is that unrelated models can be composed together with a bit of extra training for the glue model. And that open AI is just as powerful as "Open"AI sometimes. It's breathing down their necks, just one step behind. This model is also significant for applications - it can power many automations in a flexible way.
I thought they were creating image tokens based on the queries during finetuning and appending them to the language model. They are not text tokens.
The above means that the approach is web-dev like gluing, almost literally just,
from existingliba import someop
from existinglibb import anotherop
from someaifw import glue
a = someop(X)
b = glue(a)
Y = anotherop(b)I ask because we are about to produce a similar policy at work. We can see the advantages of it, but likewise, we can't have company data held in their systems.
I get the IP angst, but some companies think their GetGenericObjectFromDB() REST bs is secret sauce.
1. plant seed
2. ...wait a very long time...
3. observe completely unexpected but cool result
The unexpected part of step 3 is what makes this very different from any kind of engineering, even webdev.Of course, there is a lot of engineering involved in good ML, but that is more comparable to agricultural engineering in the sense that it's just a lot of dumb plumbing that any engineer can do without knowledge of the actual application.
As you get better at programming you have to take on harder problems to create the surprise of something working, because you gain confidence, and as you gain confidence, you start expecting your code to work. It's only when you've compiled the thing 6 times with small corrections and gotten segfaults each time and the 7th time you finally find the place you weren't updating the pointer and you correct it, but this is the 7th error you've corrected without the segfault going away, so you don't really expect it to fix the problem, but then you run it and it's fixed!
And then you get a job and the reality is that most of the jobs you're just writing CRUD apps and for a little while you can get some surprise out of learning the frameworks, but eventually you actually get really, really knowledgeable about the Postrgres/Django/React stack and nothing surprises you any more, but because nothing surprises you any more, you're really effective and you start being able to bill the big bucks but only for work on that stack because it takes time to struggle enough to get surprised, and the time that takes means your time is worth less to your clients. Money ruins everything. And if you don't do anything non-billable, it's easy to forget what programming felt like when you didn't know how your tools all worked inside and out. Not everyone takes this path but it's certainly the easiest path to take.
I think for a lot of folks who have been doing this for a long time, the reason ML is so exciting is it's getting them back out of their comfort zone, and into a space where they can experience surprise again.
But that surprise has always been available if you continue to find areas of programming that push you out of your comfort zone. For me it's been writing compilers/interpreters for programming languages. Crafting Interpreters was awesome: for the first time I benchmarked a program written in my language against a Python program, and my program was faster: I never expected I'd be able to do that! More recently, I wrote a generational GC. It's... way too memory-intensive to be used in my language which uses one-GC-per-thread for potentially millions of threads, but it certainly was a surprise when that worked.
Personally, I'm keeping track of ML enough to know broad strokes of things but I'm not getting my hands dirty with code until there are some giants to stand on the shoulders of. Those may already exist but it's not clear who they are yet. And I've got very little interest in plugging together opaque API components; I know how to make an API call. I want to write the model code and train it myself.
Becoming great at a particular technology stack means modelling it in great detail in your head, so you can move through it without external assistance. But that leaves an arena without discovery, where you just reinforce the same synapses, leading to rigidity and an absence of awe.
It's just a bunch of black boxes AKA "pure functions".
BLIP2's ViT-L+Q-former AKA
//I give you a picture of a plate of lobster it will say "A plate of lobster".
getTextFromImage(image) -> Text
Vicuna-13B AKA //I give you a prompt and you return completion ChatGPT style
getCompletionFromPrompt(text) -> Text
We want to take the output of the first one and then feed in a prompt to the LLM (Vicuna) that will help answer a question about the image. However the datatypes don't match. Lets add in a mapper. getAnswerToQuestion(image, question) -> answer
text = getTextFromImage(image)
prompt = mapTextToPrompt(text)
return getCompletionForPrompt(prompt)
Now where did this mapTextToPrompt come from ?This is the magic of ML. We can just "learn" this function from data. And they plugged in a "simple" layer and learned it from a few examples of (image , question) -> answer. This is what frameworks like Keras, Pytorch allow you to do. You can wire up these black boxes with some intermediate layers and pass in a bunch of data and voila you have a new model. This is called differentiable programming.
The thing is you don't need to convert to text and then map back into numbers to feed into the LLM. You skip that and use the numbers it outputs and multiply directly with an intermediate matrix.
getAnswerToQuestion(image, question) -> answer
text = getEmbeddingFromImage(image)
embedding = mapEmbeddingToInputEmbeddingForLLM(text)
return getCompletionForEmbedding(embedding)
Congratulations you now understood that sentence.More precisely - It gets the question After irs passed through a matrix that transforms the text description of the image so the LLM can “understand” it.
It maps from the space of one ML model to the other.
https://writings.stephenwolfram.com/2023/02/what-is-chatgpt-...
Maybe someone has resources to understand machine-learning on an ELI5 level.
I pick that up in above video and also in the post above.
Definitely healthy for him which just to be clear I’m a huge Wolfram fan and the ego doesn’t really bother me, it’s just part of who he is, however I do find it nice that LLMs are having him self reflect more than typical.
I find it funny how despite being completely uninvolved in ChatGPT he felt the need to inject himself into the conversation and write a book about it. I guess it's the sort of important stuff that he felt an important person like himself should be educating the plebes on.
Predictably he had no insight into it and will have left the plebes thinking it's something related to MNIST and cat-detection.
You just have to do it every day. It's fun!
If you can hold attention span over several days (I can't), work on a project bit-by-bit. Just make sure it uses modern AI stuff, and that you have smart people to talk around with.
this got a chuckle out loud from me. great visual.
ML is just a different field, using a different set of technologies from those you’re familiar with.