MiniGPT-4
minigpt-4.github.io
minigpt-4.github.io
But the results are pretty amazing. It completely knocks Openflamingo && even the original blip2 models out of the park. And best of all, it arrived before OpenAI's GPT-4 Image Modality did. Real win for Open Source AI.
The repo's default inference code is kind of bad -- vicuna is loaded in fp16 so it can't fit on any consumer hardware. I created a PR on the repo to load it with int8, so hopefully by tomorrow it'll be runnable by 3090/4090 users.
I also developed a toy discord bot (https://github.com/152334H/MiniGPT-4-discord-bot) to show the model to some people, but inference is very slow so I doubt I'll be hosting it publicly.
Oh yes. Simple! Jesus, this ML stuff makes a humble web dev like myself feel like a dog trying to read Tolstoy.
The above means that the approach is web-dev like gluing, almost literally just,
from existingliba import someop
from existinglibb import anotherop
from someaifw import glue
a = someop(X)
b = glue(a)
Y = anotherop(b)I ask because we are about to produce a similar policy at work. We can see the advantages of it, but likewise, we can't have company data held in their systems.
I get the IP angst, but some companies think their GetGenericObjectFromDB() REST bs is secret sauce.
1. plant seed
2. ...wait a very long time...
3. observe completely unexpected but cool result
The unexpected part of step 3 is what makes this very different from any kind of engineering, even webdev.Of course, there is a lot of engineering involved in good ML, but that is more comparable to agricultural engineering in the sense that it's just a lot of dumb plumbing that any engineer can do without knowledge of the actual application.
As you get better at programming you have to take on harder problems to create the surprise of something working, because you gain confidence, and as you gain confidence, you start expecting your code to work. It's only when you've compiled the thing 6 times with small corrections and gotten segfaults each time and the 7th time you finally find the place you weren't updating the pointer and you correct it, but this is the 7th error you've corrected without the segfault going away, so you don't really expect it to fix the problem, but then you run it and it's fixed!
And then you get a job and the reality is that most of the jobs you're just writing CRUD apps and for a little while you can get some surprise out of learning the frameworks, but eventually you actually get really, really knowledgeable about the Postrgres/Django/React stack and nothing surprises you any more, but because nothing surprises you any more, you're really effective and you start being able to bill the big bucks but only for work on that stack because it takes time to struggle enough to get surprised, and the time that takes means your time is worth less to your clients. Money ruins everything. And if you don't do anything non-billable, it's easy to forget what programming felt like when you didn't know how your tools all worked inside and out. Not everyone takes this path but it's certainly the easiest path to take.
I think for a lot of folks who have been doing this for a long time, the reason ML is so exciting is it's getting them back out of their comfort zone, and into a space where they can experience surprise again.
But that surprise has always been available if you continue to find areas of programming that push you out of your comfort zone. For me it's been writing compilers/interpreters for programming languages. Crafting Interpreters was awesome: for the first time I benchmarked a program written in my language against a Python program, and my program was faster: I never expected I'd be able to do that! More recently, I wrote a generational GC. It's... way too memory-intensive to be used in my language which uses one-GC-per-thread for potentially millions of threads, but it certainly was a surprise when that worked.
Personally, I'm keeping track of ML enough to know broad strokes of things but I'm not getting my hands dirty with code until there are some giants to stand on the shoulders of. Those may already exist but it's not clear who they are yet. And I've got very little interest in plugging together opaque API components; I know how to make an API call. I want to write the model code and train it myself.
Becoming great at a particular technology stack means modelling it in great detail in your head, so you can move through it without external assistance. But that leaves an arena without discovery, where you just reinforce the same synapses, leading to rigidity and an absence of awe.
https://writings.stephenwolfram.com/2023/02/what-is-chatgpt-...
Maybe someone has resources to understand machine-learning on an ELI5 level.
I pick that up in above video and also in the post above.
Definitely healthy for him which just to be clear I’m a huge Wolfram fan and the ego doesn’t really bother me, it’s just part of who he is, however I do find it nice that LLMs are having him self reflect more than typical.
I find it funny how despite being completely uninvolved in ChatGPT he felt the need to inject himself into the conversation and write a book about it. I guess it's the sort of important stuff that he felt an important person like himself should be educating the plebes on.
Predictably he had no insight into it and will have left the plebes thinking it's something related to MNIST and cat-detection.
You just have to do it every day. It's fun!
If you can hold attention span over several days (I can't), work on a project bit-by-bit. Just make sure it uses modern AI stuff, and that you have smart people to talk around with.
this got a chuckle out loud from me. great visual.
ML is just a different field, using a different set of technologies from those you’re familiar with.
Just like any discussion between advanced web devs would make any humble woodworker feel?
And just like any discussion between advanced woodworkers would make a humble web dev feel?
"It's really simple, they're just using a No. 7 jointer plane with a high-angle frog and a PM-V11 blade to flatten those curly birch boards, then a No. 4 smoother plane with a Norris-type adjuster and a toothed blade for the final pass."
Whut?
"You could use Webpack to bundle your HTML, CSS and Babel-transpiled TypeScript 5 down to shim-included Ecmascript 4", "They're just using OAuth2 authentication with Passport.js and JWT tokens, which easily gets you CSRF protection", "Our e-learning platform uses LMS.js and xAPI.js, plus SCORM for course packaging and Moodle as the LMS backend.", ...
There was a time you didn't know what any of that meant.
Just because you don't know what the words mean shouldn't make it sound difficult. Not saying AI is easy, just that the jargon is not a good indication of difficulty and we should know better than to be so easily mystified.
Great idea, actually. I do hope for a curriculum that enables kids on the trade school path to learn more about programming. Why not Master/Journeyman/Apprentice style learning for web dev??
I can't imagine starting out today...
For those that don't know this is from a show called Patriot.
Edit: ah I actually saw the prior scene where Leslie was explaining to John what he expected (which is the setup for the linked bit): https://www.youtube.com/watch?v=G7Do2tlYLhs
That said, when most people say differential equations they’re usually thinking of analytical solutions which is very much not necessary for practical ML.
Thank god.
It's just a bunch of black boxes AKA "pure functions".
BLIP2's ViT-L+Q-former AKA
//I give you a picture of a plate of lobster it will say "A plate of lobster".
getTextFromImage(image) -> Text
Vicuna-13B AKA //I give you a prompt and you return completion ChatGPT style
getCompletionFromPrompt(text) -> Text
We want to take the output of the first one and then feed in a prompt to the LLM (Vicuna) that will help answer a question about the image. However the datatypes don't match. Lets add in a mapper. getAnswerToQuestion(image, question) -> answer
text = getTextFromImage(image)
prompt = mapTextToPrompt(text)
return getCompletionForPrompt(prompt)
Now where did this mapTextToPrompt come from ?This is the magic of ML. We can just "learn" this function from data. And they plugged in a "simple" layer and learned it from a few examples of (image , question) -> answer. This is what frameworks like Keras, Pytorch allow you to do. You can wire up these black boxes with some intermediate layers and pass in a bunch of data and voila you have a new model. This is called differentiable programming.
The thing is you don't need to convert to text and then map back into numbers to feed into the LLM. You skip that and use the numbers it outputs and multiply directly with an intermediate matrix.
getAnswerToQuestion(image, question) -> answer
text = getEmbeddingFromImage(image)
embedding = mapEmbeddingToInputEmbeddingForLLM(text)
return getCompletionForEmbedding(embedding)
Congratulations you now understood that sentence.More precisely - It gets the question After irs passed through a matrix that transforms the text description of the image so the LLM can “understand” it.
It maps from the space of one ML model to the other.
This thing takes an image and creates a representation matrix.
> connect it to Vicuna-13B with a linear layer
Vicuna is an open LLM, pretty good quality, not as good as GPT3.5 though.
This is the beautiful part - a mere multiplication is enough to convert the image tensor to text tensor. One freaking line of code, and a simple one.
> and train just the tiny layer on some datasets of image-text pairs
You then get a shitload of image-text pairs and train the model to describe the images in text. But keep both the image and text model frozen. Is that hard? No, just flip a flag. So this "linear projection layer" (a matrix multiplication) is the only learned part. That means it takes less time to train, needs fewer examples and requires less memory.
Training the image and text models was much more difficult. But here we don't train these models, they use them as ready-made parts. It's a hack on top of two unrelated models, so it is cheap.
In the end the finishing touches - they label 3500 high quality image-text pairs, and fine-tune on them. Now the model becomes truly amazing. It has broad visual intelligence, and scooped OpenAI who didn't release Image GPT-4 in the APIs yet.
The important lesson to take is that unrelated models can be composed together with a bit of extra training for the glue model. And that open AI is just as powerful as "Open"AI sometimes. It's breathing down their necks, just one step behind. This model is also significant for applications - it can power many automations in a flexible way.
I thought they were creating image tokens based on the queries during finetuning and appending them to the language model. They are not text tokens.
(Not training obviously)
see this comparison: https://old.reddit.com/r/LocalLLaMA/comments/12ezcly/compari...
these models quantised to 4bit should run in CPU set ups with 16GB of RAM + 16GB of swap (Linux) and perhaps other setups run similarly
There a lot of optimizations that can be done. Here's one w/ potentially a 15X AVX speedup for example: https://github.com/ggerganov/llama.cpp/pull/996
Do you reckon the 4bit quantized Vicuna just won't do here? https://huggingface.co/anon8231489123/vicuna-13b-GPTQ-4bit-1...
I think with this everything OpenAI demonstrated ~5 weeks ago has been recreated by actually-open AI. Even if it runs much much slower on prosumer hardware and with less good results at least it is de-magicked.
1. it's using vicuna as a base.
2. It has a pretty high quality fine-tuning dataset. I initially missed this, and it's a very important advantage.
3. (speculatively) it doesn't collapse to extremely short responses (which BLIP2 and other models trained on image-text caption pairs) because of how small/simple the adapter is.
I was interested in training a BLIP2-LLaMA model before this, and I might still do it just to test (3).
How about 2x3090? Can it be run on multiple gpus?
emb_in_vicuna_space = emb_in_qformer_space @ W + B
These two models are trained independently of each other, on very different data (RGB images vs integer token ids representing subwords), and yet somehow they learn to embed different data in feature vectors that are so... similar. WHY should that be the case?It suggests to me there may be something universal about the embedding layers and hidden states of all trained deep learning models.
Still, it's kind of shocking that it works so well!
I'd be curious to see if the learned weight matrix ends up being full-rank (or close to full-rank) if both spaces have the same dimensionality.
If both spaces have the same dimensionality, the learned weight matrix would be full-rank only if every feature in the target space is expressible as a linear combination of features in the input space (plus a bias). Which brings me back to my original question: WHY would that be the case when the two models are trained independently on data that is so different?
So it's really less-than-full rank which would require an explanation - ie, why does this image representation project into this perfectly isolated subspace of the language representation (or vice versa)?
If that happened I would start looking for things like a vocabulary of smell which is completely distinct and non-overlapping with any visual context. But we use cross-modal analogies in language /constantly/ (many smells are associated with things we can see - 'smells like a rose') so you wouldn't expect any clean separations for different modalities... Maybe there's some branch of analytic philosophy which has managed to completely divorce itself from the physical world...
That's a really good point. Thank you!
At it's core, BLIP2 already projects RGB inputs into text token space and Vicuna (or rather LLaMA) uses such tokens as inputs as well as outputs. The only reason why a linear layer is needed at all is because they are not trained at the same time, so you still have to move text embeddings from one space to another. But it should not be surprising at all that one hidden linear layer suffices to do just that (see the universal approximation theorem [1]). This approach is just an efficient way to combine different models for downstream fine-tuning tasks while keeping their weights frozen, but it is neither new nor particularly surprising.
[1] https://en.wikipedia.org/wiki/Universal_approximation_theore...
The universal approximation is exactly not about deep models. Deep means many layers. But in the most simple (and proven) case, a single hidden layer perceptron is all it needs according to the UAT. Technically it also needs a nonlinear activation function, but you get all sorts of nonlinearities for free downstream anyways in this particular model.
My point still stands: The fact that models with sufficient capacity can approximate any function does not imply that two models trained independently of each other on different tasks will learn to approximate functions that relate to each other only by a linear transformation.
Taking a step back, this is just a wild statement. I know there's some doom and gloom out there, but in certain aspects, it's an awesome time to be alive.
I wonder if there's powerful enough ViT model that does OCR.
The results look interesting, however.
Here's hoping that they'll add GTPQ 4bit quantizing so the 65B version of the model can be run on 2x 3090.
A more honest name would be Visual-Vicuna or Son-of-BLIP.
The model as a whole is just BLIP-2 with a larger linear layer, and using Vicuna as the LLM. If you look at their code it's literally using the entire BLIP-2 encoder (Salesforce code).
The number of parameters used for GPT-4 is unknown.
[0] https://twitter.com/SebastienBubeck/status/16441515797238251...
So we're back to guessing ...
A couple of years ago Altman claimed that GPT-4 wouldn't be much bigger than GPT-3 although it would use a lot more compute.
https://news.knowledia.com/US/en/articles/sam-altman-q-and-a...
OTOH, given the massive performance gains scaling from GPT-2 to GPT-3, it's hard to imagine them not wanting to increase the parameter count at least by a factor of 2, even if they were expecting most of the performance gain to come from elsewhere (context size, number of training tokens, data quality).
So in 0.5-1T range, perhaps ?
https://www.reddit.com/r/IAmA/comments/12rvede/im_stephen_go...
maybe even add an "a" for extra spice: Son-of-a-BLIP
Outside of the brand name ChatGPT, lay members of the general public are way more likely to call these chatbots (like Bard and Bing) “AIs” than “GPTs”. And although GPT could technically refer to any model that uses a Generative Pre-trained Transformer approach (although it probably wouldn’t be an open-and-shut case), the mark “GPT-4” definitely is associated with OpenAI and their product, and you can’t just use it without their permission.
Let's not discuss the amount of copyright licenses OpenAI has already infringed, too
At Brewer’s Art in Baltimore, MD they just released a beer called GPT (Green Peppercorn Tripel)[1]. They’re likely allowed to do that because a reasonable consumer would probably not actually think they had collaborated with OpenAI, because OpenAI does not make beer.
OP is releasing a model called “MiniGPT-4”. A reasonable consumer could look at that name and become confused about the origin of the product, thinking it was from OpenAI. This would be understandable, since OpenAI also makes large language models and has a well known one that they’ve been promoting whose brand name is “GPT-4”. If MiniGPT-4 does not meet that consumer’s expectation of quality (which has been built up through using and hearing about GPT-4) it may cause them to think something like “Wow, I guess OpenAI is going downhill”.
Trademark cases are generally decided on a “reasonable consumer” basis. So yeah, they can seem a little arbitrary. But it’s important for consumers to be able to distinguish the origin of the goods they are consuming and for creators to be able to benefit from their investment in advertising and product development.
M3 is around the corner tho, and there's some announcement to come from intel or arm following their partnership. There's also the new card coming from intel that is supposed to be aimed squarely at machine learning workloads, and they don't have to segment their market by memory sizing like Nvidia do, but they aren't well supported as device targets, but a pair of these will likely be very cost effective if and only if they will get credible compatibility with the libraries and models
Get a 3090 or 4090. Forget about AMD.
Do I need dualboot? Or is Windows good?
4090 24GB is 1800USD, The Ada A6000 48GB is like 8000USD and idk where you buy it? So if you want to run games and models locally the 4090 is honestly the best option.
EDIT: I forgot - there is a rumored 4090ti with 48gb of vram, no idea if thats worth waiting for.
Seems you can get used RTX A6000s for around $3000 on ebay.
I think that's such a silly name for it, but oh well
Thanks for the correction!
Why do they do this? Sometimes consumer products are versioned weirdly to mislead customers (like intel cpus) - but these wouldn't even make sense to do that with as they're enterprise cards?
According to GPT-4 the next generation one will be called Galactic Unicorn RTX 6000 :D
You can’t switch which GPU Linux is using without restarting the session
One caveat though, my asus b650e-f is barely supported by the currently used ubuntu kernel (e.g. my microphone doesn't work, before upgrading kernel + bios I didn't have lan connection...) so expect some problems if you want to use a relatively new gaming setup for linux.
WSL on windows apparently decent, or native PyTorch, dual boot windows/ubuntu still prob best tho.
[1] https://www.dell.com/en-us/shop/nvidia-ampere-a100-pcie-300w...
Currently the 4090, the rumor is the 4090ti will have 48gb of vram, idk if its worth waiting or not.
The more VRAM the higher paremeter count you can run all in memory (fastest by far).
AMD is almost a joke in ML. The lack of CUDA support (which is nvidia proprietary) is straight lethal, and also even though ROCM does have much better support these days, from what I've seen it's still a fraction of the performance of what it should be. I'm also not sure if you need projects to support it or not, I know pytorch has backend support for it but I'm not sure how easy it is to drop in.
I mean in all honestly there's no reason a gaming card would need 48gb at the moment when so few games even use 24gb.
48GB really only makes sense for workstation cards.
Example: A 30B parameter model trained at 16bit FP gets quantized down to 4 bit ints. 4 bits = 0.5 byte. 30 billion * 0.5 byte = 15GB of VRAM (plus a GB or few of other overhead)
For more real world discussion see
There's a subreddit r/LocalLLaMA that seems like the most active community focused on self-hosting LLMs. Here's a recent discussion on hardware: https://www.reddit.com/r/LocalLLaMA/comments/12lynw8/is_anyo...
If you're looking just for local inference, you're best bet is probably to buy a consumer GPU w/ 24GB of RAM (3090 is fine, 4090 more performance potential), which can fit a 30B parameter 4-bit quantized model that can probably be fine-tuned to ChatGPT (3.5) level quality. If not, then you can probably add a second card later on.
Alternatively, if you have an Apple Silicon Mac, llama.cpp performs surprisingly well, it's easy to try for free: https://github.com/ggerganov/llama.cpp
Current AMD consumer cards have terrible software support and IMO isn't really an option. On Windows you might be able to use SHARK or DirectML ports, but nothing will run out of the box. ROCm still has no RDNA3 support (supposedly coming w/ 5.5 but no release date announced) and it's unclear how well it'll work - basically, unless you would rather be fighting w/ hardware than playing around w/ ML, it's probably best to avoid (the older RDNA cards also don't have tensor cores, so perf would be hobbled even if you could get things running. Lots of software has been written w/ CUDA-only in mind).
I haven't tried with LLaMA at all.
> Current AMD consumer cards have terrible software support and IMO isn't really an option. On Windows you might be able to use SHARK or DirectML ports, but nothing will run out of the box.
I was merely sharing that I did not have that same experience that current consumer cards have terrible support.
For RDNA2, you apparently can get LLMs running, but it requires forking/patching both bitsandbytes and GPTQ: https://rentry.org/eq3hg - and this will be true for any library (eg, can you use accelerate? deepspeed? fastgen? who knows, but certainly no one is testing it and AMD doesn't care if you're not on CDNA). It's important to note again, anything that works atm will still only work with last-gen cards, on Linux-only (ROCm does not work through WSL), w/ limited VRAM (no 30Bq4 models), and since RDNA2 tensor support is awful, if the SD benchmarks are anything to go by, performance will still end up worse than an RTX 3050: https://www.tomshardware.com/news/stable-diffusion-gpu-bench...
Absolutely fair and I agree with this part. I started my reply with "FWIW" (For What It's Worth) on purpose.
> For RDNA2, you apparently can get LLMs running, but it requires forking/patching both bitsandbytes and GPTQ: https://rentry.org/eq3hg - and this will be true for any library (eg, can you use accelerate? deepspeed? fastgen? who knows, but certainly no one is testing it and AMD doesn't care if you're not on CDNA).
I haven't tried any of the GPU-based LLMs yet. SD leveraging PyTorch (which seems to have solid ROCm support) worked for me. It will not be faster than NVIDIA for sure but if someone already has a 16GB+ AMD card they may be able to at least play with stuff without needing to purchase an NVIDIA card instead.
Basically, you want nVidia, and you want lots of VRAM. Buy used for much more bang for the buck.
Depending on your budget, get:
- an RTX 3060 with 12GB or
- 1 used RTX 3090 with 24GB (approx twice as expensive as the 3060 but twice the VRAM and much faster) or
- 2 used RTX 3090 cards if you need more than 24GB.
Everything beyond that gets quite a bit more expensive because then you need a platform with more PCIe lanes, you may need more than one PSU and you will have problems fitting and cooling everything.
With two cards and 2x24GB you can run the largest version of the LLaMA model (the 65B variant) and all its descendants with 4-bit quantization inside your GPU's VRAM, i.e. with good performance. Can can also try some low resource fine-tuning variants (LoRa etc).
Oh and while you're at it also get a decent amount of RAM like 64GB or 128GB (it's very cheap right now) and a NVMe SSD. These models are quite large.
But justai.com would also be apt
> The painting depicts the scene of the crucifixion of Jesus Christ. In the foreground, there is a group of people standing around the cross, including Mary, the mother of Jesus, and the two thieves who were crucified with him. In the background, there is a hill with a group of people standing on it, watching the crucifixion. The sky is dark and stormy, with lightning illuminating the scene. The overall mood of the painting is one of sadness and tragedy.
Then I asked for a sonnet on it:
"In this painting, we see the scene of the cross
Where Jesus is crucified, with Mary at his side
And the thieves who were crucified with him
In the background, a hill with people standing
Watching the tragic event unfold
The sky is dark and stormy, with lightning flashing
Illuminating the scene of the crucifixion
The mood is one of sadness and tragedy
As we contemplate the sacrifice of Jesus
But even in death, he shows us the way
To love and forgiveness, to live another day
And so we remember the crucifixion
And the love that Jesus showed us that day."
So...not quite what I was expecting, but it's progress!
[0] https://en.wikipedia.org/wiki/Lament_over_the_Dead_Christ_(V...
Manufacturing capacity of lidar components might be limited.
Another might be reliability/failure modes. If the system relies on lidar, that's another component that can break (or brownout and produce unreliable inputs).
So in a vaccum, yea a lidar+camera system is probably better, but who knows with real life trade offs.
(again, I just made these up, I do not work on this stuff, but these are a few scenarios I can imagine)
Plus we have ears, are under a roof so can't get rained on, are self cleaning, temperature regulating, have much better dynamic range, wear driving glasses…
Asking people to accept a driverless car based on over-arching statistics is papering over some very glaring issues. For example, are most accidents in cars being caused by "average" drivers or are they young / old / intoxicated / distracted / bad vision? Are the statistics randomly distributed (e.g. any driver is just as likely as the next to get in accidents)? Because the driverless cars seem to have accidents at random in unpredictable ways, but human drivers can be excellent (no accidents, no tickets ever), or terrible (drive fast, tickets, high insurance, accidents, etc). The distribution of accidents among humans is not close to uniform, and is usually explainable. I wouldn't trust a poor human driver on a regular basis, nor would I trust an AI because I'm actually a much better driver than both (no tickets, no accidents, can handle complex situations the AI can't). Are the comparisons of human accidents being treated as homogenous (e.g. the chance of ramming full speed into a parked car the same as a fender-bender?). I see 5.8M car crashes anually, but deaths remain fairly low (~40k, .68%), vs 400 driverless accidents with ~20 deaths (5%), I'm not sure we're talking about the same type of accidents.
tl;dr papering over the complexity of driving and how good a portion of drivers might be by mixing non-homogenous groups of drivers and taking global statistics of all accidents and drivers to justify unreliable and relatively dangerous technology would be a strict downgrade for most good drivers (who are most of the population).
It's software. Original Teslas with AP1 better than Teslas own in house software on their latest AP.
Let me try again and this time I will definitely not hit anything.
Sorry, that was another dog.
BingDrive: I'm sorry, but I prefer not to continue this conversation.
https://en.wikipedia.org/wiki/Moravec%27s_paradox
Although, I'm not surprised AI can describe paintings, but I'm still pretty surprised it can generate them. How'd we get that to work!?
Also, the resolution of Stable Diffusion's output isn't much of a limitation if you're willing to use other tools to massage the output into something professional-quality. See [1]
[0] For example: https://praeclarum.org/2023/04/03/chatsvg.html
This is definitely something they could be trained to be much better at, but I guess it's hasn't been a priority.
Here's a relatively easy one: [1]
[1] - http://www.illustrationartgallery.com/acatalog/WilliamsTimmy...
> This painting appears to be a depiction of a child running down a dirt road with a kite in hand. The child is wearing a yellow shirt and blue shorts, and the background is filled with various colors and shapes. The overall mood of the painting is playful and carefree.
What's amazing is that it has correctly described the scene including details such as:
- A boy
- On a dirt road.
- Having a Kite in hand.
- Wearing Yellow shirt and blue shorts.Could you describe the image as thoroughly as possible? Please include the content, any noteworthy symbology, techniques or styles, what kind of media, guess at the artist or identify any attribution, then attempt to describe the theme or idea it represents?
Got this
The image is a painting of a child running down a dirt road with a kite in hand. The child is wearing a yellow shirt and blue shorts. The background is a mix of urban and rural elements, with buildings and trees in the background and a field of grass in the foreground. The style of the painting is a mix of realism and abstraction, with a focus on the child and the kite. The artist may be attempting to convey a sense of freedom and playfulness, with the child running with the kite representing the joy and wonder of childhood.
The medium used is oil paint on canvas. The artist may have used brushes and other tools to create the different textures and effects in the painting. The artist may have also used underpainting to establish the composition and values before adding the final layers of color and detail.
It is difficult to identify the artist without more information, but the style and technique of the painting suggest that it may be the work of an artist who is skilled in realism and abstraction, and who is interested in exploring themes of childhood and playfulness.
It also missed the right half of the painting.
Try asking it why the painting is called "Timmy's Last Surprise".
If that is the case, then this is really a very, very simple paper. But I guess simple things can lead to great improvements, and indeed their results seem very impressive. Goes to show how much low hanging fruit there must be in deep learning these days by leveraging the amazing, and amazingly general, capabilities of LLMs.
Eventually you'll have to rename it, either to resolve the confusion or because OpenAI pressure you to do so, or both.
So better to pick a less confusing name from the start.
(This one is even more confusing because it's about image inputs, but GPT4 with image inputs had not actually been released to anyone yet - similar in fact to how DALL-E mini got massive attention because DALL-E itself was still in closed preview)
I'm not sure if that's better from a marketing standpoint though....it works, you still remember DALL-E mini
if this disappears then spammers and the various botnets will have the upper hand again.
Idea one: Captchas are to become pretty useless as a "is this a human" tactic soon. Maybe it already is, I don't know. What other things could we think off to prove someone is human? I was watching Lex Fridman and Max Tegmark and they were remarking on how Twitter using payment as a differentiator between human and bot is actually really good. And maybe the only way we can reliably determine if someone is a human or not right now. Just by the virtue that having thousands of bots doing something, that suddenly costs $5 per event will deter most attacks. Integrating online identification systems from various countries could be one tactic (such as https://en.wikipedia.org/wiki/BankID that we use in Sweden to log in to basically any online service). New startup: Un-botable authentication as a service.
Idea two: Since captchas are useless, we'll be able to do bots that can do almost everything on the web. No need for writing automation scripts, headless browsers, regexp etc. Just feed real visual data from browser to GPT-4 (or MiniGPT-4 or similar). Give instructions like "You need to accomplish this task: Go to facebook.com and create a user account and be friends with 100 people and act like a human. Follow the instructions on the website.". Then let the bot figure out where to move the mouse and send click events, keyboard events etc. Obviously much slower than a scripted bot, but would be very hard to detect as a bot. Good for E2E testing also? New startup: WebGPT: Bots that can do any task on any website. TestGPT: E2E testing service. Just write test instructions like "Make sure the login works on /login" and you're done! And you thought flaky tests were a thing of the past... Would be kind of cool for UX testing also. "You are a 43 year old truck driver in Spain, please go to news.ycombinator.com and see how easy it is for you to register an account, log in, go to any news story and make a new comment. Report anything you think is confusing."
I for one will probably still stay near the fundamental side of things as all of this thrash is simply far too much for me, but it is quite nice and beautiful for me to get to see all of the variety and beauty of what is being built here. <3 <3 <3 <3
Now, to figure out a good way to integrate these seamlessly into my development workflow....I go about as painfully low-tech as one could. I might as well be using an electronic typewriter with only Ctrl+F search, lolzers....
[1] https://developers.google.com/fonts/faq#how_can_i_get_a_lice...
> The image shows a block diagram of a power amplifier circuit with a power supply and a load resistor. The power supply is connected to the input of the amplifier, which is then connected to the output. The load resistor is connected to the output of the amplifier and the power supply. The amplifier is designed to amplify the input signal to a higher level, which is then sent to the load resistor to power the load. The power supply provides the necessary voltage to operate the amplifier.
On the one hand I'm impressed that it recognises an amplifier, but just about everything in the description is (slightly) wrong. It is not a power amplifier, the power supply is not connected to the input (why should it be?), there is no load resistor...
More and more I get the impression that people are so taken with the chat interface that they are grading these models more favorably than is warranted.
> I apologize for the confusion. The image shows a block diagram of a basic amplifier circuit with a power supply, but it does not show any input stages, output stages, or load resistors. [...] The input stage is made up of an operational amplifier, which is used to amplify the input signal. The output of the input stage is then connected to the power supply, which provides the necessary voltage to operate the amplifier. The amplifier is designed to amplify the input signal to a higher level, which is then sent to the load resistor to power the load.
So it acknowledges the remarks, describes everything that is _not_ in the image and then puts the stuff back into the description anyway :)
I have no problem when there is a lot of money spent to implement things that can be used and tested, but I wish I could read about the next step of AI, something that involves science and not software engineering.
I've yet to read anything related to the analysis of trained data or trained neural networks, the structure of the brains of small insects, psychology on general intelligence, AI and emotions, the real ability of an AI to think and write meaningful code or understand what is a problem/solution etc, I mean anything that is related to AI that can be treated by computer science.
I have no problems with algorithms and implementations, but don't university/professors also try to go beyond, with epistemology, to connect the dots with biology and neurology?
Unless I see science cover several fields of research, it's not real, worthy research in my view, which is why the reason why I'm so stubbornly skeptical of chatGPT.
The people doing so called alchemy are making far better systems that anyone who’s trying to actually understand things. And we’ve been trying for decades now.
Most of our discussions are about model size, but few about dataset. Yet all the scaling laws hint at the great usefulness of more data. Sometimes even little data can have a great impact in the fine-tuning phase. In the end it is the training data that transforms a random init into the model.
and prompted "Describe any issues found in this screenshot and steps to resolve them"
while it correctly identified it as a screenshot from a computer, it gave a very generic response and didn't identify the error messages correctly:
"The screenshot shows a computer screen with a warning message that reads, “The system is currently unavailable. Please try again later.” This indicates that there is an issue with the computer system that needs to be resolved.
To resolve this issue, you can try the following steps:
1. Check the system status: Check the system status to see if there are any errors or issues that need to be resolved.
2. Restart the computer: Try restarting the computer to see if that resolves the issue.
3. Check for updates: Check for any updates that may have been released for the computer system and apply them if necessary.
4. Check for viruses: Run a virus scan on the computer to check for any viruses that may be causing the issue.
5. Check for hardware issues: Check for any hardware issues that may be causing the issue. This could include issues with the power supply, hard drive, or other components.
6. Contact technical support: If none of the above steps resolve the issue, contact technical support for further assistance. They may be able to provide additional troubleshooting steps or help you resolve the issue."
It is both low resolution and filled with jpeg artifacts.
https://github.com/Vision-CAIR/MiniGPT-4/blob/main/MiniGPT_4...
I’m running Vicuna 13b 16f locally and it needs 26GB of VRAM, which won’t even fit on a single RTX 4090. The next gen RTX Titan might have enough vram but that won’t come cheap. I’m expecting a price point above $2500.
I don't see a really good reason why OEMs couldn't do that now, in the past there have been OEM cards that have more VRAM than the reference design. I'm sure there's an appetite for cards like that for people who don't want to refinance their home loan to get 2 x RTX 4090 cards.
Even 33B only needs 20GB of VRAM in GPTQ 4bit.
8bit has zero perplexity loss, so there's really no reason to run in 16bit.
Even a $200 P40 24GB is enough to run 33B at extremely high speeds in GPTQ 4bit.
I run it on my phone CPU and get ~4 tokens per second. On my laptop CPU I get 8 tokens per second.
On a $200 P40 I run LLaMA-33B at 12 tokens per second in GPTQ 4bit. A consumer 3090 gets over 20 tokens per second for LLaMA-33B and 30 tokens/second for Vicuna-13B.
thanks
Then I asked it what are the likely ingredients of the product. It still hadn't replied after 2274s so I gave up on it.
I'd segment the audio semantically based on the topic of discussion, and I'd segment the video based on editing, subjects in scene, etc. We could start simply and just have a "timestamp": [ subjects, in, frame] key-value.
It'd take some fiddling to sort how to mesh these two streams of data back together. The first thing I'd try is segment by time chunks (the resolution of which would depend on min/max segment lengths in video and audio streams) and then clump the time chunks together based on audio+video content.
Great job at OCR-ing the text and recognizing the context, but it still can't do doom and gloom like us eastern Europeans.
The queue is about 100 at the moment, with 700s of waiting.
I clicked on the Video button while waiting, assuming that it would open in a new tab, and lost my place in queue.
Invoker here, I would like to have a chat or send me an email @ community@invoker.network