Nvidia releases NVLM 1.0 72B open weight model
huggingface.co
huggingface.co
Also they seem to only train on publically available data, concluding that quality is more important than scale.
The cc-by-nc-4.0 license applies to the network weights. The only thing non-commercial about the license is that it restricts how you may reproduce the licensed material:
> reproduce and Share the Licensed Material, in whole or in part, for NonCommercial purposes only; and
As long as you are not selling the network weights themselves, nothing in the license prevents you from evaluating the neural network for commercial purposes and selling the outputs. In 'production' you will have to directly download the weights from Nvidia themselves (or another 3rd party which is distributing the network weights non-commercially in good faith) though, you can't share the network weights onto your commercial inference server from another one of your commercial deployment servers. Or at least, it gets more dicy there and may be considered commercial reproduction so better avoid it.
For similar reasons you may 3D print a CC-BY-NC model of a tool and use that tool in your commercial workshop, you may use a CC-BY-NC compiler of a language to compile commercial programs, etc.
I'm not even sure if network weights are copyrightable independently of the code and data used to generate them. In my personal (not a lawyer) view, the weights of a neural network are the product of a mechanical transformation process much like a compiler or assembler, and we don't consider a compiled binary to have a copyright independent of its source code.
I still wouldn't notoriously try to violate a purported weights license, mind you, both because it's rude to ignore the authors' wishes and because it would not be fun being used by NVidia or any other deep-pocket AI company.
Creative Commons themselves write at https://creativecommons.org/faq/#can-i-apply-a-creative-comm... :
"Can I apply a Creative Commons license to software? We recommend against using Creative Commons licenses for software. Instead, we strongly encourage you to use one of the very good software licenses which are already available."
Of course, LLM weights aren't traditional software...
The problem is if you happen to sign any agreement with NVIDIA in order to get the weights. The problem is whatever contracts you may be bound by.
Can't this flip on a dime and a billion dollar company lose billions?
Allowing the language model weights to be updated during training could potentially result in better performance on both tasks, though, if Nvidia's result replicates. I could believe that it might: after all, more diverse data is more diverse data, and the model will be forced during training to generalize more.
I see it playing out one of two ways. Either Nvidia are selling shovels in a gold rush, the rush will end, and the business will dry up (after they have made a lot of money!). Or AI sticks/takes off, and Nvidia are selling a commodity too far from the value, like most electronic component manufacturers, and they'll maintain significant market share but have their margins reduced to a fraction of what they were before (after they made a lot of money!).
The human value doesn't come from ML training or inference, it comes from taking a better photo. The business value comes from drafting a better email. Those companies closer to that value will likely do better in the long run, as they always have done.
Midjourney is profitable. All the acquired startups (i.e. Streamlit or MosaicML) who made millions per employee "made money" for the people who cared.
OP was likely talking about profitability.
FWIW I wouldn’t really count streamlit as an ai company
Also Klarna threw out 700 people, they probably make money with AI.
And i found this article: https://www.ft.com/content/a9a192e3-bfbc-461e-a4f3-112e63d0b...
There was never ever any technology like LLMs close to what chatgpt and co can do in regards of understanding random human input.
My startup doesn't need to make money with it directly, but for us it increased our data quality on text and images.
I'm also quite happy to pay 10-20$ per month for random things LLMs do quite well for different use cases like creating some scripts etc.
There is a reason why CUDA works on every NV gpu but ROCm support is spotty at best and only guaranteed on data center GPUs.
AMD and Intel insisted on selling only flimsy garden shovels.
I'm saying that Intel and AMD made single-purpose GPUs useful only for graphics. Whether that's because of the software or hardware is immaterial. Effectively, it's one product in the same sense that an iPhone is one product to a consumer, but technically it's the iPhone device + iOS the software + Apple services such as iCloud, music, etc...
The distinction is one of business strategy not technology.
only exception im excited about is the non-main characters from video games, where a lot of the random NPCs, can now actually bring some more fun to the game.
And translation and grammar/spell checking is also at a level which was unthinkable before LLMs hit.
But thats it, really. The "talking machine" aspect of it is more and more uncovered as totally useless.
you built a robot that sorts laundry? Tell us more!
Do you have any more info or links about the setup?
I would have answered earlier, but the silly HN rate limiter prevented me from passing the link to you.
I dont want to look it up yet agan.
And I dont want to use HN anymore,, this rate limit time-waster really just killed my sympathy for this site.
Sorry if I came off negatively.
I think many underestimate the true usefulness the current generation of AI has already achieved because a lot of it is in traditional, boring, bespoke or inhouse LoB systems whereas the press always focuses on public B2C
I used chatgpt 3 days ago to generate a script for me. Saved me probably an hour too.
We use it also in my startup for tasks which we wouldn't even tried without ML models because the quality of old libraries were to bad. Like pdf catalog to text, image classification and segmentation.
Simplified for posterity:
kv_bytes = kv_bits / 8
hidden_per_head = hidden_size // num_attention_heads
total_heads = hidden_per_head * num_key_value_heads
kv_bytes_per_token = 2 * kv_bytes * num_hidden_layers * total_heads
(Edit: I accidentally swapped in some of the vision config bytes in my original calculation; these are the corrected numbers.) So, for NVLM 1.0 72B, that works out to 640kb per token assuming FP16 KV cache. If you use the entire 32k context length, that's an extra ~20GB of overhead for the KV cache. Then depending on how you're running the LLM, there might be extra overhead e.g. compiled CUDA graphs.You can cut this down lower by using grouped query attention as described here: https://medium.com/@plienhar/llm-inference-series-4-kv-cachi... This allows you to divide that number by the number of grouped heads, although it trades off accuracy for VRAM usage.
But TLDR, a minimum of around 164GB of VRAM at full accuracy. To me that seems fairly low, and I think vLLM would OOM without significantly more than that, but that's about as low as you could go in theory if you're running everything at FP16. Half that, of course, for FP8.
You'll typically need to have a copy of the KV cache per GPU, if you're using multiple GPUs, so multiply the KV cache overhead by the number of GPUs you're using. This will depend on what the specs for the GPUs you're using are; for example, you'll need 3 H100s (really four, since vLLM wants the number of heads to be evenly divisible by the number of GPUs); if you're using L40Ses, you'll need eight of them; but most likely only a single AMD MI300x.
Edit: specifically ocrbench and VQAv2
I wonder if one of the reasons they released it was to respond to OpenAI's plans to enter the chipmaking market.