I've posted this before, but it seems like this genre is just getting more and more popular - and more and more untethered from any actual metrics of how good these models are.
I've posted this before, but it seems like this genre is just getting more and more popular - and more and more untethered from any actual metrics of how good these models are.
They are not comparable with late 2022 ChatGPT.
They will be able to shutdown your startup on a whim. Even if they didn't politicians and regulators would be a huge risk. Without democratization we get blade runner . Not that democratization has no problems, just that it is the way a lot of us are wanting it to go.
I get it, I understand why people like decentralization - but the open source community doesn't even have close to the capability to train an actually open-source LLaMA equivalent.
While those organizations might not be the right fit, I think they serve as an "existence proof" that a large scale open / nonprofit project is not completely inconceivable.
Could possibly see an industry consortium (of players too small to compete on their own) funding an open effort.
Last idea sounds crazy but hear me out: how much would Nvidia spending $100M on "open" models boost spending on graphics cards? I hope someone's running the numbers on that...
Probably not as much as in a zero-sum game where everyone is trying to train their own model. Every leap in CPU inference makes this an increasingly less appealing option for them.
But I agree, some industry consortium might try to do it. I think they would first have to be relying heavily on LLMs before they were willing to do that, and its possible by then that the lead will have gotten too large to easily surmount, especially now that everyone has stopped publishing.
It would be great if we could have detailed QA evaluation to show this, but of course then the open source people would train their models on it as a fine-tuning datasaet.
I mean… the horror.
langchain agents are a good starting implementation.
you can build your own prompt and get the ai to work by iself hallucinating tools, which may be cheaper to test out than going back and forth with an agent manager. not as accurate, but you can still extract useful work, i.e. https://i.imgur.com/AE4R3dR.png (gpt-35-turbo is traditionally failing this task completely, prompt get it to work at it)
these prompt all require the model to work off data within the prompt within the first shoot. model require a degree to introspection for that to work.
There's also already tools for conversation flows (which just means you prepend the conversation history to the prompt).
I'm not saying the performance is nearly as good, but the actual workflow does already exist and is massively improving. The interesting part to me is that this finetuning can be done in a couple few hours on a consumer gpu (4090).
Llama has the potential to reach ChatGPT it needs tunning to get better at responding to questions, llama if I am not worng is mostly attempting to predict what is next.
I can see it similar like Midjourney and Stable diffusion, midjorny can make any stupid prompt look like a digiatal art in the style of Midjorney but look how many stable Diffusion innovation happens, a competent person that is on top with all the new stuff can produce absolute anything in any style they want.
But I just don't see the need to mislead and suggest that these models are "on par" with ChatGPT or something like that. They just aren't.
Imagine what a math community could train, they just need access to the model and soem GUI software that can help them train.
So llama based Chat stuff is not yet comparable with ChatGPT but there are already lot of progress made. At this moment coding and math is bad in llama based but other stuff is great, like story creation, also I only could test 3-b 4bit and it is good enough to for example provide me a complex response in valid JSON format.
Then hardware will be a separate business. I think Apple might be caught off-guard by Nvidia on hardware. The latest NVlink and 400Gbps interconnects when combined with H100 next iterations and also rumors of advanced PCIe motherboards with high lane Nvidia CPUs and it looks to me that next year they can be selling $100-300K physical systems optimized for LLM inference that physically remind me of mainframes.
It is kind of idiotic that some scientist can spend years and a lot of public money to create some technology and then bilionairs are miliking all the profits.
I’m really hoping there are viable distributed and somewhat decentralized eventually consistent training algorithms we could all run in a P2P system. That would be super cool.
However I can easily see that now the framework has been established if a company builds a proprietary curated dataset for specific skills and then pays to spend resources on specialized reinforcement training.
Then they can commercialize that I would think. As people would pay for an LLM that does XYZ the best. Kinda like your Disney example but I was thinking engineering tasks in my head.