Training LLMs from ground zero as a startup
yitay.net
yitay.net
So it seems like you have the various LLM creators all doing roughly the same sort of thing (training with text and image data) with similar hardware and similar data. Each of these naturally has their own brand of "secret sauce" for distinguishing their venture. The various secret sauces can make a difference in the quality of an LLM's output.
Yet overall, this seems like a massive, energy intensive exercise in redundancy.
This is commonly refered to as a market working as intended. Yes, the waste from this type of redundency can be massive, especially if you realize that ultimately just a tiny percentage of these efforts will result in even moderate success. But it is the price to pay at the edge of progress. A planned monopoly might be more efficient (despite popular banter that just compares a megacorp or a gov, which is basically the same, to a single succesfull startup ignoring the 999 that tried and failed), but those seldom beat a market on innovation.
Is it? Seems like market is unable to separate wheat from the chaff and is just throwing money around hoping to hit the jackpot. While AI has massive chance of affecting our lives, the investment market paints a pretty similar picture to what happened during the crypto boom.
Still have many teams trying to achieve the goal, but prevent corporate secrecy - effectively allowing competitors to look over each others shoulders and copy good data and ideas.
Such a system probably wants to compensate those whose ideas were copied, but that isn't strictly necessary - another approach is to simply make it illegal not to share data/results. Your compensation is your freedom from prison.
Golliath 120b is still the best open source model and no one knows why since it's just two llama2 60b glued together.
Different people have different goals, and they don't necessarily align with yours.
In the grand scheme these alignments are harmful as they place a reality distortion field. Authors create model of what language is and then contort that model to fit an opinionated idea of what language should be. Smells a bit Orwellian, right?
No, seems perfectly fine by me. You are already shaping your results by your selection of training data. Eg do you want to train a model that speaks English, or German, or both? Do you want to run your training data past a spam filter first? Do you want to do a character based model, or one of those weird encodings that is popular with LLMs these days?
Doing some other procedures afterwards to make sure your LLM doesn't say embarrassing things is small fries by comparison.
Also it's good practice for trying to get alignment with more important values (like "don't kill all humans") later when models might get powerful enough to be able to kill all humans.
Playing some little games where OpenAI tries to keep you from making their model say embarrassing things, and people keep trying to make it say embarrassing things, is a good low stakes practice ground.
A GPT that hasn't been aligned does not work how we expect - you give it a prompt, and it will autogenerate until it reaches an end state.
To even make the GPT answer the question in the prompt, and not autocomplete it into nonsense, is an example of alignment.
It took a lot of fine tuning and data curation to get ChatGPT up to its current chat-like interface.
But this is not the only alignment you can do. The original Transformer paper was about machine translation, turning the prompt into the translated text. Once it was done it was done.
We could choose to have the model do something else, say translate the prompt into 5 languages at once instead of one, just as an example. This would be another alignment decision.
There is nothing political or selection bais or anything inherent to the original definition, its only recently "alignment" has morphed into this "align with human morals" concept.
Even in the Andrej Karpathy's build-your-own-gpt YT video, which is highly talked about around here, he uses the phase like this. The end of the video you are left with a GPT, but not a question-and-response model, and he says it would need to be aligned to answer questions like ChatGPT.
Keep in mind that this is also chaff to distract people from the real secret sauce. I imagine that just as many startups are hiring writers and photographers to create extremely well labelled uncontaminated data for training.
One only need to look at the perverts over at civitai to see how far you can go with intensive labeling on a tiny compute budget.
our conversation was recorded here https://sub.thursdai.news/p/thursdai-feb-15-2024-openai-chan...
no events planned near term but come to the big shindig in june https://ti.to/software-3/ai-engineer-worlds-fair . last year's summit was the first time i really understood how much of a reach we have and how many good AI people we've managed to gather as friends.
[1] https://hn.algolia.com/?dateRange=all&page=0&prefix=false&qu...
This post gives a lot of credit to Google's infra and hardware teams, and I'd love to read a perspective from one of those insiders who then went on to do related work elsewhere.
> I was completely taken aback by the failure rate of GPUs as opposed to my experiences on TPUs at Google
Should be "I was completely unaware of the failure modes of GPUs, because all my career I've been inside Google and used Google TPUs and was well-acquainted with those failure modes."
I've used GPUs mostly, and when I tried TPUs the jobs failed all the time for really hard-to-debug reasons. Often the indirection between the x86 chip and the TPU device caused hours of hair-pulling, stuff you never get with x86+nvidia+pytorch.
10-15 years ago, Google minted many $10m+ data scientists (aka Sawzall engineers) who also ventured "into the wilderness" and had very similar reactions. This blog post is much more about the OP hyping his company and personal brand than contributing useful notes to the community.
In my humble opinion, we never had failures of GPU even for large scale training. Our current training batch job is a 20GB json file which takes 6 hours just to load and has been running for more than 15 days with not a hiccup. And we are using the older Tesla T4.
GPUs have memory constraint issues but if you can plan and work around it, I havent seen it crash in real life.
That's an undemanding and well-debugged chip by this point (6 years ago!). So you aren't experiencing any of the pain people using A100s or H100s (never mind people who have to stand up clusters with B100s soon) are going through now.
Gwern(or anyone else) do you have any resources on this?
Err you definitely should be doing something about that.
20GB on T4s (how many?) isn’t really comparable to terabytes on thousands of A100s.
The main Reka.AI page looks like a regular ChatGPT clone, an LLM you pay for by the token. How is this different from all these other companies? Pricing seems to be comparable to ChatGPT 3.5-Turbo.
The world of LLM startups is beginning to look like the world of hedge funds and private equity firms - where the prerequisites for seed/funding are:
A) Prestigious employment history / correct pedigree.
B) Solid network of investors ready to jump before any product has even begun.
At least until compute cost drop to a cheap enough level...
I wonder what was the budget spent for the chips/cloud GPUs to achieve GPT 3.5 level LLM - at least in the order to magnitude - 2-5 millions?
It is a perfectly acceptable use of the idiom.
2. How do you think idioms are created in the first place?
3. What exactly forces you to act like this?
Google's systems are reliable because of the tens of billions of dollars that Google has invested into developing datacenter hardware, software, and processes over 25 years. Highly-competent teams at smaller and less-mature organizations will always deliver a much worse product.
Another thing to consider is priorities. Google prioritizes reliability. They retire parts that fail repeatedly, even if the failures are relatively infrequent. Smaller and less-sophisticated datacenters keep parts in service even with frequent failures, or don't even monitor failure rates of certain parts. Smaller datacenters buy and use Google's old parts and unreliable parts.
Therefore unreliable machines does not imply anything about the competency of the hardware team.
If the low reliability of the hardware is making your work slow, then how about improving the software so it can tolerate the unreliable hardware, or switching to a more reliable (more expensive) hardware provider?
> Thankfully, I (and many of us in the team) have built up this intuition quite a bit in our ML careers to get it right within a substantially short amount of tries. While we’ve trained really good models before in our previous jobs, differences in training infrastructure, data, incorporation of new ideas and other environmental issues can still cause non-trivial differences in outcomes. That said, a strong prior helps to significantly cut down the search space and is probably one of the easiest explanations to why we were able to train really strong models with so few trials, resources and experimentation.
By the time you get it all working, not only you've spend lots of your VC capital on training alone, your competitors (Google, Meta, etc) already released a more powerful model much better and quicker than you before you could your run the second training epoch.
Another example of a startup incinerating the VC pump and dump scheme for vaporware AI snake-oil.
(GIGO is what one gets when feeding LLM with "G"arbage "I"n, "G"arbage "O"ut.)
This is the current problem about making a vaccine signature fitting like a glove ... as tight as possible ... when populating the anti-malware (i.e. IDS/IPS/NDS/XNS) search pattern engine for use by Aho-Corasick-variant algorithms (such as Parallel-Failureless Aho Corasick).
However, LLM as a binary code-based detector for malware detection has a very limited benefit (it is there but only as a backend topical add-on after all other conditionals have been identified).
LLM lacks qualifying conditionals surrounding a premise data, and I have my doubts of using LLM for medical diagnosis as well: until we start having LLM denote the much-needed weighted combo-conditionals by "percentages".
- signed, a compute resource hog at FAANG
Haven't worked at Google, anyone else share this sentiment? I always feel like working with Google code is typically not idiomatic and super difficult to go "under the hood" if anything isn't precisely on the happy path.
Also, most work seemed to involve some balance of junior and more experienced people, which helped keep quality higher. Outside of Google, I've seen pretty large projects written by new grads with little supervision (and on a tight timeline). Those codebases can be pretty hairy.
It lets them extract usable SWE horsepower from pretty much anyone who steps inside and at least tries to be useful and not just coast. They can ingest a startup engineer, someone who's been a mid-tier enterprise codemonkey, yr mythical 10xer, the whole statistical gamut.
@resource0x in a sibling comment made the point that it's possible to write great code even if the program is a flawed design. I'm probably conflating those things.
Rumors of Google’s AI code superiority are vastly overblown in 2024. I’m currently at another major AI lab, and the code here can actually be understood and worked on, which I consider to be a massive advantage.
Google has superb robustness and code quality, with garbage-level usability. Once you're setup, you can kick off many massive training jobs and compare results easily. However, getting to that point is really hard. You'll never figure out how to use the ML infrastructure and libraries on your own. You can only get it to work by meeting with the teams that wrote the infra so they can find and fix every error and misconfiguration. Usually, there is one single way to get things working together, and neither the documentation nor the error messages will get you to that brittle state.
It's near impossible to get a VM with a TPU or GPU attached, so there's no way to debug issues that happen between the library and the accelerator. Plus somehow they've made Python take longer to build (??!!) and run than C++ takes, so your iteration cycle is several minutes for what would take seconds at any other place. Fun stuff! Somehow it's still one of the best places to do ML work, but they sure try to make it as difficult as possible.
I worked there, and the quality is definitely much higher and the code tends to be far more maintainable. However, there is often a cost for that, which is velocity.
Some of this is reduced by the sheer amount of automation in tooling (i.e. bots that block style violations and common bugs before a code change is submitted).
In other cases, it slows things down quite a bit.
Google's codebase is idiomatic to Google due to their strict language tooling. e.g. their C++ code stays away from advanced features. The tooling teams at Google have very strong say.
I wouldn't be surprised if it was version 2.7 too...
Tldr: there’s multiple factors to consider here and it’s more interesting to understand the pressures that cause the decisions, especially if you want to try to create a world where different decisions are made.
Outside specific cases around machine learning, it’s really not: Go is that language. It’s not like each of those platforms doesn’t have to have a similar team that understand Go anyway (for their SDK), so they could save their customers the abject pain of Python dependency management by just writing their CLIs using it.
For anything more advanced they offer language specific SDKs in Rust, Swift, Kolton, etc…
For example integrating storage in an iOS app.
EXTREME predictability (e.g. as never ever using the system's libssl), in trade for huge binaries. They go pretty damn far in this: you won't catch a Google binary even using most of libc.
Which honestly is a GOOD thing because it would make it much easier for newcomers to ramp up on existing codebases. Most people aren't used to working with spaceships and constexprs.
Readability is also far more valuable to a large team than efficiency for anything that isn't a number-crunching loop.