A BERT for laptops, from scratch
github.com
github.com
Edit - for anyone unsure about what "BERT" is or its relevance, it's a transformer based natural language model just like GPT. However, where GPT is used to generate text, BERT is used to generate embeddings for input text that you can then use for predictive models (e.g. sentiment prediction), and that process is also demonstrated in the notebook.
*Edit 2 - The 17 hours are pretraining only, not including the time to train the tokenizer, or finetuning.
A few years back cough I upgraded gcc from 3 to 4 and emerged system and then world. That was over 1200 packages. It took about a week. That was in the days when I used a Windows wifi driver and some unholy magic to get a connection. I parked the laptop on a table with the lid open and two metal rods lifting it up 6" for airflow. I left the nearby window open a bit too.
(I have to admit I stopped using Gentoo mostly because it encouraged me to endlessly fiddle with my system, and it would invariably end up broken somehow. That's entirely my fault, and not Gentoo's. I switched to Archlinux as my distribution of choice, and I manage to hold myself back enough not to destroy my installation.)
Beyond that the big win over performance even is customization. The most secure code, and the fastest code, is the code that isn't there at all. Do you really need your entire system to support LDAP authentication? Maybe... What about your local email daemon? Do you need that? Because cron does and since your mail daemon also has MySQL support built in, installing cron gets you the MySQL libraries.
I don't use it anymore because of the overhead, but there are a lot of performance and security benefits to be had there.
> Realistically, though, most software that can benefit from specialized instructions already detects their availability at runtime and uses those, even if the code was compiled with -march=x86-64.
Ok, but isn't text generation more general? E.g. you could ask it to predict the sentiment of a sentence and write the result as a sentence?
For most applications BERTs output is used to fine tune additional NN layers, eg for text classification
GPT and BERT were actually the first models published after Attention was published by Google.
its an encoder-decoder model whereas GPT is decoder only. feels like a pretty big difference, though in practice i honestly still dont have a strong grasp of how encoder-decoder is deficient to decoder-only when it comes to text generation. i get that BERT was designed for translation but why cant we scale it up and use it for textgen just the same
BERT can't be used in an autoregressive way because it doesn't output a new token, it simply generates embeddings from the existing tokens (you get one for each input token).
BERT - Encoder-only - embeddings for downstream tasks
GPT/OPT/etc - Decoder-only - language generation
T5/T0 - Encoder-decoder. Kind of does both?
(Bidirectional Encoder Representations from Transformers)
GPT predicts the next word by only look back at what we have seen so far. In other words, it's auto regressive.
> [W]e split training documents at the character level into a prefix, a middle part[,] and a suffix with the splitting locations sampled independently from a uniform distribution over the document length. We apply this transformation with a probability of 0.9 and to documents that are not cut across multiple model contexts only. We randomly format half of the splits in the prefix-suffix-middle (PSM) format and the other half in the compatible suffix-prefix-middle (SPM) format described in Bavarian et al. (2022, App. D). We extend Llama 2’s tokenizer with four special tokens that mark the beginning of the prefix, the middle part or the suffix, and the end of the infilling span
Tokens are super powerful :)
> As an example, our model would complete the string 'enu' with 'emrate' instead of 'merate' which shows awareness of the logical situation of the code but incomplete understanding of how tokens map to character-level spelling.
that doesn't really feel like a failure of language modeling to me
> Note, however, that the results in random span infilling are significantly worse in suffix-prefix-middle (SPM) format than in prefix-suffix-middle (PSM) format as it would require token healing (Microsoft, 2023),
Talk about a negative endorsement. I am continually disappointed in the auto correction implementation.
https://en.wikipedia.org/wiki/Bit_error_rate#Bit_error_rate_...
No. There are at least two kinds of costs. First, It takes time to search 'adjacent' domains. Second, by reducing your available acronyms/initialisms, you make it harder to map your architecture name onto those letters.
It is fun to think of some of the alternative BERT names that "could have been", such as BIDET = BIDirectional Encoder representations from Transformers.
Now my disappointment is immeassurable and my day is ruined.
"If you want to run the full notebook on a full size model, expect training the tokenizer to take ~15 hours, pretraining with the MLM objective to take ~17 hours (on a 3070 RTX, adjust expectations for your own system), and finetuning to take about an hour"
I wonder how hard it would be to modify this code to run on a 64GB M2 Mac.
It's frustrating how much potential that platform has for this kind of thing (given the way the GPU shares memory with the CPU) that isn't yet harnessed because most of the ecosystem is built around NVIDIA and CUDA.
I'm sure it's frustrating from a consumer perspective, but it should be no surprise why Nvidia won here. CUDA shipped unified memory addressing ten years before the M1 hit shelves. On top of that, their architecture and OS support is top-notch, you can ship your CUDA code on anything from a $250 Jetson to a $300,000 DGX system, and their hardware is relatively ubiquitous.
The frustrating thing is how companies like Apple and Nvidia insist on being each other's enemies. Only consumers feel the pain when researchers discover cool stuff like this and want to share.
I've personally got an 8GB M1 Macbook as my work development machine, and while I'm having a lot of fun with llama.cpp it does feel somewhat disconnected from the bulk of the ML ecosystem.
It isn't that hard, I was able to run in on M1. The changes are:
remove or modify multiprocessing - it doesn't work on Mac the same way as in the code;
replace `device = "cuda"` with `device = "mps"`
In this line ` att_idxs = (torch.clamp(torch.arange(context_size)[None, :] - torch.arange(context_size)[:, None], -pos_emb_radius, pos_emb_radius-1) % pos_emb_size).to("cuda")` replace cuda with "mps"
in `optim.AdamW` remove `fused=True` - we can't do it without CUDA
Replace ```with autocast(device_type='cuda', dtype=torch.float16): _, loss = mlm_head(bert(batch_data_torch_xs[mb_start_idx:mb_end_idx]), batch_data_torch_ys[mb_start_idx:mb_end_idx]) ```
with simply `_, loss = mlm_head(bert(batch_data_torch_xs[mb_start_idx:mb_end_idx]), batch_data_torch_ys[mb_start_idx:mb_end_idx])`
replace `scaler.scale(corrected_loss).backward()` with `corrected_loss.backward()`
replace ``` scaler.unscale_(optimizer) scaler.step(optimizer) scaler.update() ``` with `optimizer.step()`
It should work.