In order to run large language models, we should all be buying a fully loaded Mac Studio (128GB of ram, 20 CPU cores, a lot of GPU and Neural cores.) and putting Linux on it to remove the artificial restrictions.
Yes, we will be running them soon in low end hardware, but we need to get at least to GPT-3.5-turbo level of inference speed and quality before we try to make it small.
I already started.
$ neofetch
-` x@decpti
.o+` -------
`ooo/ OS: Arch Linux ARM aarch64
`+oooo: Host: Apple Mac Studio (M1 Ultra, 2022)
`+oooooo: Kernel: 6.1.0-rc6-asahi-4-1-ARCH
-+oooooo+: Uptime: 4 hours, 23 mins
`/:-:++oooo+: Packages: 177 (pacman)
`/++++/+++++++: Shell: bash 5.1.16
`/++++++++++++++: Resolution: 1920x1080
`/+++ooooooooooooo/` Terminal: /dev/pts/0
./ooosssso++osssssso+` CPU: (20) @ 2.064GHz
.oossssso-````/ossssss+` Memory: 717MiB / 129540MiB
-osssssso. :ssssssso.
:osssssss/ osssso+++.
/ossssssss/ +ssssooo/-