4,769 karma · joined July 26, 2010
[ my public key: https://keybase.io/kkwok; my proof: https://keybase.io/kkwok/sigs/5qGM4MGK91WsEYGjQRq6h0oHCn4i8PZaN8gq4FonTRM ]
In particular there was some prior art that I found for doing it from the OpenQwaQ project, which was a GPLv2 3D virtual world project in Squeak/Smalltalk started by Alan Kay[1] back in 2011.
If I recall correctly, it worked well for native apps, but didn't work well for Chromium/Electron apps because they would use an API for grabbing the global mouse position rather than reading coordinates from events.
[0]: https://github.com/antimatter15/microtask/blob/master/cocoa/... [1]: https://github.com/OpenFora/openqwaq/blob/189d6b0da1fb136118...
Assuming the platter is 100g, 42mm, spinning at 7200RPM, there is about 25J of rotational kinetic energy, whose mass equivalent is 2.8x10^-13g (0.28 femtograms).
Assuming 200 electrons per NAND floating gate with 3bits/cell TLC on a 2TB SSD, there would be 5.3x10^14 electrons, weighing about 0.5 femtograms.
The paper only applies to "special integers" where the prime factors are known to only differ by two bits.
This figure compares the prediction performance of GPT and quantitative models based on machine learning. Stepwise Logistic follows Ou and Penman (1989)’s structure with their 59 financial predictors. ANN is a three-layer artificial neural network model using the same set of variables as in Ou and Penman (1989). GPT (with CoT) provides the model with financial statement information and detailed chain-of-thought prompts. We report average accuracy (the percentage of correct predictions out of total predictions) for each method (left) and F1 score (right). We obtain bootstrapped standard errors by randomly sampling 1,000 observations 1,000 times and include 95% confidence intervals.If they had tried to replicate Llama 2 identically with their hardware setup, it'd cost a little bit less than twice their MoE model.
[1] https://github.com/meta-llama/llama/blob/main/MODEL_CARD.md#...
[2] https://blog.ori.co/choosing-between-nvidia-h100-vs-a100-per...
The original idea was to be able to navigate around with just arrow keys (conceptually by turning yourself around in place and being able to walk back and forward).
If true, it seems wild to sit on this kind of discovery for over two years.
Update: Seems like there might be even more history given the name LK-99 apparently comes from the names of its discoverers Dr. Lee and Dr. Kim, and the year of its discovery, 1999 (https://kr.linkedin.com/in/ji-hoon-kim-03508b80).
It's unfortunately not fully open source
User: What is 12 + 12?
StableLM: Yes, that is correct! 12 + 12 is equal to 18.
User: Write me a python program that calculates the nth fibonacci number
StableLM: Here is a python program that calculates the nth fibonacci number
def nth_fibonacci(n):
# base case
if n==0:
return 0
else:
return 1 + n - 1I quantized the weights to 4-bit and uploaded it to HuggingFace: https://huggingface.co/cakewalk/ggml-q4_0-stablelm-tuned-alp...
Here are instructions for running a little CLI interface on the 7B instruction tuned variant with llama.cpp-style quantized CPU inference.
pip install transformers wget
git clone https://github.com/antimatter15/cformers.git
cd cformers/cformers/cpp && make && cd ..
python chat.py -m stability
That said, I'm getting pretty poor performance out of the instruction tuned variant of this model. Even without quantization and just running their official Quickstart, it doesn't give a particularly coherent answer to "What is 2 + 2" This is a basic arithmetic operation that is 2 times the result of 2 plus the result of one plus the result of 2. In other words, 2 + 2 is equal to 2 + (2 x 2) + 1 + (2 x 1).I wish rather than stopping training early they would have run more data through a small model so we could have something more competitive with LLaMA 7B.
I suspect a big part of why stable diffusion managed to consume so much mindshare is that it can run on ordinary consumer hardware. On that point, I would be excited about an open-source RETRO (https://arxiv.org/pdf/2112.04426.pdf) model with comparable performance to GPT-3 that could run on consumer hardware with an NVMe SSD.