There's a Reddit forum[1] for discussing personal bankruptcy, btw, if you're curious what the process looks like. (I found it interesting myself.)
808 karma · joined October 23, 2014
There's a Reddit forum[1] for discussing personal bankruptcy, btw, if you're curious what the process looks like. (I found it interesting myself.)
[1] https://www.harvey.ai/blog See second story announcing their Series A.
I just heard the Liszt Funerailles played on the Alexander piano (YouTube on the embedded main page), and yes it sounds very unusually melodic in the bass.
Not much on empirical observations, though.
>This vector seems to get taller every model year, for example the recent LLaMA 2 model from Meta uses an embedding vector of length 3,204, which works out to 6KB+ in half-precision floating-point, just to represent one word in the vocabulary, which typically contains 30,000 - 50,000 entries.
>Now if you’re a memory-miserly C programmer like me, you might wonder, why in the world are these AI goobers using 6KB to represent something that ought to take, like 2 bytes tops? If their vocabulary is less than 2^16=65,384, we only need 16 bits to represent an entry, yeah?
>Well, here is what the Transformer is actually doing: it transforms (eh?) that input vector to an output vector of the same size, and that final 6KB output vector needs to encode absolutely everything needed to predict the token after the current one. The job of each layer of the Transformer is quite literally adding information to the original, single-word vector. This is where the residual (née skip) connections come in: all of the attention machinery is just adding supplementary material to that original two bytes’ worth of information, analyzing the larger context to indicate, for instance, that the word pupil is referring to a student, and not to the hole in your eye.
Firstly, he is confusing representation with encoding--he's right that 2 bytes is enough to encode any token. That is in fact approximately how it's done: a code book is indexed into (with a longint in pytorch, at least last I worked with it ~6 months ago). The purpose of the embedding is to allow the model to learn a representation of the token, a la word2vec. (Though this representation is purely based on the characters comprising the token and does not distinguish between "student" and "eye" in the case of "pupil" as in his example.)
Secondly, his description of each layer's function as adding information to the original vector misses the mark IMO--it is more like the original input is convolved with the weights of the transformer into the output. I am probably missing the mark a bit here as well.
Lastly, his statement that the embedding vector of the final token output needs all the info for the next token is plainly incorrect. The final decoder layer, when predicting the next token, uses all the information from the previous layer's hidden layer, which is the size of the hidden units times the number of tokens so far.
Just googled it for a refresher and found this[1] comprehensive but readable explanation of the above idea. The obituary refers to it briefly as well.
A recent note I took from Ch 7: Unrestraint and Pleasure
"...virtue keeps the source safe, while vice destroys it, and in actions the source is that for the sake of which one acts ... so neither there nor here is reason able to teach anyone the sources, but here it is virtue, either natural or habituated, that directs one to right opinion about the source."
Reminded me of Buddhist teachings that say that moral conduct is necessary and you can't just think/meditate your way to wisdom/enlightenment[2].
[1] Nicomachean Ethics, Aristotle. tr. Joe Sachs [2] cf. https://puredhamma.net/living-dhamma/transition-to-noble-eig...
If the inputs are images, you may find that some dimension scores e.g. how much blue there is in the image. Though often it's not that simple (there could be multiple dimensions that relate to how blue the image is, especially if the embedding dimensionality is large, which it does tend to be these days. Though you could reduce the embedding dimensionality first using PCA, and see what input images correspond to high/low values of the first principal component, etc.).
Not necessarily true. If you have a good pair of speakers, a good amplifier, but a bad DAC[1], a CD can sound worse than vinyl (whose output does not need to go through a DAC). Old CD players (like mine) have dated built-in DACs, so this is not too exceptional a situation.
The above comment holds even if the vinyl was made from a CD source, since the vinyl maker could have used a quality DAC that's better than your CD player's.
For a long time, I didn't understand why my FM radio channel (WQXR) sounded better than my CDs. Turns out my CD player's DAC was poor in comparison to what the radio station was using to play their CDs.
[1]Digital to Analog Converter.
A more recent academic but high-level explanation of transformers, very good for detail on the different flow flavors (e.g. encoder-decoder vs decoder only), is Formal Algorithms for Transformers[2], from DeepMind.
[1] https://jalammar.github.io/illustrated-transformer/ [2] https://arxiv.org/abs/2207.09238
We are slowly working through the issues reported in this thread. Thanks for the kind and constructive feedback!
[1] https://chrome.google.com/webstore/detail/qa-with-klavier/jb...
And given your explanation of your usecase, this feature looks more compelling to build out. Would you consider messaging me (email in my profile)? We'd love to chat, and maybe roll out a solution for you as a pilot customer.
2: We are thinking of the website integration. Do you think OpenAI may release this too? Questions received by email is a new idea that sounds interesting!
3: Thanks for the suggestion – we will look into it.
We'll look into adding memory (either as a default or as an option).
(And yes, I feel the anxiety too--to keep up with what people are doing with the tech!)
Our site is meant for Q&A, and has a layer of tech that finds the sections in the large document that are relevant to the Q first. This will not work well in general for summarization on unstructured content. But most content tends to be structured and in practice we are finding that the approach still works on e.g. news articles, wikipedia articles, blog posts. It's almost as if where it doesn't work, a human would have trouble too. (e.g., on a long rambling HN thread).
Given your comment we are going to consider retaining question history (or offer an option to do so)!