307 karma · joined November 16, 2011
Transformer is a patterning probabilistic machine for a sequence of identities[1]. These identities are fed to the transformer in lanes. The transformer is conditioned to shift lanes one position to the left until they make it to the output, and make a prediction in the right-most lane that got freed up. Attention adds an exponential amount of layer interconnectivity, when we compare it with a simple densely connected layers. The attention mask serves as a high-dimensional dropout, without which it would be extremely easy for the Transformer to simply repeat the inputs (and then fail to generalize when making the prediction). Each layer up until the vertical middle of the Transformer works with a higher contextual representation than the previous one, and this is again being unwound back to lower contexts from the middle layer back to the original identities (integers) on the outputs. This means that you have raw identities on the input and output which span a certain width/window of the input sequence, but in comparison the middle-most layer has a sequence of high level contexts spanning extreme lengths of the original input sequence, knowledge-wise. [1]It's important to know that modification (learning) by the Transformer, of the vector embeddings which represent the input/output identities/integers that the Transformer works with, constitute big portion of the Transformer's power, and the practical implication of that is that it's impractical to try to tell the Transformer that e.g. some of our identities are similar or there's some logical system in their similarity, because all the Transformer really cares about is the occurrence of these identities in the sequence we train the Transformer on, and the Transformer will figure out the similarities or any kind of logic in the sequence by itself.
Like, for example, releasing a weaponized virus from a lab. Oh wait...
I don't have IBS but more like a colitis and I was really good for months but it came back strong after one of my dogs passed away and I got terribly sad.
What I'm getting at is that more than a vitamin D, I suspect a strong link to happiness and life satisfaction.
2. If a crypto protocol doesn't evolve at the pace of available innovation, that particular blockchain will be superseded by a new one. That said, a (truly democratic) evolutionary process is a core part of every blockchain specification.
3. You can get blockchain data via public (and federated/proxied) API, but you can always cryptographically verify its veracity, and your edge device (e.g. your smartphone) can do that. The same the other way around, you cryptographically sign the inputs you send to the networks, so that no federated API can tamper them, because the secret key stays on your device. This is referred to as the "trust-less model".
For comparison, my understanding of Transformers, after going through Peter Bloem's "Transformers from scratch" [1], implementing/understanding the code and the actual flow of the mathematical quantities, my understanding is that:
- Transformers consist of 3 main parts: 1. Encoders/Decoders (I/O conversion), 2. Self-attention (Indexing), 3. Feed-forward trainable network (Memory).
- The Feed-forward is the most simple kind of (an input->single-layer) neural net, actually often implemented by a Conv1d layer, which is a simple matrix multiply plus a bias and activation.
- The most interesting part is the Multi-head self-attention, which I understand as [2] a randomly-initialized multi-dimensional indexing system where different heads focus on different variations of the indexed token instance (token = initially e.g. a word or a part of a word) with respect to its containing sequence/context. Such encoded token instance contains information about all other tokens of the input sequence = a.k.a. self-attention, and these tokens vary based on how the given "attention head" was (randomly) initialized.
The part that really hits you is when you understand that for a Transformer, a token is not unique only due to its content/identity (and due to all other tokens in the given context/sentence), but also due to its position in the context -- e.g. to the Transformer, the word "the" at the first position is a completely different word to the word "the" on e.g. the second position (even if the rest of the context would be the same). (Which is obviously a massive waste of space if you think about it, but at the same time, at the moment, the only/best way of doing it, because it moves a massive amount of processing from inference time to the training time - which is what our current von-Neumann hardware architectures require.)
So I did some googling[1] and apparently, until 2030 there's only 1 million EVs planned to hit the roads in the US, which is quite shocking given the dire trends of climate change and toxic pollution on this planet.
Another interesting fact is that the costs of extending the power grid will be passed on to all its consumers in the form of increased rates (as it should, because if you don't drive an EV, you should be "taxed" on the extra pollution your car creates).
[1] https://www.bcg.com/publications/2019/costs-revving-up-the-g...
Globally destroying web experience by mandating super-useless, super-obnoxious Cookie prompts on all web pages is one thing, but setting things up for "Equilibrium" or "1984" type of society is a whole new level, and I'm not sure if sheer stupidity can be claimed on this one.
Thinking that super-intelligence is containable is like thinking you can beat AlphaGo in the game of Go. And coming up with reasons why that's no so it's just you being in denial of your eventual mortality, as well as of the eventual extinction of the species. Best that we can hope for is that that extinction will be some form of transmutation to a higher level of evolution.
You could be better equipped for deep meditation and deep thinking than most of us, but maybe you just need to do some practice.
I, for one, am glad that my social bubble is utterly boring and un-engaging, and my level of engagement with Facebook is inversely proportional to my engagement/passion in life as such.
Sadly, AI will object this in court and win on all counts.
Just a hint - each of these "facts" or "rules" mean slightly something else depending on the context they are used in. This context is not hard-encodable because it slightly mutates with every new information learned or forgotten.
Interesting, last time I checked (years ago), Linux was abhorrent when it came to trackpad support.
And you are actually right that the trackpad surface on XPS in particular is actually one of the better/best ones. However options for tweaking it in Windows settings is non-existent, and so it is much slower than the mac trackpad, and my finger end up fatigued from casual scrolling, which never happens on mac. (This (trackpad scrolling speed) is tweakable eg in VSCode but not in Chrome etc.)
Macbook now has the M1 going for it.