278 karma · joined June 18, 2019
Is there a good overview to learn about the current models, which all just seem like cryptic acronyms to me? in apps like Windy etc. WRF, TRRM, IK-HRRR-3km, ECMWF-9km,...
I understand by now that small grid cells are better for local prediction and that thermic winds are mostly missing from them all.
If every use of AI was like this, I perhaps wouldn't have this slight allergic reaction to it, but as it stands, this voice and rhythm has become associated with bad lazy grifting writing.
load-bearing this load-bearing that
@dang I would welcome a small secondary button that one can vote on to community-driven mark a comment as AI, just so we know.
Additionally, I am really missing support for stacked diffs, ie, easily pushing a number of commits into one PR on github each such that they all show their incremental diff.
ezyang's gh stack was pretty useful, if a little bit fragile [0] and graphite.dev is also very nice, but paid software with a strong VC based motivation to become everyone's everything instead of a nice focused tool.
[0] https://github.com/ezyang/ghstack
I'm also not super happy with the default 3-way merge editor, but often cannot use vscode or other GUIs.
the one thing I really want beyond persistent over whatsapp is threads. I hope matrix has a similarly trivial app/web/pc story as mattermost has, because the other users are not necessarily able to handle anything more complex than "download an app and sign in".
Really enjoyed the part talking about Tolkien. It reminds me of my own LOTR experience: I finished the trilogy in three consecutive summers in the hilly countryside of Italy near Rome. The first summer I made it through the fellowship of the ring with a lot of patience and trying for the slower moving parts. The summer after that I started over and read book 1 and 2, and in year three I felt I was finally in sync with the pace of the book and enjoyed reading through book 1, 2, and 3 in a few weeks.
I feel that this idea is now in jeopardy, if I understand the 10k message history is the limit correctly.
And there I thought I had a solution to slowly bring over project channels, family related things etc. that was as reliable as "my linux box will be reachable on the public internet" and I am willing to manage that it does.
Seems I was wrong, but I don't know which other software has better future proofing.
I feel that this idea is now in jeopardy, if I understand the 10k message history is the limit correctly.
And there I thought I had a solution to slowly bring over project channels, family related things etc. that was as reliable as "my linux box will be reachable on the public internet" and I am willing to manage that it does.
Seems I was wrong, but I don't know which other software has better future proofing.
- apply learned knowledge from its parameters to every part of the input representation („tokenized“, ie, chunkified text).
- apply mixing of the input representation with other parts of itself. This is called „attention“ for historical reasons. The original attention computes mixing of (roughly) every token (say N) with every other (N). Thus we pay a compute cost relative to N squared.
The attention cost therefore grows quickly in terms of compute and memory requirements when the input / conversation becomes long (or may even contain documents).
It is a very active field of research to reduce the quadratic part to something cheaper, but so far this has been rather difficult, because as you readily see this means that you have to give up mixing every part of the input with every other.
Most of the time mixing token representations close to each other is more important than those that are far apart, but not always. That’s why there are many attempts now to do away with most of the quadratic attention layers but keeping some.
What to do during mixing when you give up all-to-all attention is the big research question because many approaches seem to behave well only under some conditions and we haven’t established something as good and versatile as all-to-all attention.
If you forgo all-to-all you also open up so many options (eg. all-to-something followed by something-to-all as a pattern, where something serves as a sort of memory or state that summarizes all inputs at once. You can imagine that summarizing all inputs well is a lossy abstraction though, etc.)
Hot spring baths usually top out around 42-43C