563 karma · joined December 3, 2021
That's a great list of existing embeddings models (in addition the SentenceBERT models https://www.sbert.net/docs/pretrained_models.html).
Transformer Feed-Forward Layers Are Key-Value Memories https://arxiv.org/abs/2012.14913
The Dual Form of Neural Networks Revisited: Connecting Test Time Predictions to Training Patterns via Spotlights of Attention https://arxiv.org/abs/2202.05798
One of the most interesting presentations in the last session of the workshop is this talk by David Bau titled "Direct Model Editing and Mechanistic Interpretability". David and his team locate exact information in the model, and edit it. So for example they edit the location of the Eiffel Tower to be in Rome. So whenever the model generates anything involving location (e.g., the view from the top of the tower), it actually describes Rome
Talk: https://www.youtube.com/watch?v=I1ELSZNFeHc
Paper: https://rome.baulab.info/
Follow-up work: https://memit.baulab.info/
There is also work on "Probing" the representation vectors inside the model and investigating what information is encoded at the various layers. One early Transformer Explainability paper (BERT Rediscovers the Classical NLP Pipeline https://arxiv.org/abs/1905.05950) found that "the model represents the steps of the traditional NLP pipeline in an interpretable and localizable way: POS tagging, parsing, NER, semantic roles, then coreference". Meaning that the representations in the earlier layers encode things like whether a token is a verb or noun, and later layers encode other, higher-level information. I've made an intro to these probing methods here: https://www.youtube.com/watch?v=HJn-OTNLnoE
A lot of applied work doesn't require interpretability and explainability at the moment, but I suspect the interest will continue to increase.
I appreciate you elaborating on your feedback. Thank you.
https://peterbloem.nl/blog/transformers
https://e2eml.school/transformers.html
I would also add Luis Serrano's article here: https://txt.cohere.com/what-are-transformer-models/ (HN discussion: https://news.ycombinator.com/item?id=35576918).
Looking back at The Illustrated Transformer, when I introduce people to the topic now, I find I can hide some complexity by omitting the encoder-decoder architecture and focusing only on one. Decoders are great because now a lot of people come to Transformers having heard of GPT models (which are decoder only). So for me, my canonical intro to Transformers now only touches on a decoder model. You can see this narrative here: https://www.youtube.com/watch?v=MQnJZuBGmSQ
If you want to query for a search term, you can use a trial API key which is free to use for prototyping. The embedding model itself is not open source, though. [co-author of the post here]
Your prompt suggestion is a good one for LLMs as a whole. Any information added to the context informs the model and nudges it towards the expected answer format.
Otherwise, you may be prompting a Base LLM expecting the behavior of a different kind of LLM (an instruction-tuned chat model).
Cohere's Command model builds on top of the base model, giving it the capability to follow instructions and user commands.
This article has the first four points (out of thirteen) that are useful to keep in mind. They are:
1- Recent AI developments are awe-inspiring and promise to change the world. But when?
2- Make a distinction between impressive cherry-picked demos, and reliable use cases that are ready for the marketplace
3- Think of models as components of intelligent systems, not minds
4- Generative AI alone is only the tip of the iceberg
---
The article goes into each one with more detail. Welcoming all feedback and perspectives as we collectively figure out this new frontier.
In a context like this, we use tensor because it allows for any number of dimensions (while vector/ array is only one, matrix is two). When you get into ML libraries, both popular packages PyTorch and TensorFlow use the "tensor" terminology.
It's a good point. Hope it's clearer for devs with "array" terminology.
1- Forward Diffusion (adding noise, and training the Unet to predict how much noise is added in each step)
2- Generating the image by denoising. This doesn't predict the final image, each step only predicts a small slice of noise (the removal of which leads to images similar to what the model encountered in step 1).
So it is indeed an iterative processes in that way, each step taking one step towards the final image.
Thank you!
1- Clustering by UMAP. Here the plot would show clean separation of topics. But the clustering algorithm would be working on highly compressed data (from the 1024 dimensions of the embedding down to the 2 of UMAP).
2- BERTopic's approach of doing UMAP down to 5 dimensions, using this dimensionality for clustering, then UMAP again from 5 to 2. Which is an interesting approach.
I've heard people having good results with all three. It's kinda hard to objectively compare, but my leaning was to give the clustering algorithm the representation containing the most information about the text.
[2] https://assets.cohere.ai/blog/text-clustering/askhn_cluster_...
[3] https://assets.cohere.ai/blog/text-clustering/askhn_cluster_...
[4] https://assets.cohere.ai/blog/text-clustering/askhn_cluster_...
[5] https://assets.cohere.ai/blog/text-clustering/askhn_cluster_...
[6] https://assets.cohere.ai/blog/text-clustering/hn10k_clustere...
[7] https://assets.cohere.ai/blog/text-clustering/askhn-3k.html
[8] https://storage.googleapis.com/cohere-assets/blog/text-clust...
[9] https://storage.googleapis.com/cohere-assets/blog/text-clust...
[10] https://colab.research.google.com/github/cohere-ai/notebooks...
- https://www.assemblyai.com/blog/how-dall-e-2-actually-works/
For this one in particular, here are a few more results for Battlestar and The Office:
https://twitter.com/Miles_Brundage/status/153247388947686195...