Meta AI: Fist high-performance self-supervised algorithm for multiple modalities
ai.facebook.com
ai.facebook.com
From what I understand, human validation (supervision) is not happening while algorithm is training on data. Is that right? Will this be open to the public via standard ML frameworks or proprietary?
The code is what we're interested in here, not "hold onto your papers" talk that tells us how to be excited about it.
I'd imagine on the flip side the potential problems would be when gains in one modality are offset by a decrease in performance in another or you are prevented from trying new things because of the choices made to support a unified architecture.
The multi-modal part is a big deal because we could have a model be able to understand all inputs human needs.
Facebook marketplace:
- Seller posts a listing with image + description
- FB uses data2vec to transform the image + description into one single vector (e.g. you could average the two vectors)
- Buyer searches for a product using text search - the text is also encoded into a vector with data2vec
- To serve product search results, you find the closest match to the text query vector from your pool of product vectors
This pattern described above is generally called vector search and is very common for recommendation algorithms and much more. There's a shift towards algorithms like data2vec that can combine different types of data into one vector. The aim is that an image of a dog and the word "dog" would map to the same vector, meaning that the vector represents the concept of a dog, regardless of the input data type.
The advantage of these "multi-modal" algorithms (i.e. they can take multiple data types as input) is that you can (in theory) use them across all of your ML algorithm needs. If you're Facebook, you have 100s of teams and services that have this need. A few examples:
- Instagram ads prioritization
- Instagram search
- Harmful content moderation
- Facebook content search
- Facebook marketplace search
Each of these is likely a separate team, very likely using a separate embedding algorithm. As approaches like data2vec improve, there will be some consolidation.
N.B. - I've made a lot of assumptions based on what I've seen at my current employer. If anyone from Meta/Facebook is reading this, please chime in!
Directly from the paper: "Our work does not perform multimodal training but aims to unify the learning objective for self-supervised learning in different modalities"
It is still great work though, a robust masking representation architecture that works across modalities.
> We apply data2vec separately to speech, images and text and it outperformed the previous best single-purpose algorithms for computer vision and speech and it is competitive on NLP tasks.
1. Do the cutout in feature space, not the original input space. (edit: cutout is actually in input space)
2. The above would likely just collapse the features to 0, so they use the same network that does the reconstruction to produce the features (!). In their own words:
"We first encode a masked version of the training sample (model in student mode) and then construct training targets by encoding the unmasked version of the input sample with the same model but when parameterized as an exponentially moving average of the model weights (model in teacher mode)"
> Do the cutout in feature space, not the original input space.
I think they do the cutout in the original input space, based on examples they show, e.g. they mask parts of text and grey out parts of an image.
> and make the network predict the missing part
I think predict the latent representation (as you say in (2)), but not of the missing part, but of the whole corrupted sample, and require it to be close to latent representation of the original input done by the teacher.
If I am not mistaken, the masking occurs in the input modality, not the feature-space, even though it is the feature-space that is used for the reconstruction task.
Regarding how the feature space is kept uncollapsed, it seems like a hyperparameter-tweaking (ie unsolved?) problem; quoting the paper:
" Representation collapse.
A common issue with algorithms which create and predict their own targets is representation collapse. This occurs when the model produces very similar representations for all masked segments making the problem trivial to solve. Different strategies have been proposed to address this issue, e.g., contrastive models such as wav2vec 2.0 (Baevski et al., 2020b) use the same target representation both as a positive and a negative example, preventing collapse. Algorithms such as BYOL (Grill et al., 2020) do not optimize the teacher parameters to minimize the loss. VicReg (Bardes et al., 2021) adds an explicit loss encouraging variance among different representations.
In our experiments we found that collapse is most likely to happen in the following scenarios:
First, the learning rate is too large or the learning rate warmup is too short which can often be solved by tuning the respective hyper-parameters.
Second, the EMA decay rate is too low which leads to student model collapse which is propagated to the teacher due to parameter tracking. This can be addressed by carefully tuning τ0, τe and τn.
Third, we found collapse to be more likely for modalities where adjacent targets are very correlated and where longer spans need to be masked, such as for speech. We address this by either explicitly penalizing the lack of variance (Bardes et al., 2021), or by promoting variance through normalizing target representations over the current sequence or batch (Grill et al., 2020). The former worked well for small models but is less reliable for larger models and it also requires tuning additional hyper-parameters. In contrast, we found applying instance or batch normalization before or after averaging targets to work well while being simpler. For models where targets are less correlated such as for vision and NLP, momentum tracking is sufficient to prevent representation collapse. "
That seems to explain why representations don't collapse into a constant, but not why they don't collapse to the same feature...
Even with large/labeled/cleaned dataset, it seems that each domain change or even formatting/encoding forces you change the the architecture.
This is not some random nitpicking. This is a great mystery worthy of its own detective TV show so I don't appreciate the downvotes. All my friends are extremely puzzled by this whole situation.
In practice, you spend time and expertise to reform the data into previously-known-to-work form.
FWIW, our datasets are huge, with dense data/noise ratio.
One area of research is extracting an effective type theory from data that is viewed as a semantic model, eg sensor data of a phenomenon would lead to a type theory describing it.
You’re essentially taking a TDA persistent homology/covering and reinterpreting that through the lens of homotopy type theory to “decompile” your data. There’s some early results, eg connecting convolutions and type division.
But that’s overall at really early stages of research.
Type division is like convolution from ML, which is why we can recognize “shapes of shapes” and undo the product structure. (And arguably, another avenue towards arriving at the manifold hypothesis.)
I’m currently working on the “easy” direction of encoding the type statements to matrices, with the hope most steps are reversible. (So far, so good.)
Still rough white paper:
https://www.zmgsabstract.com/whitepapers/shapes-as-digital-i...
GitHub repo for encoding type theory models, still VERY early:
Well, much simpler stuff.
You have a logistic dataset of objects/GPS/time. Fine. Now you add historic truck data which is location time series. If is not obvious how you can learn the delivery times between 2 items.
You need human expertise, and design a way to extract usable routes, and also solve multi-hop routes, and then maybe you can learn typical speed between 2 items.
It is doable. But it not with a generic "multi-modal architecture".
Detecting routes sounds like a persistent homology problem — and then you’re detecting flow across that inferred structure.
I think Michael Robinson has work in that area:
From the related works section of the paper.
Sheeeeeeeeeeeeeeeeeeit