AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling
junzhan2000.github.io
junzhan2000.github.io
Large Sequence Model (LSM) is OK too.
To elaborate a bit, that's effectively what happens in a natural-language-only LLM. Text tokens are immediately mapped to wide embeddings, which are processed by a transformer, producing another embedding, which is decoded back to a discrete token. To a first approximation, unified multimodal models substitute f: text -> [embedding] with g: image -> [embedding,] and act on a sequence like [*f(t), *g(i), ...]. So... the multimodal model maps text/image parts to wide embeddings, which are processed by a transformer producing another embedding, which is decoded back to a discrete text/image part. Which sounds somewhat similar to the former :).
Of course the function g is a bit more complicated, and decoding takes more work, but the underlying machinery is unchanged. The ability to handle multi-modal inputs/outputs uniformly is exactly what's so exciting about work like this!
Whether data is continuous or discrete, no matter its modality (text, video, music, etc.), we now have an array of proven methods for representing it with discrete tokens, enabling us to use existing sequence modeling architectures (Transformers, linear RNNs).
We live in interesting times!