Large Sequence Model (LSM) is OK too.
To elaborate a bit, that's effectively what happens in a natural-language-only LLM. Text tokens are immediately mapped to wide embeddings, which are processed by a transformer, producing another embedding, which is decoded back to a discrete token. To a first approximation, unified multimodal models substitute f: text -> [embedding] with g: image -> [embedding,] and act on a sequence like [*f(t), *g(i), ...]. So... the multimodal model maps text/image parts to wide embeddings, which are processed by a transformer producing another embedding, which is decoded back to a discrete text/image part. Which sounds somewhat similar to the former :).
Of course the function g is a bit more complicated, and decoding takes more work, but the underlying machinery is unchanged. The ability to handle multi-modal inputs/outputs uniformly is exactly what's so exciting about work like this!