This is undocumented (frustrating) but it looks like it's chunking them, running embeddings on the chunks and storing the results in a https://qdrant.tech/ vector database.
We know it's Qdrant because an error message leaked that detail: https://twitter.com/altryne/status/1721989500291989585
It only applies that mechanism to some file types though - PDFs and .md files for example.
Other file formats that you upload are stored and made available to Code Interpreter but are not embedded for vector search.