I am working in Mistral robotics team. I confirm this is map-less. The only inputs are the text prompt and the front camera rgb image.
Or, I don't know, make your own vacuums.
One could maybe autogenerate these text planning commands, but it would require a map and the robot's current location, so it doesn't really solve that, unless it can find a specific thing completely on its own. How much of a planning horizon does it have?
The advantage over traditional approaches is presumably flexibility. LIDAR isn't going to solve an instruction like "find the man with the pink shirt".
A in-model memory approach is probably still deep research but maybe a Rag-like pipeline could work in some instances