I tried accessing Gemini 2.0 Flash through Google AI Studio in the Safari browser on my iPhone, and to my surprise it worked. After I gave it access to my microphone and camera, I was able to have a pretty smooth conversation with it about what it saw through the camera. I pointed the camera at things in my room and asked what they were, and it identified them accurately. It was also able to read text in both English and Japanese. It correctly named a note I played on a piano when I showed it the keyboard with my finger playing the note, but it couldn’t identify notes by sound alone.
The latency was low, though the conversation got cut off a few times.