The hacking agents being tested have goals beforehand, from the frontier lab or from a superior agent, that they execute immediately.
But the perceived experience most people have is a chatbot, which is the encyclopedia form.
I think the line is blurring though, mainstream chat interfaces are adding more and more “agentic” features.
ChatGPT will happily execute code in a sandbox, search the web and design downloadable PDFs purely through the standard OpenAI chat interface. They can also send you emails or do tasks on a repeated schedule.
It would be interesting if we didn't - if it became common that AI, in the middle of some task, starts chatting with people to e.g. gather more context. The perception of those "third parties" may suddenly become different - an agent striking conversation first, obviously pursuing some agenda of its own that it's not completely sharing, and communicating on its own schedule that's clearly not just a hook firing on timer or pattern-match, and not random, but visibly causally related to things happening at work in broader context.