I'm saying the client can take the action, but the user never sees the client because they interact with it through e.g. their AdBuster box, a glorified PiKVM that interacts with client, OCRs it, and produces simple filtered HTML for the user.
As long as you can attach a display and USB inputs, your "monitor" or "braille device" can go straight to a vision model, which can then send inputs from your "keyboard". There are already off the shelf devices that can do this if you install an agent harness.
Like an ad blocking DVR for the 2020s.