Clicking through the website with computer use isn't currently part of the website generation chain for most models, and while there are gaze prediction models, to my knowledge none are productized or universal to the point where you can chuck them in and consistently predict what draws attention on any given page
(Generally I think it is alright to interpret "can't do" in terms of "the common workflow that current models will afford" and not "definitionally incapable of")