I understand how annoying, painful, and disruptive to the user experience it is to have part of your app taken over by browser or OS default widgets and behavior is. It's sheer hell on UX. Yet the custom styling that would make it integrate smoothly would be a godsend for attackers and phishers.
The role that UX - and misuses of UX - play in security is often not considered deeply by highly sophisticated users.
Compared to that, one icon, that is the same as that of the company, is not that threatening.
Not only that, but if you consider a sign in form that wouldn't have a logo, it would be way easier to trick user into putting their credentials in, because the user wouldn't be able to differentiate them. Also OAuth is always branded AFAIK.
Users may also notice discrepancies in the logo, if it was cloned poorly. Though I can't think of a way someone couldn't forge a logo given all the possibilities. Adobe Illustrator can trace images into svg and there's plenty of companies' svg logos just in the google search.
I think the real answer is that allowing this kind of deep and arbitrary styling of interfaces is far more dangerous than it is helpful.
Unless you're somehow able to get everyone to transition at once, never support a non-NFT fallback and, ensure that no sufficiently visually similar logo ever gets registered... which IMO sounds both challenging and like a hacky re-implementation of a trademark system.
The stated (and valid) concern is that malicious actors would use fraudulent branding on those browser auth popups — let’s tell the user we are MSFT or AAPL and steal their password.
One solution is for the branding to consist entirely of NFT assets that can all be tracked to a definitive owner, and use some DNS-based glue (ala DKIM/SPIF for email) to link the NFT to the TLD.
Then your browser can refuse to show the MSFT logo (and show a big red fraud alert page) if the owner of the branding can’t be reliably traced back to Microsoft (owner of the site).
Of course, neither the NFT nor the sig-in-DNS approach actually solves the problem of a visually identical but technically different image (use a slightly different color in a few places, etc.) being used to trick people. I'm not sure what we've gained. The malicious use case would seem like it's not effectively prevented.
I can’t help but think there must be a way to make it work.
I suspect there's a lot of complexity hidden in the perceptual model requirement, though.
As always the only defense is to compare what is being asked for with the FQDN.