You can combine multiple queries into one request in such a way that transport layer caching just isn’t effective.
If you’re making a content and read heavy website, you can run GraphQL over GET, and cache traditionally.
The spec says nothing about the transport layer itself.
Why return 200 for a request, when part of it failed?
Separation of concerns. Network layer errors are one thing, application layer errors are another.
They require different resolutions.
With a network layer issue, you could retry, or perhaps you need to fetch a new access token then retry.
For application layer issues, say you request data on 4 different entities, but the service for one of those types is down. Should you chuck out the whole request, or return everything and something like a Problem or Error type for the failed one?
Perhaps you tried to access a field you don’t have permissions for, or require elevated permissions. Should you fail the whole request, or return an error on the field itself, allowing you to inform the user what you need them to do to continue?
The point is precisely that the client processes the error closest to where it’s relevant. This works especially well with component based rendering.