I take the point that it makes caching harder, but I don’t think that should overrule ergonomics concerns.
I take the point that it makes caching harder, but I don’t think that should overrule ergonomics concerns.
But for HTML and Markdown in particular, there's been so much useful work done in the semantic HTML space and microformats, that I don't know why anyone interested in this wouldn't just improve their HTML markup and leave it to the agent to do the rest. Convert to markdown or extract the useful HTML before handing it to model.
It's different content representations. A text in a markdown file is conceptually the same content as the same text in HTML (or PDF).
Link: /some-page rel="canonical"
Link: /some-page.json rel="alternate" type="application/json"
Link: /some-page.html rel="alternate" type="text/html"
Link: /some-page.md rel="alternate" type="text/markdown"On the flip side, I'd argue that the current centralisation of user agents (in the classical sense here) that benefit from programmatic content negotiation in form of a handful of harnesses like Claude or Codex is a great lever toward forcing the ecosystem to adopt better practices: If Anthropic added content negotiation as described in this thread to Claude, many sites would be incentivised to improve their web servers.
Maybe it will notice those, and maybe it will figure out the pattern for follow-up page requests, but there's no guarantee and it won't help the first request.
I am well aware that few sites are taking that much care of their API in terms of HTTP features, but all of the problems discussed here have solid and battle-tested answers.
The content at a URL should always match, the format in which its represented can be different based on the request. Its a bit like buying a book in hard copy or paperback, same book different format.