The "book as an MCP server" framing is the part that interests me. I run a small MCP server myself (semantic search over government documents), and the thing I keep noticing is that the interface changes what people ask — nobody asks a PDF "what should I do in my situation," but they'll ask a tool that. Shipping specialized tools (draft_decision_doc etc.) over one generic ask_the_book tool also seems like the right call — clearer intent per call, less context spent on tool descriptions. I noticed you went with keyword search over a summary corpus plus citations rather than embedding the full text — curious whether that was a licensing decision (keeping the manuscript out of the package) or you found summaries-with-citations just work better for this kind of prescriptive content than semantic retrieval over the full prose?
Cloudflare Workers — the whole thing is a single stateless worker in front of a vector index, so I get TLS, DDoS filtering and bot scoring at the edge without running any infra myself. The tradeoff is you're limited to what the platform exposes; something like a gateway with pluggable security policies would matter more once there are write-capable tools involved. Mine is read-only search, which keeps the attack surface pretty boring — probably why my logs are mostly scanners looking for a WordPress that doesn't exist. Thanks for the Kong pointer, will take a look.
Small data point from the operator side: I run a tiny public MCP server, and when I finally turned on request logging, almost none of the traffic was what I expected. Mostly link-preview bots, keepalive pings, and scanners probing for wp-admin on an endpoint that isn't even WordPress. Made me realize most small MCP deployments probably have zero visibility into this — people ship a server and never look at what's actually hitting it. So the "you can't flag risky AI-generated commands without context" point resonates; at the low end the problem is even more basic, there's no context at all.