You're not alone. We've seen "legit" crawlers (FB, etc.) go pathological when they discover a large static corpus. First thing I'd check is whether you're accidentally making it easy to enumerate: sitemap(s), predictable URLs, internal links that expose the whole space, or query params that explode the cache key.
For Facebook specifically, don't rely on generic bot blocks. Verify via reverse DNS + IP allowlist/denylist, and consider serving a minimal "crawler view" (or even 429/410 for nonessential paths) rather than the full page. Also check if they're ignoring cache headers because you send no-store/private or vary on cookies.
Add hard cost controls at the edge: strict rate limits per IP / per /24 / per ASN, request token buckets by path prefix, and a "crawl budget" for unauth'd traffic. Even if they rotate residential IPs, limiting by ASN/geo + per-account quotas can cap blast radius.
If they're creating accounts, treat signup as an attack surface: tighten friction (email verification, domain allow/deny, delayed access to bulk pages for new accounts, per-account request caps, progressive challenges after N pages/min). CAPTCHA alone won't hold; you need behavioral throttles.
Make sure caching is actually working: confirm CDN cache hit ratio, ensure pages are cacheable (public, long max-age, s-maxage), strip cookies on public pages, normalize query strings, and consider serving a single pre-rendered HTML shell with deferred data for humans to reduce origin work.
If the content is "for humans," consider changing how it's delivered: chunk pages behind search or pagination instead of fully enumerable endpoints, require an interaction token per page view, and watermark or slightly personalize responses so bulk scraping is less valuable and easier to detect.