My use case is that I have around +20M documents with non-unique hashes and each query should return an arbitrary amount of documents matching the query / filter as well as calculated meta-data based on the results in the aggregation field.
Now the issues is that if you want to have only one document per hash, you need to use a TermAggregation on the hash field followed by a TopHitsAggregation of size 1 to obtain the actual document rather than the hash field.
At this point you have many many buckets containing a single document but:
- you can't paginate them since TermAggregation doesn't let you do pagination for the reasons you explained and linked to above
- you can't calculate an aggregation on all the returned documents since all of them are in their own separate bucket (by hash)