Slicing by ID is how Github does commit history and it drives me nuts that I can't jump several pages, for example, to see when the first commit was, or to guess whereabouts some commit is given a known time range. IDs make it impossible to do anything but iterate step by step.
I much prefer an interface that exposes:
a) how many items exist in total
b) which offset it's starting from
c) how many items in the slice
Then if I want a slice from 300-8000th items, I can type exactly that in the URL. Yes, I understand this will render a huge page, just let me do it the one time so I won't be spamming your server with requests over the next hour trying to find something while fighting against bad UX.
Fair point. I agree, that's a nice side effect. But what if you're looking at page 1, and an item is removed from that page? Then you'll never see the first item at page 2, because it's now the last on page 1.
> Then if I want a slice from 300-8000th items, I can type exactly that in the URL.
That's a nice feature. But it can put a lot of load on your backend if you paginate over 10 of thousands of items.
How common are concurrent removals in practice though? I can't think of a single instance off the top of my head where this is problematic and you can't just go back to the previous page if you're really paranoid or confused that something is "missing"
> it can put a lot of load on your backend if you paginate over 10 of thousands of items
Look at it in aggregate. Fiddling w/ URL params to paginate (instead of clicking pagination links) is a power user move. If I want the date of first commit in a repo, I'll only look at the last page (vs paging through the entire history). For guessing, I can click a page, and if I went too far, binary search from there (again, vs linear search). Etc.
Even for the most degenerate use case (e.g. some jerk trying to crawl over the entire dataset), the load is smaller with a single request than the overhead of multiple requests. Paginating is not an appropriate mitigation strategy against this type of traffic, and you arguably can implement detection/caching/blocking mechanisms much more easily for naive huge queries than if you need to differentiate regular traffic from bots.
The whole issue is that you're never going to know about it. Sure, you can write some convoluted automated process to double-check previous page(s) but most people aren't going to do that.
Another thing to consider when paginating via id cursors: if the deleted item is the one your cursor is sitting on, then you no longer have a frame of reference at all. For example, what happens to commit history pagination in github's implementation if I rewrite git history? Chibicc for example deliberately rewrites its history for didactic purposes.
It's difficult to do in most HTTP deployments where we try to limit HTTP response duration.
> if the deleted item is the one your cursor is sitting on, then you no longer have a frame of reference at all
Not a problem. If your order your results by timestamp and ID, then you provide the timestamp and ID of the last item returned, and take anything after. Doesn't matter if the item is still there or not.
To prevent that, usually maximum page size is enforced on the server anyway and the client is informed about the actual page size in the metadata in the reply.
It's a trade-off.
1. APIs consumed by other backends lets call them API2B
2. APIs consumed by frontends lets call them API2C
In this case cursors are better for API2B but not for API2C as in case of API2C most users expect to be able to jump directly to a specific page.
At least when I am designing an API i take different decisions based on this split. For example in case of API2C I always want to see FE design even if I work on backend.
Thus using page numbers is probably a pretty poor proxy for what you're actually trying to do when you say "getting to arbitrary pages." Presumably you are wanting to skip to a specific place in the list, perhaps specified as a percentage ("take me halfway through the list") or as some predicate on the data ("take me to items from 2 weeks ago"). APIs should provide ways of expressing these specific places in a list, instead of requiring you to either guess page numbers or do extra work to calculate them.
You'd still present the user with a set of results (a "page") and would let them seek forwards/backwards to the adjacent subsets of results.
AFAICT cursors can't support this.
And even if you go halfways into the list you can't display all subsequent items, so you'd still have to paginate in some way.
The page metaphor is there for a good reason.
The other concept is using an actual page number to make requests, e.g. requesting {page: 1} and then subsequently requesting {page: 2}. This concept is the one I was claiming is less desirable than some alternatives.
As for cursors, I don't see any reason why you couldn't make requests like {listPosition: "50%"} or {createdBefore: "2020-02-15"} and then still use cursors in the response to request the previous or next page. (Those two examples probably aren't actually good API naming conventions, but it should demonstrate the idea.)
Which ones?
for example GET /customer/1/orders/
response:
{'orders': [order1, order2, order3],
'navigation': {
'firstpage': '/get/customer/1/orders/',
'nextpage': '/get/customer/1/orders/?query=orderId^GT4',
'lastpage': '/get/customer/1/orders/?query=orderId^GT990',
'totalrows': 1000
}
GT means greater than
Edit: obviously the proper impl totally depends on your app. If the result set is enormous, caching doesn't make sense. If it's rapidly changing, pagination probably doesn't make sense. etc, etc.
If we don't care about added items (they can be deduped client-side), and only care about removed items, maybe the backend can maintain a tombstone timestamp on each deleted item, instead of deleting them, and then the client can provide a "snapshot" timestamp in the query that can be compared with the tombstone timestamp.
A query ID is super easy to implement, and yes, for any application you run you have to be aware of the resource requirements. Tune your eviction policy, and this is totally feasible.
Deduping and tombstones is messier to implement IMO, and hitting a cached result for the next page (e.g. a redis LIST) is probably less expensive than a "query" (whatever that means for your backend).
That said, using “page of results preceding <first id if following page>” and “page of results following <last id of preceding page>” reduces obvious pagination artifacts compared to <page number n>. Whether it's better to beat consumers over the head with inconsistency due to concurrent changes probably depends on the application or provide a nearer illusion of consistency probably depends on application domain and use case.
Padding the offset would solve for the problem mentioned where deleting an item would mean some non deleted items are not included in the paging results because they got moved up a page after that page was requested, but before the next page was requested. For example If I request 100 items at a time but set my offset to be 90 more than what i have received so far, i can expect my response to have duplicates, if it does not then i know my offset was not padded enough and i can request from a different offset. Of course you would adjust the numbers based on knowledge of the data.
Edit: If you used the ID of the last item instead of an offset, then you could get errors if your last item is in fact the one that was deleted.
> Edit: If you used the ID of the last item instead of an offset, then you could get errors if your last item is in fact the one that was deleted.
If the items are sorted by ID, then we use the ID of the last item x. But for example if they are sorted by timestamp, then we use the timestamp of the last item (and maybe the ID to break ties). Then it doesn't matter if the last item is still there or not. We are only considering values before or after depending on the sort order.
for instance your comment is 26227524
and the parent post is 26225373
edit: hacker news pagination is not high priority apparently since they just use the easiest way with p= some number.