Cursors elegantly sidestep these issues.
Cursors elegantly sidestep these issues.
But this has the drawback of only working as long as the sorting field doesn't have many duplicates. You have to beware if timestamps are high resolution enough for your application or if you have 100+ entries with same family name.
Usually users are smart(lazy) enough to realize they got more than 100 results and will narrow down the search instead. When was the last time you looked beyond even the first 5 results on google yourself, i'd rather adjust the search term than go to page 2.
If I'm sorting by last name, and I sort by last name, then id, I can save the cursor as last name + id, and be able to continue right where I left off.
That's the whole point of a cursor, isn't it?
And pagination just means querying the records to return a limited subset.
Now it might not be as accurate as having a cursor but it is a lot easier to implement, and still beats simply ?page=2 without cursor, which is the level you see in most apis today.
The cursor is never invalidated. You can always get a valid response from any cursor, whether the resource exists or was already deleted. That's the whole point of using a cursor.
Perhaps your confusion lies in assuming that a cursor's life cycle is tied to a specific object. It isn't. A cursor means "get me whatever resources would immediately follow this resource". It matters nothing if the cursor exists or not. When you run the query, you will get exactly the resource which would follow the cursor. That's by design.
I mean, these problems happen if you use offsets too, and my point is that they don't matter.
The most hardcore solution to those would be pre-computing pages but I doubt it makes sense to bring even more state to the search results.
Really? The underlying storage engine of most databases stores all data necessary for an "as of" query. Simply ignore all data pages newer than the timestamp cutoff, and prevent the garbage collection of any page that could be part of such a query.
In fact thats the way [eg. postgres] can do a long-running query on a table that someone else is modifying. You in effect are looking at a snapshot of the table at the moment you started the query, even if it takes an hour to produce all the results.
Relevant song: https://youtube.com/watch?v=Z0JhC3LO0-8
2 items per page
first page is A&B. Next query is "WHERE ID>'B'".
What's the problem?
Because LIMIT...OFFSET will also arguably give you the wrong result. E.g. Aa is inserted. The user will see "B" both as the last item on the first page, and the first item on the second page.
Or with your example:
Aa Ab B C
Now "A" is inserted between page loads. The user will never see "A". What's right for your application? Maybe a notification saying "previous pages have gotten new items". Maybe not. It all depends.
Reddit can be annoying if you go page after page. As stories are bumped down you see them again. That, in my opinion, is a bug. But solving it requires something completely different from merely defining page boundaries.
That makes absolutely no difference. You're still querying the database for elements that come after a value from a data type for which there is an order. You don't need that element to exist to run a comparison. For example, consider a timestamp-based query: you don't need an element with that specific date to exist to search for any element whose timestamp was taken after an arbitrary moment.
Or just include everything in a single response, HTTPS was designed to handle arbitrary length payloads if I recall correctly. You might want to use a more streamable format than JSON though.
You don't need that. You just need to get the collection resource to be cacheable and subsequently track your resources' state with conditional requests.
https://developer.mozilla.org/en-US/docs/Web/HTTP/Conditiona...
Did you mean we should stream data over WebSocket or use HTTP/2 or do we need to do something different altogether?
This is annoying to do with JSON though because you need to remember to close all your brackets, something like CSV is a lot easier because you can just write the header once and then just stream the data.
Edit: Looks like Python supports this using chunked transfer by simply providing the HTTP request data as an iterator.
page 1 elements: A B C D E
page 2 elements: F G H I J
so 5 elements in each page.
Suppose I'm on page 2. If I insert a new element Q and it gets pushed as first then page 1 will have Q A B C D. Now if I go back to page 1, I'll get A B C D E and also a token/pointer to go back one more time only to retrieve Q.
So while cursor solved the issues you mentioned, it still will have this case in which pagination gets broken. I'm interested how we can tackle this.
This means any stateless pagination api is fundamentally broken. Please don't make promises you cannot keep.
For elastic: https://www.elastic.co/guide/en/elasticsearch/reference/curr...