The author filtered with:
CreateAt > $1 OR (CreateAt = $1 AND Id > $2)
That can't use one index for both conditions. Instead, make it faster and simpler: (CreateAt, Id) > ($1, $2)
End of post.The author filtered with:
CreateAt > $1 OR (CreateAt = $1 AND Id > $2)
That can't use one index for both conditions. Instead, make it faster and simpler: (CreateAt, Id) > ($1, $2)
End of post.1. Always use BUFFERS when running an EXPLAIN. It gives some data that may be crucial for the investigation.
2. Always, always try to get an Index Cond (called Index range scan in MySQL) instead of a Filter.
3. Always, always, always assume PostgreSQL and MySQL will behave differently. Because they do.
The article felt like it was fumbling around when the initial explain pointed right at the issue. I didn’t know this specific trick, so I did learn something I guess.
The query should be CreateAt > $1 OR (CreateAt = $1 (NOT $2 like in your sample) AND Id > $2)
So the idea is here about paginating through posts that might be constantly changing so you can't use a simple offset, as that would give you duplications along the way. So you try to use CreateAt, but it could be possible that CreateAt is equal to another one so you fallback to ID.
But here I stopped reading the blog post, because I now think why not use Id in the first place since it also seems to be auto increment since otherwise you couldn't really rely on it to be a fallback like this? I don't have time to investigate it further, but tldr; that still left me confused - why not use ID in the first place.
[edit] CreatedAt timestamp could be something from the client when the post is submitted or tagged from the ingest server and not when they actually are processed and hit the database
And I agree I shouldn't have said "auto increment".
(my guess is time based index offer faster search performance or lower overhead than a string based search index which doesn't understand it is representing encoded time data)
And with fallback they would end up reusing that index anyway.
But here the case should be batch indexing, processing, so it seems like auto incr with a timeout of assignment if those auto incr ranges are cached would still be suitable.
Like as I understand the problem, there is one service (ElasticSearch) that is working on indexing, and it's getting batched rows from postgres, to then index, but make sure at the same time to not miss any in those batches. And it's fine that it doesn't immediately at this second or minute do the indexing, so it should be fine to wait for the IDs to have been allocated.
Because it's ">" it might be missing that one record during what I think is pagination.
https://dba.stackexchange.com/questions/266405/does-ordering...
Or besides that, if there are odds of CreateAt collision, and you are fallbacking to ID, you are still possibly not getting it chronologically?
And also if CreateAt does happen to equal to another record that is exactly the case where Postgres might most likely not have the auto incr chronological.
So still it seems like the edge case it tries to prevent it would still happen at least at similar magnitude of odds.
edit: on second review if live insertions were occurring then this code would have an edge case, however the indexing job has an endtime presumably chosen where they can be sure no more inserts will occur. Given that the choice to use a timestamp probably has to do with the fact that there are 4 different tables being indexed and you would otherwise have to track their IDs individually.
original:
id, timestamp
1 , 1000
2 , 1001
4 , 1002
--------- limit stops here
5 , 1002
3 , 1003
ID only: If we use ID > 4 as our start point for next time and ID 3 was not inserted yet then we will have missed itcreatedAt only: If we use createdAt > 1002 then we will skip ID 5 next batch
OPs strategy: Even if we use createdAt > 1002 and it skips ID 5 it will be caught by createdAt = 1002 AND ID > 4. The Order by createdAt asc, id asc guarantees that if the limit splits 4 and 5 that we see 4 first and thus don't miss 5. I think this does still miss the case where ID 4 is inserted after ID 5 however.
Yeah, and I would think that is very likely to happen given it would be the same timestamp.
So the whole thing still seems like a flawed, and unnecessarily complex solution to me, which should just use one simple unique sortable field to do all of it.
Like my intuitive guess is that maybe the solution could save maximally 20% - 40% of the same edge case, which doesn't seem like a good solution. It is not going to solve the problem. It's just adding complexity that can cause other problems.
So if postgres does the type of caching where it allocates 10 auto incr IDs to each process, which causes sometimes IDs being out of order, then normally it would be just enough to wait after these allocations have performed and index then, you are not going to miss any rows.
I would assume these processes have some form of timeout if there was a case where they couldn't assign one of those IDs, and then this ID would just maybe not exist or if there was a mechanism to reallocate that would work too, but none the less, I think some form of postgres sortable unique id would have to work by itself.