Once you introduce any mildly interesting vector or typesetting feature, you will not get pixel perfection. Examples:
* You want to draw a circle. Antialiasing is desirable because jaggies look really ugly. Do you want to prescribe the exact antialiasing algorithm? What about other ones that are faster or more accurate? Which AA convolution kernel are you using - box, Gaussian, sinc?
* Which approximation algorithm will you use to draw cubic Bézier curves? I don't think there is a closed-form solution for them.
* You support affine transformations. Define a 10×10 square. Scale it by 1/100. Scale it by 100. What if, due to floating-point rounding errors, your square is now 9.999×9.999 and now renders to 9×9 pixels without antialiasing?
* You support automatic text wrapping. One renderer uses float32, looks at how wide each word is, and decides to break at some point in the sentence. Another renderer uses float64, looks at how wide each word is, and decides to break at a different point in the sentence.
I downloaded their example file and it's entirely binary, unlike PDF which is just pseudo-binary, but Evince did open it and it seems unlike PDF it's entirely raster based and would require a separate OCR layer on top of the text to make it eligible for copy-paste, if that's one of the goals
Not really all that different from PDF:
"Like PDF, DjVu can contain an OCR text layer, making it easy to perform copy and paste and text search operations." https://en.wikipedia.org/wiki/DjVu
DjVu however really does seem to be biased toward scanned-origin documents, not digitally produced ones.
That is not correct, PDF supports text spans natively; perhaps you're thinking of scanner software that merely uses PDF as a convenient packaging for their JPEGs?
I cannot defend DjVu as I've only had tangential contact with it, and for sure have never tried to author any such file. I was just raising awareness that there are competing standards that appear to be libre and are designed for pixel perfect output