I don't really feel like playing the quotes game... but, sure.
"All of the issues we found had to do with tablet migrations"
"ndeed, the work Dgraph has undertaken in the last 18 months has dramatically improved safety. In 1.0.2, Jepsen tests routinely observed safety issues even in healthy clusters. In 1.1.1, tests with healthy clusters, clock skew, process kills, and network partitions all passed. Only tablet moves appeared susceptible to safety problems."
No one is here to claim that anyone is getting through any kind of rigorous testing without bugs found. But there is a huge difference between "My extremely common write path + a partition = dropped transactional writes" and "Under very specific circumstances, with worst case testing, multiple partitions, and the db in a specific state, we drop writes".
There is an ocean between, say, mongodb's test results, and Dgraph's.
Read Redis's evaluation, for example:
https://aphyr.com/posts/283-call-me-maybe-redis
"If you use Redis as a queue, it can drop enqueued items. However, it can also re-enqueue items which were removed. "
"f you use Redis as a database, be prepared for clients to disagree about the state of the system. Batch operations will still be atomic (I think), but you’ll have no inter-write linearizability, which almost all applications implicitly rely on."
"Because Redis does not have a consensus protocol for writes, it can’t be CP. Because it relies on quorums to promote secondaries, it can’t be AP. What it can be is fast, and that’s an excellent property for a weakly consistent best-effort service, like a cache."
Again, Redis is a very different type of database, so expectations should be aligned. Further, this test is quite old.
But that's a huge difference from DGraph's results.
Basically, saying "Well no one does well on Jepsen" isn't really true. Lots of databases do well, but you have to adjust your definition of "do well".