could the long s `ſ` be throwing the model off?
I think the simple explanation is the likely one (the reason I deliberately chose this specific benchmark): the model isn't intelligent enough to figure out use/mention distinctions. It understands Voltaire is discussing injustice, violence, tolerance; but it doesn't understand which side he's on.
I guess that would make sense for such a small model to be missing the subtlety