The Order of the JSON
blog.almaer.com
blog.almaer.com
You can’t just flip that switch, run ”a client to post the JSON to that instance”, see that test “worked just fine”, and call it a day.
The POST might just store it, for later processing to wreak havoc (say by ignoring a value that isn’t in the expected place in the JSON), or only rarely seen JSONs might cause problems, or it might ‘only’ break the yearly run, etc.
This change made parsing more lenient which is generally considered fine [1] and made the software actually follow the spec, since it specifies object entries are unordered.
How do you know? It makes the outermost layer of the system accept more liberal in what it accepts, but that doesn’t guarantee that the inner layers can handle that more liberal content.
For example, the code may assume that “<user-ID>” always is the first node in the json. If you start sending “<car-ID>” instead, things may go fine until you get two customers who share a car.
”and made the software actually follow the spec”
How do you know? The spec of this piece of software may state it has a JSON-like API that requires the “<user-ID>” to be the first node in every request.
And yes, most of its code may handle that change fine, but it only takes one piece of code to break things.
”How do you justify ever changing anything with that attitude?”
In the case of ”a service running IBM DataPower Gateway which sat on top of WebSphere which sat on top of the COBOL.” where ”Much of the system was so old that it was hard to find anyone who knew how it actually worked, and it’s maintenance had been outsourced ”: very carefully.
This change may have been fine, but you have to check that.
The robustness principle is a powerful guideline for interface design; not an excuse to turn off validations that someone else has put in place. If the person who designed the interface wasn't ready for something, principles won't make the software work.
OrderedJSON and JSON might seem compatible on the surface but are you sure you caught every little code path that might have been assuming OrderedJSON?
This isn't an issue of spec compliance or not because the data format wasn't JSON in the first place. OrderedJSON is not a subset of JSON even though it has the property that all OrderedJSON documents are syntactically valid JSON -- they represent different abstract values and JSON is the lossy interpretation.
By rolling out the change over a period of time to a small sample of users and comparing the experimental effect to a control sample.
I was struggling with how to do it correctly, and then one of the people in the r7rs working group simplified the problem by a very large factor by pointing out to me that I was trying to solve a non-existant problem, because I tried to add definitions to expression contexts where it simply made no sense. In fact, instead of supporting all forms, I could get a better result by just focusing on one of them and add simple wrappers for another 5.
It was all quite humbling. Had I taken a step back and actually analyzed the problem I would have come to the same solution, but I immediately tried the, to me, most obvious and also hard-to-get-right solution.
It's not a paradox, it's a matter if excess focus - once I get into a rut I just keep bulldozing[0], but coming to a problem fresh it's often trivial because the rut hasn't formed.
I'd say it gets better with experience, and you're probably better than you think - my guess is you don't see your successes as clearly as your failures. I guess you spend so much time on the failures, but if you instantly perceive the right approach to a hard problem and solve it in a shot, well, it wasn't a hard problem, right? A kind of bias of perception.
[0]I get the impression that's a bit of a man thing generally
This condescending perspective on problem solving is part of the problem with this entire industry and ultimately turns people into tools for disposal by businesses.
I'm not bitter at all...
(And if things did go wrong, we had pretty much the ultimate backup plan: revert!)
It gradually dawned on me that his own code (which I thought was generally low quality) did generally reflect his values here: it was write-once. Any subsequent change just layer on and patched around.
> I think stuff likes this is the real price of not hiring really good people.
Yup. And I still am not sure I'm good enough to tell them apart in hiring without asking questions that the candidate will just tell you what you want to hear.
I think stuff like this is why decent developers don't become "really good" when working in these environments. If you have a culture of letting people figure things out, and forgiving mistakes, then you'll end up with better engineers. On the other hand if you have a culture of protecting "territory", and blame then you'll get 6 months for 9 people to check a box.
Totally agree. I feel bad for a lot of the young devs who are exposed to daily micromanagement and stand ups and the pressure to constantly deliver something. They don’t get the opportunity to make mistakes and learn from them.
I must admit I was scratching my head on this.
The JSON spec might not specify order, but the serialized JSON is ordered by nature. JWT needs it for example (and I'd assume many signature models). Or you might have a caching layer that needs it. Maybe unchecking this causes the legacy backend to get hammered? There are valid non-spec concerns with order.
Replies here seem to assume stupidity here. It's a valid reason, but it's not the only one. Equally, the author doesn't ask "why" - why would it take so long, why was that option enabled?
>There exists in such a case a certain institution or law; let us say, for the sake of simplicity, a fence or gate erected across a road. The more modern type of reformer goes gaily up to it and says, 'I don't see the use of this; let us clear it away.' To which the more intelligent type of reformer will do well to answer: 'If you don't see the use of it, I certainly won't let you clear it away. Go away and think. Then, when you can come back and tell me that you do see the use of it, I may allow you to destroy it.'
[1]: https://en.m.wikipedia.org/wiki/Wikipedia:Chesterton%27s_fen...
Agree it's a leaky abstraction. Equally agree it should be documented. Often legacy systems have weirdness after seeing decades of edge-cases. Weirdness that makes them robust in all sorts of unlikely ways. Equally makes them poorly documented, leaky, opaque, and frustrating.
But I certainly have a lot of pause with the sentiment of "some idiot checked the JSON order checkbox on DataPower". My first thought is instead "I wonder why someone thought this was necessary."
If it's not JSON, it's undefined when it uses JSON tools and libraries, and anything goes.
If it's not JSON, document it as your own dialect, and use your own tools, and try to never talk to the outside world based on it with systems that treat it as JSON.
>But I certainly have a lot of pause with the sentiment of "some idiot checked the JSON order checkbox on DataPower". My first thought is instead "I wonder why someone thought this was necessary."
It might not be, it could be the BS default setting...
1) the gateway or more likely something behind it starts crashing for memory errors when someone is serializing some stupidly large array into it
2) using ordered vs non-ordered keys on input changes the behavior and fixes it
3) enabling this checkbox to prevent it fixes the issue and guards against it happening again
4) everyone has forgotten about this web weenies inquiry into the issue
5) due to #4, it takes like 3 weeks of system crashes and hairy debugging to find the cause
6) this guy is long gone, having ridden off into the sunset of smugness
Also:
1) fixing the serialization-dependent memory allocation 4 layers deep on the back-end to allow things to operate safely without this checkbox requires a complex change to several other components and associated system-wide validation testing which would take: drumroll "9 people 6 months to complete"
2) What is the point of this javadoc reference and how exactly does it relate to the issue at hand? Here's a random code doc reference too!: https://pymotw.com/3/collections/ordereddict.html and here also a discussion about impacts: https://mail.python.org/pipermail/python-dev/2016-September/...
1) Nothing happened, this already went down years ago, as the author writes.
2) The setting was only stopping the BS IBM system layer, wasn't really needed elsewhere.
3) The BS setting was probably even default, e.g. not consciously enabled by someone for the specific system.
>What is the point of this javadoc reference and how exactly does it relate to the issue at hand?
The point is that the IBM system who had this setting was using that BS, brain-dead, not-JSON, implementation to handle the order, when the setting was enabled.
You almost never need canonical representations for signing things. I would even say that if you need a canonical representation to sign your things, then that is a design smell of your cryptographic protocol.
Whether the file has an order is an implementation detail. There might be no static file at all, for example, it could be an endpoint returning you the JSON, and it could return different orders for the exact same items every time you call it. That would still be perfectly valid.
If you "need it", don't put your data into an object, put it in a list of objects.
If the underlying layers need "ordered objects" (in other words, ordered hashmaps) they don't really use JSON.
The mistake is assuming that two syntactically compatible protocols were actually the same protocol. uint =! int.
Was it a HTML Gui with file upload and everything? Or just some server, port and proprietary protocol?
Maybe these things are on Ebay for cheap. Oh, that looks like fun.
I already had at least four solid reasons not to use JWT. Still, adding another just keeps the dumpster fire burning.
By which I mean, in the most oblique and backhanded fashion, that there’s not going to be any reason for validating JSON order that doesn’t ultimately reveal or confirm a design gaffe somewhere down the rabbit hole.
As soon as you read IBM you know you have bigger problems than COBOL in your system.
My own theory is that most (not all) of technology is there to support BS jobs and they are irrelevant for the normal functioning of society. It simply doesn't matter when (most) technology breaks.
And, in fact, I do use ordered JSON for comparison in testing, as you describe.
However... comparison of JSON as objects (i.e. in memory) is order independent. Hashcode is also order independent (the trick is to sum the elements' hashcodes e.g. https://docs.oracle.com/javase/7/docs/api/java/util/Set.html...
Diff for ordered JSON is possible using longest common subsequence for trees, but has terrible complexity, and lacks diff's clever optimizations both general and specific to typical input.
And if it doesn't impact performance significantly, these are all pretty good reasons for JSON outputters to default to sorting objects deterministically by keys, or at least to provide a flag to do so. (Even if there's no canonical sort order between JSON libraries, all that matters is it's deterministic for each library.)
BUT... I can't imagine any scenario where you'd want to validate that JSON content was ordered on the input, which is what was enabled in this article. Why does IBM even have that as an option?!
Be strict in what you emit and liberal in what you accept, and all that...
If you're going to rely on a specific order for comparisons, it makes sense to alert the user to any JSON in a different order (instead of silently, liberally accepting it), or you'll get false negatives elsewhere. Easier to check for a sorted order, but also possible to define a specific order. IDK what IBM did here.
funfact: jq used to sort keys; now it retains ordering.
So the rework time might be to write a general purpose re-order layer that can re-order any imcoming message.
Because it's a precondition for something else down the line (a dependency)
This assumes of course that you're using proper hashes that make use of the full domain of the output type (a proper hash will have a 50% chance of any arbitrary bit being flipped by any change to the input). But if you're not using proper hashes, you're doing something wrong.
If I have a 32 bit current hash value-- for any possible 32 bit value I add, I get a different 32 bit value out.
XORing is effectively adding each bit and throwing away the carry bit. Adding just cascades carries to the left.
It still feels wrong to say this, it feels like since adding will effectively shove bits off the high end and drop them on the floor that you're losing information, but I can't actually justify that feeling with reasoning.
(XORing is effectively adding with all of the carry information lost/falling off).
Similarly, sometimes it might be useful to store some data in a different character encoding, or maybe multi-character glyphs are stored in a different normalization form. (I can't tell "ü" from "ü" just by looking.) Or maybe your sample data has 1.0 but your program generates 1.000. There's a million ways that serialized structures can be functionally identical but quite different. Easy: don't compare raw bytes.
If you want functionally-equivalent data to be accepted, you just need an equality-tester (and hashing algorithm) that's agnostic to such issues. They're not hard to write.
Some years ago, I worked on an antitrust case involving a firm that had outsourced relevant systems. Multiple times, to different IT firms. And there was literally nobody left who knew how they worked.
After considerable negotiation, they agreed to provide documentation. And what that ended up being was a report by IT company 2 about their understanding of what IT company 1 had done with the firm's systems. Because, I gather, IT company 1 had evaporated.
And yes, the core of it was COBOL.
HMAC-SHA256(
b64(reencoded_header) + '.' + b64(payload),
secret
)
You should really just verify the signature for the provided header + payload in their base64 encoded form.One example I struggled with for days recently was USPS. Not only they use xml in the url parameter, the order of the elements also matters. Unfortunately, the order in the documentation is incorrect.
I haven't had a chance to look at why yet.
[1] https://stackoverflow.com/questions/5525795/does-javascript-...
Or perhaps they're so used to getting away with quoting 6 month for a 6 hour job, it never crossed their mind not to.
and with app like fluented, we can decorate it with whatever metadata we want (instance name, machine type, container name, etc..)
It's a security issue, not just convenience. With unsorted maps the internal hash seed can be exposed, together with timing information.
Another famous omission from the specs.
On the other hand, accepting unsorted maps seems like it could introduce covert channels?
They can be DOS'ed.
> What timing information would be leaked and why would that be problem?
You misunderstood. By 1. leaking the order, and 2. by checking the timing of getting certain keys of a map you do have two independent infos to get at the secret seed.
> accepting unsorted maps seems like it could introduce covert channels?
A covert channel might be exposed by the sender, if he uses some non-random but unsorted map order. The receiver needs to accept any order. There's no risk in accepting unsorted maps. The only risk at the receiver side is another famous omission from the spec: How to deal with duplicate keys. Accept (overwrite or drop) or reject? All 3 cases can be seen in the wild.
This led to hash randomization on by default since Python 3.3.
The backend polled an API that served XML updates on game scores etc. Note that I didn't previously know anything about baseball (and I still don't, not really).
So let's say that a baseball team scores a Double. We'd see some XML like <doubles><double/><double/></doubles>.
Now, suppose a team scores a tripple... You know that there's a <triples></triples> entity. What would you expect to see inside the node?
If you're me, you'd expect to parse <triples><triple/><triple/></triples>.
I got the call during dinner. There was an important game happening, and the app suddenly broke. People were uptight. They loved our app and they were complaining.
In the end, it turns out that we would need to process <triples><double/><double/></triples>. Why? "Oh, it's always been that way." (Says the brusque developer at the service charging $50k/month for access to this feed.)
> This so far out of the spec it makes my ankles hurt.
This is not in fact out of spec. The JSON spec does not define semantics here but instead quite explicitly leaves it up to the JSON processor and data interchange spec for what to do about ordering of objects.
What? It is literally on the first page of json.org and section 1 of RFC 7159 [1]
> An object is an unordered set of name/value pairs.
> An object is an unordered collection of zero or more name/value pairs
>The JSON syntax does not impose any restrictions on the strings used as names, does not require that name strings be unique, and does not assign any significance to the ordering of name/value pairs. These are all semantic considerations that may be defined by JSON processors or in specifications defining specific uses of JSON for data interchange.
Also it's pretty clear that the keys must have an order going over the wire. And JavaScript objects (which aren't unrelated to JSON have an order now) https://www.stefanjudis.com/today-i-learned/property-order-i...
https://fasterxml.github.io/jackson-annotations/javadoc/2.3....
The result is that all implementations are broken, f.ex. this is how a tree structure has to be implemented with JSON:
http://root.rupy.se/meta/user/task/eelzter/44953781393584543...
It's inefficient, ugly and just wrong.
Every time the issue surfaces it elicits a large amount of eye rolling among the cognoscenti. However, I'm a firm believer in the idea that recurring demands from the user base need to be addressed, rather then mocked. The coding community appears to me notably tone deaf on this.
'Course anybody can use <insert suitable technique> to send a JSON map in the desired order, except the parser on the other hand will blissfully disregard it. Or you can devise any ordered solution on both ends, which will leave you outside of the accepted standard and open you up to the situation described in the article (where the requirement may well have been frivolous).
I can remember very few - if any - instances of somebody implying that the standard should somehow make room for this kind of scenarios. The standard reply was "use a different format" or "change the requirement". (Somehow remembers me of people asking what can one do to have whitespace-preserving XML, and at least one amusing story about that)
[["Key1", "value1"], ["key2", ...
You just need to deserialise it into a specific collection on the receiver. It's within standards and there are no weird parser issues.Consider serialized JSON already has a natural ordering, which is thrown away for some reason (probably parsing convenience).
It's got nothing to do with parsing. The json comes from JavaScript syntax, where objects represent unordered mapping. Json naturally does the same.
If you want expressive power for reading, json is very poor in comparison to pretty much everything else. Just use it for simple serialisation.
That's not true. The spec at http://www.json.org/ says "An object is an unordered set of name/value pairs".
If your application expects JSON input to conform to one particular set of decision points here, that’s fine and useful but you’re not using JSON anymore. You’re using a JSON-compatible subset. Document that new thing you’ve made and stop calling it JSON; because it’s not.