Redis as a JSON store
redislabs.com
redislabs.com
For example, AFAIK there is still no agreement on how to handle binary data, e.g. what do you do with:
>>> json.dumps(["\xb8\xc3\xb6\xbb"])
Traceback (most recent call last):
...
UnicodeDecodeError: 'utf8' codec can't decode byte 0xb8 in position 0: invalid start byte
You can base64-encode it or do a number of other things, but then that's outside the JSON standard, and suddenly it isn't the flexible encoding standard it seemed at first...For the binary data: JSON is not a binary format, one should not try to mix those. To encode it and store it as a string is exactly right. Otherwise a database as blob store is the better place for that.
You can't easily append something to an already existing JSON file. You have to decode the whole thing first and then write it back whole. XML has the same problem. That's why I said we replaced one bad format with another.
JSON is not an improvement over XML. It replaced one set of weaknesses with another.
[1] I think they are cool because of their simplicity - it is just a two-level hierarchy - sections and options within sections. And options are key-value pairs. For many applications, this is enough, and any format that supports more may be overkill (though would still work).
Edited to add the [1] footnote.
And works very well as a document format, and JSON works very poorly as a document format! So if you have documents (a novel, an article, a manual, etc.), XML is still a very nice choice. But was a mistake to use XML as a data format.
(Just elaborating on your point.)
So something meant to be read by humans. XML is nice for marking up the actual text with structure like chapters, sections, footnotes, references, formatting, etc. (thus, the ML in XML comes from "markup language")
With JSON, there isn't really a notion of prose text to "mark up". It is structured data meant to be processed by machines.
(Of course it is useful for humans to read and write JSON sometimes, but JSON is primarily meant to be input to or output from some computing process. You wouldn't sit down to read a JSON doucment unless you are some super weird programmer dude.)
Want comments? Pipe through json5. Still vastly simpler than XML.
It's hard to take people's criticism seriously every time JSON comes up. There's a sibling comment that even suggests that it's a mistake to use JSON for config in the first place. C'mon.
The creator of JSON was explicitly against comments for a few reasons. The biggest one was that comments are often used to "extend" a format. Look at doc-block, annotations, etc...
In order for the format to work, everyone had to use the exact same set of rules, and that meant that if people started adding their own "crap" to the format in a "backwards compatible" way you'd lose what makes it special.
So he didn't add comments. And I personally think it worked out.
Right now, there are a number of different ways to encode both types of data, and you see them all in the wild. And extra custom serialization and deserialization steps are always required to handle them.
Actually, I was thinking something more like Protocol Buffers or Avro if you are looking for a binary format for transmitting structured data.
Unix timestamps are lossy.
1. No time zones: you can't recover the sender's time-zone from it. (Not the end of the world, obviously.)
2. No defined calendar: if you convert an arbitrary datetime into a timestamp, you don't actually have enough information to convert it back into a datetime, because the same instant of continuous time has different representations in different calendars.
3. The unix epoch is very recent, and there is no such thing as a "negative timestamp": you can't convert historical datetimes to timestamps. If you're trying to pass around e.g. medical records tagged with their dates, what do you do with the ones about things that happened before 1970?
(And without problem 3, problem 2 would be even worse: you'd have to care about calendars beyond the Julian and Gregorian ones, to represent those historical dates. Some calendars don't map monotonically to continuous time! Some instants have multiple representations in the same calendar! Some instants have no representation in a given calendar!)
But the worst problem of all, that affects you even if you don't care about encoding weird historical dates:
4. Timestamps don't "care about" leap-seconds [i.e. the standard doesn't force timestamps to either include or exclude them.] Therefore, every time there's a leap second, Unix time becomes less precise by one second (the integer representing a timestamp after a leap second could map to dt n, or dt n-1, depending on if [and exactly when!] the system that generated it adjusted its clock for the leap-second.) That means that right now, if an unknown system handed you a timestamp that purportedly represents the current time, you'd not actually be able to know that with less than a ±27 second error-bar. That number will only keep climbing.
> to encode it and store it as a string is exactly right.
That's not actually the problem. It'd be fine if JSON took binary data and then represented it in some human-readable encoding like Base64.
The problem is that JSON doesn't specify the representation of binary data, or the mapping between binary data and such a representation. So you can't take a binary buffer, drop it into a JSON serializer, and expect it to pop out the other side on any random JSON-speaking system as a binary buffer. Instead, both sides have to have known characteristics (an agreement to represent binaries in JSON a certain way.)
The whole point of a serialization format is to bundle up those guaranteed characteristics, so that once you know "this system speaks JSON", you don't have to ask any more questions. JSON fails at being a serialization format because you still have to ask the other side how it "expects to see" binaries represented, or dates represented.
> JSON fails at being a serialization format because you still have to ask the other side how it "expects to see" binaries represented, or dates represented.
That's not a failure, it is a feature. The moment you store binary stuff you always have to communicate with the other side about what exactly is represented and how. JSON does not make false promises. It stores a few main data types and a structure, and that's it.
This is the whole "Just use a string" vs "We need to have data types for everything". People who like Unix vs people who like PowerShell, dynamic vs static typing, hackers vs business types. It's a philosophy thing, not a "this is the right way" (apart from mine, obviously).
Other examples of presentation-layer encodings: ASN.1's encodings [DER, XER]; YAML; the Erlang External Term Format.
These encodings take a set of native types, and represent them with canonical encoded/serialized forms. They don't have to allow for faithful bijective encoding of application or domain types (although YAML does), but they are expected to be able to represent any reasonably-common "part of the runtime" scalar or container-ADT type which you might want to use to build your domain types out of.
Presentation-layer protocols are the basis of RPC-like protocols like REST or SOAP: you define an application-level message encoding to give semantic meaning (i.e. an application-level type) to terms which have passed through an RPC channel, where those terms arrive in your runtime after RPC decoding with types guaranteed by the presentation-layer protocol you've chosen.
JSON is very limited as a presentation-layer protocol, because it guarantees representations for very few types, and many types that are needed to effectively implement RPC-like protocols are not in that small set. This is recognized by some, as there have been a good few attempts to define a presentation-layer encoding "on top of" JSON rather than just using JSON itself... but almost everybody ignores these [fine] layer-6 encodings in favor of continuing to create "JSON APIs" (i.e. layer-7 protocols that specify JSON itself as their layer-6 carrier.)
The creators of these JSON APIs then find themselves forced to specify how their particular API represents {binaries, datetimes, sets, strict maps, exceptions, UUIDs, ...}—in other words, to define their own one-off layer-6 encoding atop JSON. And the authors of clients for these APIs find themselves forced to write their own logic to parse and generate representations fitting these specifications—even though they're almost always exactly the same choices as every other API creator made.
Choosing a different, richer layer-6 encoding allows both the API creator, and the client authors, to just avoid all that work. Instead, the people who write tooling conforming to the standard for each runtime will do that work, once, and all the API designers and API-client implementors get to benefit by relying on that tooling. Which is rather the point of a standard.
A nitpick, but, REST isn't RPC-like, and is an architectural style, not a protocol.
There's a sort of de-facto layer-6 protocol in the way people use REST in combination with JSON [usually using AJAX calls in browser SPAs] to do RPC, by PUTing or POSTing JSON documents to REST endpoints, and then reading back JSON response bodies. It's a "protocol" where HTTP is doing some of the presentation-layer type-encoding, the "application/x-www-form-urlencoded" encoder is doing some more (for e.g. GET query parameters), and JSON is doing the rest.
The "browser-like JSON over REST" RPC approach is so common that it's nearly synonymous with "REST." If you say "we've created a RESTful API", you usually mean that you've defined a layer-7 protocol where browsers use AJAX to make JSON-bodied requests to a hierarchy of REST-architected endpoints, expecting JSON-bodied responses.
(There's also JSON-RPC, which a clearer example of a standalone layer-6 protocol with the same problems, but it's not in wide-enough use to really matter.)
The real benefit to JSON is that it's so lean and easy to parse. I can get a hashmap or a plain object out of some JSON text in a few lines of code in almost any language, the same can't really be said of XML, especially not the kind of XML usually used as a data interchange. Additionally, that lean-ness means it has significantly less data to transfer, I've seen JSON versions of data be 1/10 the size of the XML equivalent, which can matter at the far end of the spectrum.
The real strength of JSON, though, is arrays. List structures have always been a major weak point of XML, as there's no way to do it without a million tags. God help you if your list has anything other than a primitive in it.
So, what do or would [1] you consider a good format. Asking because I'm interested in data formats. They are inputs for my xtopdf toolkit, plus I'm generally interested in data munging and have done it a good amount.
[1] "do" for existing ones, "would" for non-existent ones.
Note that I think many people skim that too quickly and miss the point. It is more an explanation of why the "worse" things so often seem to beat out the "better" things than simply a cynical statement or something.
I'm sure there is a ton of code out there that doesn't take this precision issue into account. It seems like Twitter ran into that issue at some point, because in their APIs they have both id and id_str (see [2]).
[2] https://dev.twitter.com/overview/api/twitter-ids-json-and-sn...
Strings in javascript/json are UTF-8, UTF-16, or UTF-32 encoded[0]. So providing a non-Unicode string is going to fail.
Instead, use an array of numbers limited to the byte range (as you would for any other binary data).
console.log(JSON.stringify({buffer: [0xb8, 0xc3, 0xb6, 0xbb]}))
//{"buffer":[184,195,182,187]}
[0] - http://rfc7159.net/rfc7159#rfc.section.8.1However, the original commenter was worried about stuff "that's outside the JSON standard,". Storing byte-sized numbers in an array is well within the JSON standard, and any system that could consume json would be able to consume it without issue, and without additional dependencies.
Uh, then don't use it. Nobody said JSON was the answer to everything.
And even if you get to pick your own infrastructure, a lot of development effort is "stolen" by JSON-focused tooling and infrastructure, such that there's very little development effort given to protocols that aren't as limited.
Consider: how many languages that offer a "batteries-included" JSON parser+generator, also offer a "batteries-included" parser+generator for e.g. ASN.1?
Also, much like shoving large blobs in most databases, maybe you shouldn't be shoving large blobs into JSON?
JSON's fine if you don't have any requirements around data serialization and you want it to "just work" for your webapp, but there's a lot of tech debt inherent in it.
So you dump a report in JSON format and back it up to S3. S3 costs are growing faster than you thought, so you gzip deflate all of it. Everyone has to go patch their JSON deserialization to detect gzip extensions. Whatever, just growing pains.
Then another team tries to read the reports, and they're getting errors because your definition of an interface is "we'll just use JSON, the keys are human-readable".
You define a formal API for your report format and in doing so you realize the need for versioning attached to your report schema, so you wrap all your JSON objects with types and version annotations. You could define a central repository for these schemas, but it's easier to just bake them into the top-level response. Everyone agrees that this is "lightweight" and not "centralized".
Now you're storing reports where each sub-object has its own annotations, or you're defining an entire schema at the object level. Object deserialization is taking 200ms even for small payloads, because of all the validation callbacks you're firing, and developers are now "performance hacking" their components by disabling validation callbacks. Now you have all the space overhead of schematic annotations with none of the benefits.
In order to adhere to the API, either teams are writing separate serialization libraries, or you form a team to maintain them as infrastructure, which is a great idea except that the horse already left the barn 2 years ago.
Without even realizing it, you've reinvented XML and XSD. And I don't really like XML either but at least you have to be honest about what you're getting yourself into.
Wouldn't it be annoying if numbers and booleans didn't exist in JSON, you just used strings? Same thing, JSON has a few data types it just missing a very common one.
Somewhat surprisingly, human readable dates are an absolute mess of historical convention and complicated geophysics. To write a human readable date, at the most basic level, requires us to account for both calendar for and time zone.
But even that isn't enough. To get these display dates to match up with everything else in the world we also have to worry about a whole smorgasboard of national holidays, differing leap year conventions, leap seconds, historical mishaps, typographic conventions, and host of even more obscure minutia.
All this complexity obscures from the programmer that human readable dates, for the above reasons, are almost useless at reliably determining time sequences.
Instead, we should think of dates as UI details on the same level as a color scheme. Usually what we want is better thought of as a point in a time sequence which is best represented as ITA timestamps, e.g. Unix Time.
I have seen so many takes on xml and soap over the years. Every one of them made me sneer in disgust.
- it's a lot like XML but with less red tape. Looks nicer to humans. Also allows expressing the semantic difference between an (ordered) array and a map.
- it maps extremely well to Javascript, and a lot of APIs are consumed by web clients written in Javascript these days
I think the latter is the real explanation.
That being said, I agree that JSON is not a panacea. My personal pain points:
1) Missing an elegant way to express an unordered set. IMHO the three basic collection data types are array, set, and map. I think JSON would be stronger if it could express a set like you can do in ES these days: { "a", "b " }
2) Not really DRY if you encode a large array of objects and sned them over the wire. So yeah, your API should support more compact formats as well.
3) Allowing comments would be really nice when we use JSON for e.g. package.json. I don't think it would make the parsers any slower so why not? XML has it.
Well there's this: JavaScript Object Notation (JSON) Pointer RFC https://tools.ietf.org/html/rfc6901
Once we'll add indexing I'm guessing the path expressions will need to become more complicated, and we'll try to choose a "standard" and stick with it.
Until now you couldn't do it without loading the entire object into memory. ReJSON just allows you to manipulate that object or retrieve just a part of it efficiently. For this use case we do not need and will never need indexing.
That is, I have some application support from Lua scripts manipulating native types today. But if I can read and write into JSON stored this way from Lua, that might be very handy!
> Windows
> Yeah, right :)
How about just "There are no plans to build a Windows version"?
But if you're storing say, a 1K user profile, and most of the time you just want to update the access time, or get some auth token - this both allows atomicity and speeds up what you are doing.
The bigger the JSON object is vs. the how small a piece you want to retrieve or manipulate - the more efficient this becomes.
Also, manipulating a field in a JSON object atomically is currently possible only with Lua scripts - and even then, Lua will have to parse, manipulate and serialize the entire object. Where as in this case, if you are just updating something - there is zero parsing and zero serialization going on in the module.
A talk by the author on youtube: https://www.youtube.com/watch?v=NLRbq2FtcIk
The ReJSON website: https://redislabsmodules.github.io/rejson/
Instead of disable all of ublock you can toggle the cosmetic filtering when you click on the ublock icon. http://i.imgur.com/Ne9mHBd.png