Ren: a lightweight data-exchange text format
pointillistic.com
pointillistic.com
The Ren site (ren-data.org) redirects there, is quite old now, and was a playground and experimental area. Hence the empty links and such.
In addition to Bolek's Humanistic repo, I had set up https://github.com/Ren-data/Ren to discuss ideas. Ren is, effectively, Redbol (Red+Rebol). One of the initial goals was to define a subset of values and normalize the syntax (Rebol never formalized its format spec), which could be shared across Redbol langs as they evolved and went in different directions. And also as a bridge for loaders in other languages. JSON has taught us that a small spec is important. The balance between simplicity and expressive value types is key.
It may come back to life at some point, but my time was better spent elsewhere for a while. I'm focused on Red now (red-lang.org). It has a native bridge feature for embedding in other langs (https://doc.red-lang.org/en/libred.html), along with a lot more. I believe there is still value in formalizing the grammar so others can create their own implementations, but it's not a priority at this time. In the meantime, you can find the active Red community at https://gitter.im/red/red to get more information and examples of what it looks like in use.
Cheers.
There could be a base set of standard types and ideally an organizational process to add more standard types -- like mimetypes. Mimetypes also might be considered included by default perhaps like "application/json; charset=utf-8:[1, 2, 3]").
It is more characters to include a type for each primitive, but it is more expandable. Likely these files will mostly be generated and read by code anyway, with humans just looking at them now and then for debugging.
If that string version seems too cluttery, another option is something like: rational:12/17 and string:foo and string:"this has spaces".
Or given possible confusion with maps and colons, another option is rational/12/17 and string/foo and string/"this has spaces" and complex/−1+3i and real:30.564 and "application/json; charset=utf-8"/[1,2,3] and maybe even xml/<foo>bar\ baz</foo> and javascript/console.log("hello") and so on.
Or maybe pipes? Like: rational|12/17 and string|foo and string|"this has spaces" and complex|−1+3i and real|30.564 and "application/json; charset=utf-8"|[1,2,3] and maybe even xml|<foo>bar\ baz</foo> and javascript|console.log("hello") and so on.
Children: none
Opinions: none
That certainly looks pleasant and human-readable, but a bit of a nightmare to interpret! Couldn't it just as well be "Children: 0" or "Children: []"? If the idea is to let non-technical people edit configuration files, the code reading the file will have to be very flexible and forgiving.(Edit to add: maybe that's something you mostly get for free in REBOL? It would be a major headache in most other languages though)
'unknown' is almost like saying "We don't know what should be here", while 'none' is closer to saying "We didn't get passed any information here".
Ultimately, I feel like 'unknown' is more specific, it conveys intent beyond that of 'none'. Such as, it could imply we actually don't know what kind of data 'Opinions' should hold, rather than that there merely happens to be no data to put there.
Cookies in the jar: none
and
Cookies in the jar: unknown
The language concept is a bit more fundamental than config files. Rebol (and Ren-C and cousin Red) is a fully homoiconic language that itself has tools for parsing and interpreting DSLs ('Dialects' to use the language's nomenclature) that use the same language rules. In essence you do indeed get the lexer for free.
Would recommend a wee read of: http://blog.hostilefork.com/why-rebol-red-parse-cool/
I've argued that Ren is somewhat redundant as it fills the same space that Rebol does though Rebol lacks a formal specification at this time.
Sometimes you want to comment a section out of a JSON without deleting it. Other times you want to annotate some generated JSON.
And because this is a common need, you have ad-hoc unofficial solutions which are not supported by all parsers.
Quite frankly, I think this is an area he made the wrong choice on. Which is fine, but still annoying. Like checked exceptions. Logic for the choice was sound, if misguided. End result sucks.
These are all integers. There only needs to be one integer type for all of: 0 6 -6 8000000000000000000000. Python gets this right.
Or you can go the JSON route, mumble your way through the numbers spec, let every language do its own thing, and tell the handful of people seriously interested in moving integers too large to be precisely specified by a 64-bit float to encode as strings or something and stop worrying everyone else with the complexity....
I don't follow this one. If your language, static or otherwise, is capable of turning this JSON:
"100"
into a string, and this: 100
into a number of whatever type, then it's also capable of turning this hypothetical input: 100
into one type of integer, and this: 100000000000000000000000000000000000000000000000000000000
into another type of integer.OTOH, exact decimals (of which integers are a subset) differ fundamentally in meaning from limited precision binary (or decimal, though that's more rarely encountered) floating point approximations. Which matters in interchange as well as schema.
JSON.parse('[9223372036854775805]')[0] === JSON.parse('[9223372036854775806]')[0]
You should get: <- true
Is chrome's implementation broken? Those are both numbers that a signed int64 can store precisely but a double float cannot.Postgres is the only popular system with built-in JSON functionality that I know of to correctly round-trip JSON numeric data. Python comes close but fails for any number with a decimal point or in E-notation.
One can argue that the JSON spec is too permissive (and I would disagree, though I'm somewhat a purist). Or a pedant could note that the JSON spec doesn't actually say whether any of the digits in a number are considered to carry information (and does in fact note that many parsers are faulty). But it's unfortunately true that most popular JSON parsers fail to round-trip valid JSON data due to flawed design.
[1] http://json.org/
RFC 7159 not only specifies that all of the digits carry information, it specified exactly what information they carry.
However, it also expressly permits implementations to limit the range and precision of numbers accepted, recommending (but not requiring) range and precision at least equivalent to IEEE 754 float64 be supported.
Can you quote? I don't see where it does, except by reference to common knowledge.
But I should clarify. There are several inefficiencies in JSON numeric representation which may or may not be considered significant by an application. The ones I can think of, in order from "obviously not" to "well, maybe":
1. the case of "e" vs. "E"
2. the optional "+" sign after "e" or "E"
3. presence/absence of decimal point and/or E-notation (shouldn't make a difference, but does in many parsers, such as Python's)
4. the value of the exponent itself (e.g. 3.14 vs. 314e-2)
5. "-" sign in front of any number with a value of 0 (IEEE floats and ones-complement integers do have a negative zero)
6. excess trailing 0 digits after the decimal place (may be used to represent significant figures in scientific applications)
7. digits of lesser significance (obviously the most contentious)
JavaScript ignores all but #5. Python ignores all but #3, and due to #3, sometimes #5 and #7. PostgreSQL `json` ignores none; `jsonb` ignores all but #6 and #7. Personally I would draw the line between #4 and #5. But neither the RFC nor the ECMA spec tell us.
Chrome’s implementation is fine (within spec) and makes the natural choice for a general purpose JSON implementation for JS.
Javascript, for example, has this problem because everything is just a "number" and the number data type is a IEEE-754 double precision float. A double precision float can't represent all int64 values, so transmitting an int64 via JSON to Javascript and using the standard JSON parser is lossy. You could say that Javascript screwed up by making number have a specific precision, and you might be right. Python, for example, gives you seamless arbitrary precision arithmetic.
But the decision to use an arbitrary precision type for all numbers in every language is definitely not appropriate. Arbitrary precision has trade offs, and sometimes it's important to work with machine integers.
"Deep embedding" of JSON data into the host language is the cause of this issue. Recognition that JSON is a separate data type and does not always have an obvious encoding in the native language (including JSON arrays, objects, and null, even if they look like native arrays, objects, and null when you squint at them) is key.
The only widespread JSON implementation I know of to get this right is Postgres's: JSON values live as the "json" or "jsonb" type until you explicitly convert them. No data loss is incurred, with the exception that "jsonb" normalizes E-notation, conflates 0 and -0, drops duplicate object keys, and normalizes their order. Even numeric significant 0 digits are preserved.
(one of my favorite things about Perl 6 is that non-integer literals are rationals by default; floating-point is only used if the literal is specified in scientific notation)
Real world values are still typed.
"If you need more control ... metadata is your friend."
aka implicit user typedefs. Explicit is better.
{ "int": 1, "float": 1.0 }
Of course it's up to the parser implementation to interpret 1.0 as a float.
The spec is very relaxed about the number type:
I think first version of Rebol interpreter was written in Scheme ;)
It means humaneness, benevolence, kindness, and is pronounced exactly the same as 人 which means person/people.
Also, the links under Implementations all result in a 404 page.
Of course "$" is used as the symbol for several currencies, but its not clear if you have a better symbol in mind for the generic concept of money.
(The idea that in Rebol one might write EUR$1.00 to denote the value that would usually be written 1.00€ is also pretty horrible)
1) Do you agree that a notation should support currency values? That is, they are useful to identify in data, as atomic units?
2) If so, do you agree there should be a single standard symbol and lexical form used to identify them? Because if we don't do that, we have to support every localized notation, correct?
3) If you agree to both of those questions, what symbol do you suggest? ¤ is generic, but not on any keyboard layout I know of. Also, see Chris's note about ASCII priority.
Finnish (/ Swedish) keyboards have ¤ on `<Shift> 4` ($ is `<Alt Gr> 4`).
With that background in mind, I imagine the scenario here you have two data in a Ren document which are both the literal $1.00, but one is actually 1.00USD and the other is actually 1.00EUR: it doesn't prevent errors (for instance, when you want to perform an operation like + on the two data), because you still don't know what the data means. You have gained very little over just using the literal 1.00 instead.
So if I were making a proposal I'd be tempted to suggest a syntax like [1.00 USD], and maybe even giving up one of the remaining sigils ^[1.00 USD] if it is important to raise to being a special element in the syntax of a Ren file. Now that you're saying what you mean, you can use the same syntax for all units: ^[1.00 kg m -2] (1 kilogram per square meter), ^[1.00 V hz -.5] (1 volt per square-root-hertz, a typical units specification of noise in opamps).
And while this may be more beneficial in your work, a lot of software does have to deal with money, where it's important not to use floating point, but BCD or something else.
US: 325,365,189
Canada: 35,151,728
Taiwan: 23,550,077
Australia: 24,688,400
Ecuador: 16,385,068
Hong Kong: 7,374,900
El salvador: 6,344,722
Singapore: 5,607,300
New Zealand: 4,826,660
Liberia: 4,503,000
Jamaica: 2,881,355
Namibia: 2,113,077
East Timor: 1,167,242
Belize: 387,879
Micronesia: 104,937
Marshal Islands: 53,066
Palau: 21,503
Caribbean Netherlands: 25,019
460,551,122 in total. Which is still less than 50% of China's 1.4B
It's easy to forget just how much bigger China is.
Ren is intended to be a general purpose data exchange format.
I'll just leave my comment above to not destroy the comment chain.