TOML, Tom's Own Markup Language
github.com
github.com
The limitation on array types seemed fairly arbitrary at first glance, but after thinking it over I realized it aided compatibility with languages that do not support homogeneous arrays. Though as far as the types go, I would add boolean and perhaps non-quoted strings for single-word values.
Now that the technical criticism is out of the way, holy crap this guy is arrogant.
Don't you mean languages that only support homogeneous arrays (or languages that do not support non-homogeneous)?
As the spec says that the array elements must all be of the same type, thus homogeneous.
If I a mistaken, can you please explain why?
"data = [ ["gamma", "delta"], [1, 2] ] # just an update to make sure parsers support it"
so in a static language it would be like: Array<Array<???>> not sure this makes any sense
Tom's not being arrogant; he's just being irreverent.
Hash (and hash map, hash table etc) leak too much implementation detail. What if you want a tree-based mapping instead? I like how in C++ it's map (for ordered, rb-tree based maps) and unordered_map (for unordered, hash table based maps).
HashMap = HashTable = Map = Table = Dictionary = Hash
And if you're feeling adventurous, Object.send :remove_const, :Hash #define BEGIN {
#define END }
were such a bad idea, then why is it so easy? HashMap = HashTable = Map = Table = Dictionary = Hash
possibly not qualify as "obviously bad"? The only reason you've offered up is because it is easy...One addition though:
Cocktionary = HashMap
"Dictionary" never really made sense. my %hash = ();
Hash.new
So naturally people talk of Hashes etc. I understand where you're coming from, but it's really not very important, and it would be more confusing to talk of Hash Tables as learners would naturally look for HashTable in the stdlib.Though even HashMap isn't bad because typing is a solved problem - with auto completion and touch typing two words really aren't an issue in my mind.
People get used to living with all kinds of things, but that doesn't make them any better. Yes I'm aware that this applies equally to my typing comment as to you having got used to hash.
Why in golang is a function denoted by func instead of function?
I'd guess it's because programmers prefer fewer keystrokes as long as the term remains sufficiently mnemonic.
* http://www.nntp.perl.org/group/perl.perl6.language/2007/05/m...
* http://www.nntp.perl.org/group/perl.perl6.language/2007/06/m...
Because we need a decent human readable format
that maps to a hash and the YAML spec is like
600 pages long and gives me rage. No, JSON
doesn't count. You know why.
I do not know why, And would love if one can explain me?Other than comments, I see not difference between both.
Also, that human readable is not an accurate, as it should be hacker readable, you know, IT folks are the only target audience of those files.
[owner]
name = "Tom Preston-Werner"
organization = "GitHub"
bio = "GitHub Cofounder & CEO\nLikes tater tots and beer."
dob = 1979-05-27T07:32:00Z # First class dates? Why not?
{
"owner": {
"name": "Tom Preston-Werner",
"organization": "GitHub",
"bio": "GitHub Cofounder & CEO\nLikes tater tots and beer.",
"dob": "1979-05-27T07:32:00Z"
}
}{ 'because': { '80': 'percent' }, {'of': 'JSON', 'is': 'brackets' } }
[1] https://github.com/mojombo/toml/issues/2#issuecomment-140029...
{ "because": { "80": "percent" }, {"of": "JSON", "is": "brackets" } }
{ "because": [{ "80": "percent" }, {"of": "JSON", "is": "brackets" }] }
Data marshalling/transfer
Config formats
For the latter, as they are typically written by hand, it's not particularly appropriate as the syntax is noisy and multiple nesting with brackets tends to lead to errors, even if you understand it perfectly well in principle, and of course there are no comments, no datetimes etc.I imagine this is intended as a saner version of YAML for configs.
[because=[[80=percent][of=JSON;is=brackets]]]
The implementation, which is in Haxe and has an informal spec in comments, can be seen here: https://github.com/triplefox/triad/blob/master/dev/com/ludam...
I didn't view bracketing as the enemy(which seems to be the focus of a lot of config syntaxes) but rather the combination of multiple types of bracketing, plus start-and-stop usage of shift keying. I only have two types of brackets, the sequence [ type and the long string {" type, and you can "feel" when you're writing a long string because of that sudden need to use the shift.
Try this in TOML:
key = "value1", "value2"
The same mistakes can be ignorantly made in any markup.
{ "what i want": "what i really really want", "ಠ_ಠ":"Ignore the eyes" }
I like it, though. More grepable than JSON or YAML, with the way it handles nested keys using dot notation.
[1] - https://github.com/mojombo/toml/commit/aa4ac1d6df1031ebe871c...
Speaking of INI, for the longest time the killer app for INI files for me was persistent data storage in batch scripts (.bat/.cmd files in Windows 9x/NT). Using a command line utility like [2] or a similar program from IBM that sadly wasn't legally redistributable you were able to achieve persistence with minimum effort, which would otherwise be difficult to program in batch. I even wrote a portable clone of inifile.exe for MS-DOS and Linux to be able reuse my scripts more easily. TOML would sure benefit from the same.
{
// Sets the colors used within the text area
"color_scheme": "Packages/Color Scheme - Default/Monokai.tmTheme",
// Note that the font_face and font_size are overriden in the platform
// specific settings file, for example, "Preferences (Linux).sublime-settings".
// Because of this, setting them here will have no effect: you must set them
// in your User File Preferences.
"font_face": "",
"font_size": 12,
// Valid options are "no_bold", "no_italic", "no_antialias", "gray_antialias",
// "subpixel_antialias", "no_round" (OS X only) and "directwrite" (Windows only)
"font_options": [],
// Characters that are considered to separate words
"word_separators": "./\\()\"'-:,.;<>~!@#$%^&*|+=[]{}`~?",
// Set to false to prevent line numbers being drawn in the gutter
"line_numbers": true
}1- Read the contents of your JSON file.
2- Strip out the comments with some regex foo or such.
3 - Feed the remaining contents to your JSON parser.
Douglas Crockford himself suggests you "Go ahead and insert all the comments you like. Then pipe it through JSMin before handing it to your JSON parser." That sounds like a reasonable workaround.
https://plus.google.com/118095276221607585885/posts/RK8qyGVa...
[...]
> There are two ways to make keys.
I guess I haven't had enough whiskey yet.
And two ways to indent.
Tens (or possibly hundreds) of thousands of people use Jekyll now. It's interesting to note that Jekyll started out as a "brain fart" as well. Just one amongst hundreds of blog engines. I wrote it because I was dissatisfied with everything on the market, and I thought I could do something different and better, to serve my own needs. I open sourced it, because I thought others might get a kick out of it.
I'd wager that most of the great things we use today started as nearly ephemeral emanations from someone's mind, often late at night, or helped along by a snifter of brandy. The funny thing is, if you never try out your crazy ideas, you'll never know which ones might have changed the world.
https://github.com/defunkt/github-gem/pull/59
The pull has 19 people asking for integration and has some stellar comments:
"seriously? year long pull request with two lines of changes?"
"I normally would think that the github gem features for paying users would get a lot of attention from the folks at github..."
I even tried getting it pulled via pre/postsales emails to enterprise@github.com (I'm a enterprise customer) which was met with a "yeah, i'll tap him on the shoulder to integrate - year later nothing.
Scientific support for this:
http://www.psychologytoday.com/blog/choke/201204/alcohol-ben...
http://bigthink.com/ideafeed/how-alcohol-inspires-creativity
https://www.sciencedirect.com/science/article/pii/S105381001...
Note: there's nothing wrong with releasing brain farts; quite the contrary. I didn't at all mean to imply that you shouldn't do that.
<owner name="Tom Preston-Werner"
organization="GitHub"
bio="GitHub Cofounder & CEO\nLikes tater tots and beer."
dob="1979-05-27T07:32:00Z" />
<database server="192.168.1.1"
ports="8001 8001 8002"
connection_max="5000"
enabled="true" />
<servers>
<alpha ip="10.0.0.1"
dc="eqdc10" />
<beta ip="10.0.0.2"
dc="eqdc10" />
</servers>I will note this isn't a valid XML document: you have no root node.
. No native support for numbers, dates, booleans or lists. The latter can be implemented using subelements, but it's so cumbersome that you skimped on that and used a non-typed string instead (the database ports).
. Redundant verbosity. Root elements, closing tags, way too much crap to be manually inserted.
. XML parsers are huge, complex beasts which have no place in many smaller applications.
. Being XML, it leaves way too many possibilities for crappy developers. Namespaces in config files, oh joy!
How so? .Net can't magically discover the types of values or prevent developers from abusing the format.
you don't have to be zealous and put every small piece of data in a separate element.
But then you're layering a complex format with a custom application-specific parser, with an unknown syntax (e.g. spaces vs commas, are ranges supported, etc). It obviously can be done, but it's a mess.
# line 36
text.split("#").first
This will have trouble with a line like: tweet = "TOML is #awesomesauce"
-- # line 43
array = $1.split(",").map {|s| s.strip.gsub(/\"(.*)\"/, '\1')}
You should recurse into coerce here, or you'll just lose types. (Also you're assuming arrays of strings.) array = $1.split(",").map {|s| coerce(s) }
--You're also not dealing with nested key groups. (eg. [servers.alpha]).
--
That being said, naïve string parsing is a terrible way to build a new markup language implementation. It's the reason the Markdown landscape is such a mess[1]. What this really needed is a formal grammar.
[1]: I actually tried to fixed that by writing a formal lexer & informal parser for Markdown in a side-project of mine[2]. It's not quite there yet, because for practicality reasons I wrote my own parser instead of a formal AST-generating parser.
Made it into a proper project/gem here if you want to file issues: https://github.com/jm/toml
And good call on the nested key groups. Shouldn't be hard to knock that out.
Yes, being the CEO of Github does give you the power to do whatever you want.
Of course, drinking and coding is a great idea. The Ballmer peak isn't a joke, it's a way of life.
I was all set to try a translation when I hit this section:
<dependency>
<groupId>com.google.apis</groupId>
<artifactId>google-api-services-drive</artifactId>
<version>v2-rev53-1.13.2-beta</version>
</dependency>
<dependency>
<!-- A generated library for Google+ APIs. Visit here for more info:
http://code.google.com/p/google-api-java-client/wiki/APIs#Google+_API
-->
<groupId>com.google.apis</groupId>
<artifactId>google-api-services-plus</artifactId>
<version>v1-rev22-1.8.0-beta</version>
</dependency>
<dependency>
<groupId>com.google.api-client</groupId>
<artifactId>google-api-client</artifactId>
<version>1.13.2-beta</version>
</dependency>
<dependency>
<groupId>com.google.api-client</groupId>
<artifactId>google-api-client-servlet</artifactId>
<version>1.13.1-beta</version>
</dependency>
How would I represent this in TOML? [dependancy1]
groupId = "com.google.api-client"
artifactId = "google-api-client"
version = "1.13.2-beta"
[dependancy2]
groupId = "com.google.api-client"
artifactId = "google-api-client-servlet"
version = "1.13.1-beta"
That's not right, it clearly should be an array, but I don't think the standard supports it. At best I would think you'd have to use parallel arrays [dependencies]
groupIds = ["com.google.api-client", "com.google.api-client"]
artifactIds = ["google-api-client" , "google-api-client-servlet"]
versions = ["1.13.2-beta" , "1.13.1-beta"]
and that's just not pretty.[dependencies.com.google.api-client] artifactId = "google-api-client" versions = "1.13.2-beta"
[dependencies.com.google.api-client] artifactId = "google-api-client-servlet" versions = "1.13.1-beta"
Also it creates the key value maps:
dependencies.com
dependencies.com.google
which shouldn't exist, so that doesn't seem right either. npm search toml
npm http GET https://registry.npmjs.org/-/all/since?stale=update_after&startkey=1361700343737
npm http 200 https://registry.npmjs.org/-/all/since?stale=update_after&startkey=1361700343737
NAME DESCRIPTION AUTHOR DATE
node-toml TOML parser =ricardobeat 2013-02-24 10:08
toml TOML parser for Node.js =binarymuse 2013-02-24 04:19 toml parser
toml-node TOML ==== =thehydroimpulse 2013-02-24 08:01
toml-parser A TOML parser for node.js =aaronblohowiak 2013-02-24 06:41JSON has two drawbacks: a lack of comments (although you could add "#" keys in relevant places) and no binary support (arbitrary conventions include base64) but this doesn't support binary anyway.
Lack of comments makes JSON much better for data exchange than formats with comments.
It isn't a friendly form of human input. My error rate is 50%+ , you have to lint on save to catch things that are invisible to the naked eye
No ability to override, extend or reference keys. This is most useful in config objects where for eg. in a dev object you want to override the username and password for a database connection but not repeat all the other parameters
No comments
Am I missing something with that?
It's also nice to have a configuration file mean the same thing regardless of its runtime environment.
1. A way to have multi-line values for non array types
2. A more flexible number syntax (e.g. allow hex and binary integers, allow exponents on floats, allow NaN and +/-Inf)
3. Make it possible to have an extra comma after the last element on an array (as in Python)
4. Add a way to "include" another config file
#1 is important because some projects require all lines to have a max width of 80 lines, including on config files.
#2 is important for scientific/engineering projects. I think the current simple format shows that this format is a little too web centric. If this is going to be used for non-web stuff this is a must.
#3 is something that helps when putting this sort of configuration file in version control. Without this, adding an extra entry to a multi-line array creates a diff in two lines rather than 2 (since you must add a comma to the line above the one that you inserted). This is something I miss in JSON and which Python did just right (IMHO).
#4 would be useful in cases in which you want to provide a base configuration file for example.
Also, maybe I missed it but it is not super clear what would happen if you redefine an existing entry (I hope it is possible). Finally, is order important?
EDIT: typo.
It's not just in the diffs. Trailing commas make editing the list easier.
If we want simplicity, then why not make sure it is a subset of YAML?
I just need to auth and push it to npm.
(Meta: the edit link expired, hence the reply to myself)
https://github.com/ricardobeat/toml/blob/master/index.coffee
[ [1,2], ["a", "b"] ]You have an array of array, which at that level, satisfies the spec. The children individually keep types contained.
That said, I'm going to assume the intent is to not allow that.
For example, wouldn't that mean something like:
array = [ [1,2], ["a", "b"] ]
Be the same as this: [array]
0 = [1,2]
1 = ["a", "b"]If you parse the outer array as just "array of arrays" (as each element is an array), you're not "mixing". But if we're supposed to be parsing it as "arrays of arrays of _type_", then we are mixing.
I've always preferred INI over JSON for this reason.
On the other hand I'd like to mix my data types as much as I darn well please.
Edit: as mikegirouard points out it is much easier to read than (for example) serialized data, but still not as friendly as ini.
(If you don't know why this might matter, try opening your browser's Javascript console and evaluating 10000000000000001)
That's my peeve, though. I suspect that Tom is probably more concerned with readability. TOML also looks like it can be parsed a line at a time and doesn't really need to do any recursive parsing, so you could probably parse a stream of it as it arrives, which I imagine is trickier with JSON.
{"The invisible character":"really messes with javascript "}
Copy the text and paste it in console.
Using the end-of-line as a comment terminator would require significant refactoring of JSON parsers, which were previously at liberty to lump CR and LF together with SP and TAB. A starting and ending token, on the other hand, fits the pattern already required of a JSON parser.
In this decade, only a brain damaged text editor would do that.
This reminds me of a new project I'm working on called Leewh. It's based on Wheel and kinda has the same overall function, but I needed something to get my project rolling quickly and using .ini and JSON syntax separately felt... well... too square, I guess.
I figured I'll come up with something more well rounded.
https://github.com/biilmann/coffee-toml
Error handling is rough and it still doesn't handle groups with dots [alpha.beta], but appart from that it should be fairly complete.
class Tock
VERSION = "0.0.1"
# TODO: IMPLEMENT ALL THE THINGS
endParsing JSON for arbitrarily nested keys is nasty, and this makes it extremely natural.