Technical Documentation Should Be a Graph
neo4j.com
neo4j.com
Is this the new "Solution: Use Regular Expressions"?
There is a good reason tutorial style documentation is linear. Humans can keep track of what page of a book they are on, not which nodes of a graph they have visited.
Otherwise this doesn't add any new insight that wasn't there when the web was created. It's basically talking about hyperlinked documents. Identifying it as a graph doesn't help understand it any more than identifying the rotational groups of NaCl molecules lets me understand why they taste salty. It's useless structure/abstraction hunting.
I could see something like this being useful for understanding a complex and complicated document like a programming language specification. The trails would show which sections and subsections work together to produce a particular behavior users should understand.
I've also seen dependency graphs given in the preface to certain math textbooks, which is nice when you know what it is you want to know (and I wish more textbooks would be available in hypertext form).
With that said, I'm so, so guilty. :)
I've read papers where the space of possible computations for a given (typically, toy, even loop-free) nondeterministic concurrent program is represented as a directed topological space (see “directed algebraic topology” for more information), but these spaces are more general than directed graphs.
I don't think you would need to deal with infinite graphs, as "arbitrarily large" should be able to cover all (finite) possible computations and/or representations.
My guess is that every graph may correspond to an infinite number (though, not necessarily every) of possible computations and/or representations, depending on which "correspondence function" is being used to analyze the graph (though there may be additional "correspondences" between those functions and/or their results).
The precise "information" associated to them could be "anything", but would be proscribed by the computation and/or representation being disccussed and the graph in question.
I'm even more annoyed by your use of the word “isomorphic”, which, FYI, doesn't mean “vaguely similar in some way I can't articulate”. It only makes sense two speak of two mathematical objects being “isomorphic” when they belong in the same category. What category do you have in mind, that includes both computations and graphs as objects, and what are the morphisms between them?
E.g.
Computation / representation:
2 + 2
Graph [+]
/ \
[2] [2]
There is a way to produce the graph from the computation / representation and vice versa. There is a way to do that for every possible computation / representation and every possible graph (though, not necessarily with every pair).This could be described as a compiler or compiler-like. There are many interesting things to do with them, but in this example all that is needed to consider is the construction of an Abstract Syntax Tree from a string and an in-order tree walker that emits the contents of the nodes to an output string.
(1) A syntax tree isn't the same thing as a computation. Of course, you can produce computations by interpreting syntax trees, but: (a) It's perfectly possible for two distinct syntax trees to produce the same computation. [Say, by renaming all variables.] (b) It's also perfectly possible that interpreting the same syntax tree twice will produce completely different computations! [Say, if your language isn't pure.] Unsurprisingly, a syntax tree is a representation of syntax, not computation.
(2) Yes, I'm disappointed, because I expected your observation that “every computation (...) is isomorphic[sic] to some graph” to provide more insight than it turned out to.
Edit: Turned paraphrasing into literal quote.
0) You asked, "What information is associated to [nodes and edges]?". In this example, one piece of information is the direction of an edge.
1) I don't think I claimed they were the same. In fact, I am claiming that one can be represented by the other and vice versa. a) yes, that is part of the reason I conject that each graph may be associated with an "infinite number" of computations and/or representations b) also a reason -- there are perhaps an infinite number of languages, interpreters, compilers, etc. (some of which may produce the same result) for a given graph.
2) It was only an observation that the comment I responded to claimed only a subset of the truth. Anything done on a computer can be considered as a graph.
https://en.wikipedia.org/wiki/Graph_reduction https://en.wikipedia.org/wiki/Abstract_semantic_graph
Generally they're acyclic but some types of ASGs can represent recursive functions as cycles, so they are distinct from trees.
Basically, everything is in some way equivalent to some graph. Since anything is an example, it's difficult to say much without using an example, which then necessarily constrains the discussion.
It would be more useful if you presented an example computation or representation, then I could show you how to make and reverse an equivalent graph. (Which then might show you what information might be associated with nodes and edges.)
To put it another way, at best a link from a wiki page gets implied importance based on it's position within a document. A graph database edge can have many properties including "weight" or "relevance". Of course these could be added to wiki markup as properties, but there is a point at which the right tool is better than layered make-do's.
<a href="foo.html">Read more</a>
<a href="bar.html">Read more</a>
but: <a href="foo.html">Example code</a>
<a href="bar.html">Grammar</a>
If desired, one can even make this easier to parse programmatically: <a href="foo.html" class="ExampleCode">Hello world example</a>
<a href="bar.html" class="GrammarNotes">Grammar</a>
That allows for styling different types of links differently. Also, a wiki's backend software could extract separate indexes for sample code, grammar fragments, links to compiler source, etc. you could also easily include an attribute for, e.g. Click-through rate, and, if desired, style links accordingly.In fact, the difference between a classical wiki and a graph database is, IMO, an implementation choice. If you want to process complex queries rapidly or a lot of your attributes aren't textual or aren't intended primarily for visual display, a graph database is more appropriate. If you want to serve standard web pages rapidly, storing the content as HTML may be more appropriate, even if that means that processing queries such as "give me all links to source code examples" run way more slowly.
----
Programmers tend to carry over the structure of the program as the structure for its documentation. But this structure is not necessarily good for explaining how to use the program; it may be irrelevant and confusing for a user.
Instead, the right way to structure documentation is according to the concepts and questions that a user will have in mind when reading it. This principle applies at every level, from the lowest (ordering sentences in a paragraph) to the highest (ordering of chapter topics within the manual). Sometimes this structure of ideas matches the structure of the implementation of the software being documented--but often they are different. An important part of learning to write good documentation is to learn to notice when you have unthinkingly structured the documentation like the implementation, stop yourself, and look for better alternatives.
[…]
In general, a GNU manual should serve both as tutorial and reference. It should be set up for convenient access to each topic through Info, and for reading straight through (appendixes aside). A GNU manual should give a good introduction to a beginner reading through from the start, and should also provide all the details that hackers want. […]
That is not as hard as it first sounds. Arrange each chapter as a logical breakdown of its topic, but order the sections, and write their text, so that reading the chapter straight through makes sense. Do likewise when structuring the book into chapters, and when structuring a section into paragraphs. The watchword is, at each point, address the most fundamental and important issue raised by the preceding text.
https://www.gnu.org/prep/standards/standards.html#GNU-Manual...
I agree that docs should answer the user's typical questions but gnu docs seldom do. Can you imagine trying to figure out gnu tar from the man page? You never would.
Doesn't help that search engines starting with G view queries like 'tail -f' as stopwords.
These days, when looking at documentation, my go-to example for awful documentation is salt. Because, well, this is utter bullshit:
$ man salt 2>/dev/null | wc -l
133324
I find manpages to be much more convienent, even though they are less structured.
Second, is what the writer describing something that would be met with a hyperlinked table of contents, like http://www.postgresql.org/docs/9.5/interactive/index.html, along with maybe each page hyperlinking to other pages where fit?
The graph is useful for note-taking and exploration, certainly, but it produces a design constraint that isn't always situationally appropriate. Binding things into a narrative can add a lot of value.
I usually find "official" technical docs very hard to read - much harder to follow and less useful than a typical how-to book.
Writing is not code. Human elements like humour and story-telling make technical content much easier to understand and learn.
And "topic centric" should be something that comes out of testing. User testing should note the most common user taskflows empirically, and write those up in the docs.
This is kind of nit-picky, but I get bothered when someone uses the word "just" right before suggesting something that's really difficult to do.
If you think about how much energy your brain takes to unpack paragraph text into information, tree-structure should be clearer because visual boundaries correspond to topic / semantic boundaries, i.e. the brain can offload some parsing to the highly-evolved retina.
I don't have a good explanation for the preference for paragraph format in docs.
The volume of questions diminishes as the doc improves, they become more interesting, and you can get the simple ones out of the way by sending a link without remorse because you know it does contain exactly what the user was looking for.
Of course it takes a little while for this to really engage, you'll need to kick a few people in the bottom to get past the initial investment, but in my experience it definitely works.
Good docs are in everyone's interest. It's hard enough for insiders to maintain code that isn't documented properly. Worse, it's nearly impossible for outsiders to come in and make changes later - at least not without a lot of wasted time.
I'm embarrassed of the state of this thing, but I think I'd be remiss if I didn't link to https://willshake.net/about
It's a graph of a system based (almost) entirely on documents. The documents create the documents. The documents create the graph. And of course, the documents contain the code.
I can't make enough disclaimers about how rough the documents are, and even the graph is starting to collapse under its own weight. But I think it gets the idea across. If I've learned one thing, it's that this style of programming---where you're forced to document as you go---is extremely work-intensive.
I don't mean a mediocre programmer who can write, I mean an excellent writer who excels at interviewing programmers, who can comprehend source code, and most of all is trained to be a writer?
For example if you have a technical article which references some concept there will likely be a link to another article which explains that concept.
Thus each page already acts as a node connected to other nodes by the links within it. No need for a more formalized graph.