Links between Paul Graham's essays - revisited
solipsys.co.uk
solipsys.co.uk
I guess what he meant was tools which can do this for a particular site.
I can't think of any but as a previous poster said - a crawl of desired depth for all local pages, then feed the access log to "statviz" to generate a neat dot file along with the links.
Write a crawler that examines all local links, and builds a graph (it needn't be anything fancy; a collection of nodes and edges is enough, but you'll have to detect cycles when building the graph), and then print it out in graphviz .dot language.
Then you can feed it into neato [either through a command-line pipe, or construct one using Python's subprocess module], et voilà, instant pagemap! :)
You can also use some of the rendering modes and scrape position data for the rendered nodes, if you want to generate a clickable imagemap.
Output from the above is at: http://jey.kottalam.net/tmp/obgraph.out
I tried rendering this with "dot -Tcmapx -oob.map -Tgif -oob.gif" but dot segfaulted after 70 minutes. The code as posted at gist only outputs nodes for articles written by Eliezer in an attempt to make the dot file a manageable size. Tips/fixes appreciated.
1) Rendering an undirected graph via graphviz's neato tool.
2) Rendering to svg, which might let graphviz not have to keep as much rendering information in memory at any one time.
3) Asking about it on the grapviz-interest mailing list. https://mailman.research.att.com/mailman/listinfo/graphviz-i...
There's got to be standard graphing tools for things that are highly connected...
May be it would become restructured content if there are too many outward references like the... "Undergraduate" post. But even that when you look into this post the hyperlinks are for keywords like "essays", "hacker", "computer science", what you "love", hack a "blub" on windows ..
My take: - The most referenced (high in-links) posts are more dear to PG or most popular/strong in his line of thought.
- The ones with high outgoing links at times might mean they are restructured thought as 'markessien' noted.
- For a personal blog like PG's does higher interlinking reduce readability or make the content richer ? I guess the reader chooses this.
Larger numbers of outgoing links probably don't imply recycling. If anything it's a sign I think an essay is good and will get lots of readers, so it's worth trying to spread that traffic around. (These predictions are often wrong, though. I find it practically impossible to guess which essays will get lots of traffic.)
Perhaps you don't have time, but would you consider adding links that other people suggest?
It would be interesting to do to this what I do with my SiteMap, and colour nodes darker if they have more traffic.
That would be valuable ... you've made me think ... I may have a way of doing something close enough.
Hmm.
http://en.wikipedia.org/wiki/Latent_semantic_analysis
I found examples in ruby and python in this blog, which is unfortunately down just at the moment http://blog.josephwilk.net/ruby/latent-semantic-analysis-in-...
The paper is here [PDF]: http://www.arbylon.net/publications/text-est.pdf
C++ implementation: http://gibbslda.sourceforge.net/
(note, I haven't used the C++ implementation).
Thank you.
http://www.solipsys.co.uk/new/PaulGrahamEssaysRanking.html?Y...
I've just pulled the description from WikiPedia, implemented pretty much that, and that's what I did.
In essence, each page gets set to 1.0. Then distribute 80% of that down each outgoing link, and 20% to everyone equally. Lather, rinse, repeat.
I think that's pretty much the algorithm. Pages pointed to by pages with lots of juice get lots of juice. Pages pointed to by no one, or only pages with little juice, get little juice. It's linear algebra thought of as network flow.
Well, that's what I did.
1760 Stuff
770 What You'll Wish You'd Known
403 How to Start a Startup
365 Why to Start a Startup in a Bad Economy
349 Why Nerds are Unpopular
294 How to Do what You Love
274 Web 2.0
247 Great Hackers
207 How to Make Wealth
187 Lies We Tell KidsGoogle has a good PageRank implementation, though it isn't open source.
Yes, but doing what you've done here gives the entire site, whereas implementing it from the definition allows you to compute the actual values, as I did.
Of course, Google have twoke their algorithm so it's no longer exactly as they originally published, and they're not telling anyone what it actually does. Using the "clean" version is at least transparent.
I'm still interested in additional links based on semantics.
One small thing, it looks like you included the RSS link from the bottom of the page in the list of essays.