Markov chain algorithm in Go
golang.org
golang.org
One active area of research is smoothing the models, so they capture regularities without just memorizing the input text. The vanilla approach chooses a fixed order of Markov chain (e.g. 2-grams or 3-grams), where you choose n either ad-hoc, or (ideally) to balance underfitting (miss context you could use) against overfitting (end up memorizing verbatim text strings due to insufficient data). But the right "n" to choose can vary based on context as well, since different words show up more often in the training data, and are more or less variable in what kinds of words follow them. So there are a number of approaches to choose n adaptively and/or to smooth the counts ("Katz's backoff model", "Kneser-Ney smoothing", etc.).
From a more engineering point of view, there's also considerable work on pruning the models and representing them efficiently, if you want to work with things of the size of, say, the Google Books ngram dataset.
A few free-software packages that implement a good portion of the recent work:
http://cm.bell-labs.com/cm/ms/what/shannonday/shannon1948.pd...
See also the Usenet trolling program that indirectly inspired the Go example: http://en.wikipedia.org/wiki/Mark_V_Shaney
Plus the codewalk itself is cool. I'm unsure of the UX, but certainly enough for a discussion.
There's only insight to be gained if the Go version is particularly interesting due to features of Go.
I'd like to see more of this. A fresh interesting take on representing programs, especially welcome as not being served as yet another eclipse plugin or a multi-screen .vimrc.
https://github.com/gburtini/Learning-Library-for-PHP/blob/ma...
Edit: the thread above me is talking about bigram implementations, mine supports n-gram by setting the variable $degree when you train to anything you like. By default, its 1-gram.
Nothing is worse than seeing a "r" in the body of some function or as a struct member and not knowing if it is a reader, response, request so rather than using "r", I'll go with rdr, resp, req or this (were applicable). That is hardly verbose and our code is immensely more readable than something "idiomatic" like the Go standard library IMHO. I think this really should not be considered an issue of idiomaticness, if is was gofmt or go vet should care.
The short is if you like Go and not single letter var names, great use Go with longer variables names! I realize this is a religious issue so I can already feel the flames burning :)
This is how the entire Go stdlib is written, and how most packages are written.
type MyStruct struct {
wobbly int
}
func (m *MyStruct) IsReallyWobbly() bool {
return m.wobbly > 42
}
So far, I've found this convention to be straightforward to work with, and there's little reason I can see to make the receiver identifier more elaborate. Especially in light of popular languages that omit a "self" identifier entirely.