I've been playing with some clustering stuff in my free time for the past few months.
What I've found is that the problem seems to get a lot more reasonable if you know how many clusters there are.
K-Means requires this information, but afaict agglomerative techniques don't. I wonder why this tool's agglomerative clustering method requires the number of clusters as an argument.