Visualizing San Francisco Home Price Ranges In D3.js
trends.truliablog.com
trends.truliablog.com
One method to define neighborhood extents could be cultural instead of geographic by crowd-sourcing the delineation. If one can ignore the desire to draw a strict boundary at a major street, you'll find certain 'neighborhoods' spilling out of their original blocks. Perhaps adjacent areas with distinct histories start to define themselves by the same architectural historical context. Or a once strong border between ethnic communities might disappear. Gentrification might creep across a river closing the gap.
These could all be fuzzy-mapped by residents.
Myself, I might approach the problem the other way - trying to infer the existence of a neighborhood-like structure in terms of similar homes which are close to each other (apply clustering algorithms), and then measuring to what extent these structures overlap. This technique would hopefully place the 4th-and-King-at-Caltrain high-rise apartments and condos apart from the 19th-and-Arkansas single-family and converted-single-family homes.
IDK. But it is interesting none the less.
The second method is interesting and seems completely new.
How is it possible to get people to give you what neighborhood they are in? I bet we would find that good neighborhoods go much farther in to bad neighborhoods than they do in real life because of people who live their pronouncing that they live in a good neighborhood.
Either way, I'm glad I asked!
I understand wanting to fit it in the screen space. Is Log not generally reserved for large differences in magnitude and for power laws?
(I don't have an intuitive understanding of logs, so this is a genuine question).
Scaling linearly allows big values to skew the results. Imagine a neighborhood where the high end is $20m and the low end $10m, giving us an absolute difference of $10m. Another neighborhood has a $1m high end and a $500K low end, for an absolute difference of $500K. Scaled linearly, the $10m range in the first neighborhood would appear to be much much bigger than the $500K range of the second.
But if we use a log scale, instead of asking what the absolute difference is, we're asking relatively how much more expensive is the high than the low end. Using our two example neighborhoods, both would result in a 2x difference, and thus both would have the same range of prices.
It's easy to point out where all the most expensive homes are, and scaling linearly does just that. But looking at the relative differences in prices provides a much more useful way of comparing different neighborhoods (or even cities if we look nationwide), because it accounts for the natural variances in prices in different areas.
I don't suppose you have an image of the linear version?
I prefer the linear story (not for the range-index though): it shows how consistently extreme the high end is from the median relative to the distance the low end is from the median (from a dense concentration of homes just above the median?).
1.) To convert numbers, they use +number, where number is a string. That is some excellent short hand and relies on automatic type conversion in javascript. Not sure what the speed is like, but I imagine it doesn't slow down things much,
2.) They forgot that they had jQuery loaded in already. You can see this by looking at the lines when they define w and h. They use d3 to get the height and width in a awkward way. $("vis").width() would have done the same for less thought.
3.) They are using some tools to check for or create automatic clean javascript. Not a semi-colon or tab out of line. Probably jslint, because they use a forEach at some point and I don't believe coffeescript uses forEach's in it's output. Coffeescript is probably worth the time to learn then; All those function(d){return d.a} become just (d)-> d.a.
Depending on your style, this could be useful as well: getter = function(attr){ return function(d){ return d[attr]; } }
ds.attr("x",get("x"))
4.) I am embarrassed I didn't know d3.svg.axis existed. That saves time and mistakes.
All in all, well done. Learned a good amount from reading through that. Thanks!
And btw, no code cleanup tools in use, just my OCD. ;) I haven't used Coffeescript yet. Will have to check it out.
... but, not really sure what meaning i am supposed to take away from this.
"Zip codes contain heterogeneous housing units that have a spread of prices" ??
The 33% still leaves room for the selected zip code to leap out.
I wonder whether this framework can make traditional display of such a data (i.e. box plots with outliers) as pretty.