Show HN: Property Trends Scraped from Zillow
trends.pillr.io
trends.pillr.io
> Median List Price: The median price at which homes across various geographies were listed.
> Median Sale Price: The median price at which homes across various geographies were sold.
> Sale-to-List Ratio (mean/median): Ratio of sale vs. final list price.
> Percent of Sales under/over List: Ratio of sales where Sale Price below/above the final list price; excludes homes sold for exactly the list price.
It uses the Zillow data set mentioned above. I would suggest giving it a shot as well, the graphs are rudimentary but they work well enough for me. :)
Caveats :
1.This works for the USA only.
2.The report should take up to a minute to display as the data is being scraped from Zillow.
3.If the site does go down, it is due to the volume of traffic from reddit. Then I'm going to need to scale up my servers. I will try and stay on top of this.
4.If the no results are obtained, Zillow probably didn't have "sales" data for that zip code, or the connection to Zillow failed. I would just retry once again.
I'm looking for feedback on these points:
1. How do I monatise without adverts? I would hate to display ads and kill the experience.
2. Do you want similar trends? Eg. I could also show "Average days in the market for properties", "Similar zip codes for your budget." etc
3. Would you be willing to register and voluntarily give up some data ? I would like to know where you are searching from, it will help me build a database of "Where are people from and where are they looking to move?"
4. Any thing that will help me improve is good.
Well, selling it to Zillow is probably a natural fit. Or some other realtor org, but you'd probably have to change the source of your data.
One way you could monetize would be to make deals with MLSes for integration into their services, either for realtors or for consumers. Can't tell how appropriate it would be for realtors, I assume it's more oriented towards consumers.
If there's some way to get to see Northern Colorado there, I can send it around to some POI in the local MLS and see what they think. Again, you might not be looking for MLS feedback.
Disclaimer: I run coloproperty.com and the MLS. :-)
How did you find yourself running those? What's your story if you don't mind me asking?
I've known many people who got into real estate (just getting a realtor's license) on the side as a way to make some extra money, but to me the appeal has always been having access to the "raw" MLS information / hearing about properties before anybody else.
Do you have any gut feeling for what percentage of properties never hit the market because of connections like these?
Alas, you'll likely hear from a lawyer before a PM, but good luck nonetheless.
I could also argue that I’m redistributing the computed data done by my web app like the average, median, and the graph. Then I’d have to hide the Zillow results source table.
https://sco.library.emory.edu/research-data-management/publi...
https://www.justice.gov/opa/pr/justice-department-files-anti...
Contacts and connections are easy. Money is harder :)
Monetizing data acquired through scraping could run afoul of copyrights, so the love letters would probably only arrive once OP attempts monetization.
The only argument Zillow can against me (imo) is that I'm being a burden on their servers at scale.
The LinkedIn lawsuit resolved to that conclusion, though I think LI is appealing again.
It basically says the site can't take legal action against scrapers of public (non-auth protected) info. Still, there's nothing that says they have to make it easy for you, to include deploying anti-scraping measures, rate limits, redesigns and moving data behind auth.
I think folks who are house hunting will find this useful when looking for comps (they will be getting information for multiple if not all houses in their target zip code in a table; they can easily digest the information in the table). So you could potentially charge a small fee (say $5 - $20) for house hunters to run this query and have the information available to them (say a link) for a month or so. During that period, they can click on the custom generated link and it will give them the most recent result.
Same thing for 'Average days in the market' but this will be targeted at those looking to sell.
The trick would be where to find these folks so they know about your service. Maybe advertise in housing forums or reddit (find the housing related forums)
edit: One other thing, it showed "Loading property 6 of undefined..." then back to "We did not get any results" which seems quirky.
As for the “loading property of undefined” , thanks for the report. A code cleanup is long overdue
I found this interesting, and was immediately curious for more: what the data was like going back more than a month, if there are any correlations to number of bedrooms, school districts, or the like, etc. etc. I could imagine you coming up with very interesting auto-generated 40-50 page reports.
I would not pay for this now, but when I was on the real estate market last year I might have. You could target individual buyers or sellers ($5-20 or so for a single report), real estate agents (some sort of subscription), or both -- don't know what the best strategy is.
They spend a ton of money on market analysis products because most of them don't have any sort of background that helps with it.
If it helps them get 1 more client or help a listing sell for just a couple % more it has paid for itself.
Button it up a bit, paywall it, run every zipcode as a job and cache the fully rendered html report the end user sees.
If you scroll down you’ll see they made an error listing the property at 389million instead of 3.89 million. I’ll have to use some upper /lower limits in the graph to bump out outliers like this . It’s hard to pick a number, some distress sale properties get sold for very low prices
I'd say go ahead and display ads, just do it without being obnoxious. Limit ads to small static images hosted on your own domain or text links, keep them unobtrusive, and don't use Google or any other evil surveillance capitalism company.
Because the site is driven by zip code you've got a great opportunity to provide ads relevant to the area and not the specific user. Anyone who finds reasonable ads so off-putting that it would drive them away from your service will be using ad blockers anyway.
Personally, I never register for websites if they require an email address unless they have a need to know my identity to accomplish what I want from them. Your site wouldn't meet that requirement so no registration for you. I'll sign up for some sites that need a registration but don't require a legit email address but they don't get my real info (fake names, addresses, etc)
Remember that any information you collect you then have to retain, backup, secure, etc. and it puts on the hook for reporting data breaches and for gathering and handing all of that information over to governments and other parties in response to court orders. Never collect or store any more data than you absolutely need to and you'll have so much less to worry about.
If you really really want to make money without ads, limit the number of accesses to something sane and charge the people who hammer your site day after day or week after week some fee for the trouble. They obviously find your service useful so maybe a small fee would be worth it to most of them. Maybe you can even offer to automatically push the kinds of data they keep requesting over and over again to them in some way.
I'm pretty sure the splash image is London, UK. Very confusing!
https://www.housedigest.com/874861/the-iconic-san-francisco-...
I've been involved with screen-scraping of req's for <major airplane company> to produce spreadsheets for others in the same company, and I'll take FTPing the data->CSV almost any day.
* Especially when the FAA is the eventual consumer *
We do that all the time.
Also, is zipcode the lowest level of granularity (presumably not)? Do any of these services have finer granularity?
Zillow can absolutely keep up with the load. Your proxies are being blocked by the WAFs that Zillow utilizes, usually PerimeterX. Also, based on your other comments, you have no idea how to architect a web application that can handle the trickle of HN traffic, and I'd wager any problems are more likely on your side.
Find a shared proxy provider that doesn’t limit bandwidth. It’s usually about $0.50/mo per IP. Not that any of this really matters, you’ll still get blocked by Zillow. It’s not an IP limit l anyway, it’s a TLS fingerprinting or JS-based fingerprinting issue that is getting you blocked.
https://www.webscrapingapi.com/top-residential-proxy-provide... obviously SEO'd page, but it has useful info.
Rightmove, the primary property app in the UK, gives you the previous price the properties were purchased for, although perhaps not analytics across an area per se.
I can’t wait for the business study of how much wealth the Property Brothers end up destroying - should be much more interesting then the story of the guy who stopped putting olives on the airline salads.
As I already knew, houses in this zip often selling $300k above list, some are pushing $800k. Yea inflation! cough
Data use agreements with those organizations require that anyone using it make efforts to prevent its being scraped.
Source: have signed a DUA with an MLS myself for academic research.
Are there plans to filter by more params?
That being said, you still need to be resource-polite. A lot of people scrape Zillow through browser automation toolkits like Selenium, Puppeteer etc. because it's a JS heavy website and these tools are really bandwidth intensive. This could, in theory, get you in trouble for DDOS.
Instead, since Zillow is using Next.js for their backend so, you can actually retrieve the dataset for any page just by parsing the nextjs cache. This can be done by selecting data in the <script id="__NEXT_DATA__"> node which requires minimal resources from both sides. e.g. in python:
import json
import httpx
from parsel import Selector
response = httpx.get("https://www.zillow.com/b/1625-e-13th-st-brooklyn-ny-5YGKWY/")
script_data = Selector(text=response.text).css('#__NEXT_DATA__').get()
script_data = json.loads(script_data)
# all of the property data is here, for example building details:
print(script_data['building'])
I wrote a tutorial on this if you'd like to learn more: https://scrapfly.io/blog/how-to-scrape-zillow/#scraping-prop...I wish you the best of luck getting the blessing of MLS, I had an idea for something similar and couldn't get anything useful.
Also I think Zillow is litigious as part of their agreement to have access to the data (and proprietary reasons obviously)
sometimes sellers over or under estimate their home. Over asking price doesn't indicate how a market is doing. It indicates other things: are sellers under estimating their home to start bidding wars? how delusional some sellers are on the price of their home? etc. These are interesting things to know, but as a buyer, I'm more interested in putting the right offer at the right time, and there's no better indicator than $/sqft for a specific location. Of course it varies (how many bedrooms? how many bathrooms? other cool amenities?) but this could be a subfilter of this product.
https://www.redfin.com/zipcode/<ZIP_CODE>/housing-market
For example: https://www.redfin.com/zipcode/95050/housing-market