Three Months to Scale NewsBlur
blog.newsblur.com
blog.newsblur.com
I've been asked quite a few times why I open-source the code. The answer is simple. Let me use an example from Joel Spolsky.
(From http://www.joelonsoftware.com/articles/fog0000000052.html)
The dominant spreadsheet, with 100% market share, is Lotus 123. You're the
product manager for Microsoft Excel. Ask yourself: what are the barriers to
switching? What keeps users from becoming Excel customers tomorrow?
... <snip: barriers to entry> ...
That's the barrier to entry. Not how hard it is to switch in: it's how hard it
might be to switch out.
And this reminded me of Excel's tipping point, which happened around the time of
Excel 4.0. And the biggest reason was that Excel 4.0 was the first version of
Excel that could write Lotus spreadsheets transparently.
Yep, you heard me. Write. Not read. It turns out that what was stopping people
from switching to Excel was that everybody else they worked with was still using
Lotus 123. They didn't want a product that would create spreadsheets that nobody
else could read: a classic Chicken and Egg problem. When you're the lone Excel
fan in a company where everyone else is using 123, even if you love Excel, you
can't switch until you can participate in the 123 ecology.
If you know that in the absolute worst case that you can still use the product even if it's shut down, then by golly, you have even less of a reason to not switch to it.Don't get me wrong, it is wonderful and very commendable of you to open source NewsBlur! Nor do I feel entitled to the source code or free ride (I will probably end up paying for the hosted solution). Just providing some feedback, given that you indicated that low barriers to entry are important to you.
I am not a web developer so that can partially explain my struggle with fabric, Django and multitude of webdev and devops modules that you use.
I wish you good luck and I hope that you take this wonderful product even further.
The real problem with setting up your own instance of NewsBlur is that you'll have to do your own feed fetching. This is effectively what you're paying for when you pay for NewsBlur. Let me break it down.
500 feeds updated once every five minutes is 144,000 feed fetches a day. Couple that with the original page and icon fetches, you're looking at almost two feeds every second just to keep your feeds up to date. And the average feed takes 5 seconds to completely update (feed, page, differences in stories, updating unread counts). So you have to run 10 processes in parallel, and that's already beyond the capabilities of most single machines. Then good luck keeping your DBs running clean and backed up.
Or you could pay $2 / month and never have to worry about setting up a very big stack. And you get the social community of shared stories on newsblur.com.
On the other hand, you may end up with people running the app that have no way to maintain/support it.
It looks like MongoDB is being used for fetch/push history, feed icon data, and article data. It's also being used for some other stuff: https://github.com/search?l=Python&q=%22mongo.Document%2...
I suppose I could also look at the source ;)
For the cost of hosting/development/work you are MUCH better paying for newsblur and forgoing 6 outside coffee's a year.
Er... really? My laptop is three years old and I run 150 threads on it every time I launch the server I'm working on (dozens of times a day). No problems at all.
It's in Java, in case it matters.
Also there seem to be a lot of stuff that need to be installed and configured that might not be necessary for single-user/<10 user instances. Yeah the 'fetching the feed part' still needs to be done, but I guess right now it is close to impossible for anyone apart from yourself to be able to understand how to do that.
Due to the mass exodus from Google Reader, I understand that NewsBlur has faced challenges in scaling and hence is slow. But things can only get better from here.
Since it in open source software, you should at least setup a getting started developer's guide so that you will get a lot more contributions from the community.
If you have at least fifty previous entries from a feed it's easy to predict fairly accurately when that feed will need to be updated again.
That's how the indexer for http://rssident.com works. Saves you tons of cpu cycles.
GitHub is not an app store. You can't expect to "install" a software as a service's code base. Open sourcing a project is opening it up to engineers, providing a self-hosted option on the other hand should be an easy to install alternative that probably requires another full-time engineer to handle as it needs to come with ability to upgrade, get logs, support and documents...
Maybe container images that can also be used as bootable images?
I think Chef and Puppet try to fill that role, somewhat
Edit: actually, this recent HN discussion looks promising https://news.ycombinator.com/item?id=5392041
NewsBlur should distribute in an .egg form, or via PIP directly, with VirtualEnv and everything.
In our area of business, our features get copied really quickly.
For other folks hoping to build their own service, commercially or otherwise, having access to your code provides a huge savings in time to market.
I'm sure you will need all the help you can get with upcoming traffic so I've tried contacting you recently about helping you scale better with https://errormator.com, but I somehow miss you when you are online on IRC, so I've sent you mail to the mailbox i've found.
Its great since your app is python based and we maintain an awesome python client.
Drop me a line if you are interested, we love open-source at Errormator so you would get a completly free plan for non-profits :-)
I could also help you personally if you have any integration questions, you can find me on #errormator on Freenode IRC server.
Please see this inteview with us that was posted recently, the article is in polish but the screenshots should give you a better overview of what errormator can do:
http://translate.google.com/translate?sl=pl&tl=en&js...
In short:
- exception monitoring
- performance metrics
- slow request and slow call aggregation
- log aggregation
We def. need a new index page and a features page to show all the functionality in a nice way.
Cheers!
This is the kind of single founder "life style" startup that VC's would paw-paw that is Doing It Right that many of us aspire to. The other ones I can think of are Pinboard and Instapaper. I'd rather go this route, personally, here's to more like these. Cheers.
Pinboard and now Newsblur.
I wonder what would be next? Most of Yahoo's properties seem be be neglected at the moment, but what about Google? Google News perhaps?
This information is very valuable and I'm sure a good service could charge a nice premium.
That's at least $200,000/yr with high margins.
The amount of annual income from this product will now end with "million".
still should easily net him a nice recurring six figure salary.
He reaches $1 million at just over 40k paid users.
Factor in the userbase he's already been building for 4 years, the hundreds of thousands of Google Readers currently leaving, the millions that will need to find a replacement before July, and the last-minute movers on July 1st.
I don't think server costs and a couple of employees will put him below a 7 figure income.
You're clearly a talented programmer, but perhaps employee #2 should be a sysadmin.
Congratulations.
This is why I love SaaS. 6 figures revenue in a couple of days, and I can only imagine how much more will be made in the next 3 months... and beyond.
But it's not a once-off payday. This is recurring. 7 figure annual income is yours, Sam. People slave away for decades trying to reach a tiny fraction of this amount. Put it to good use, brother.
Those couple of days took 4 years of preparation.
1) Why are you open about the number of users ? (personally I think its great) 2) How do the real-time stats work, I notice that the amount of regular users was over 20k a few days ago, and now its 8k. Are these the users who have used the client/service in the past 24 hours ? (and not the actual total amount of users)
That's only temporary though.
(So existing Google Reader apps can quickly switch to NewsBlur)
And then the black swan event comes.
And then fire everywhere.
> I did not expect it to come this soon.
Maybe Google will buy them! :)
Edit: I kid, but that's actually a possibility in my mind and would be an ironic outcome. Maybe Facebook or Yahoo or somesuch will see the value of having the New Reader. Selfishly, I am hoping for something else.
If I were Palantir, I'd probably drop $250k/yr to run the World's Best RSS Reader for the kind of research analysts who become Palantir customers, and the kind of geeks who make awesome Palantir employees. There are maybe a hundred companies you could substitute for Palantir here.
Seriously ironic that both of the likely Reader successors are YC funded (Feedly and Newsblur)
For a task a parallelizable as fetching feeds, I can't recommend Picloud.com enough. You can ssh into a box, provision it however you want, save that image off, and then run arbitrary commands/scripts on instances of that image, on the CPU type you want, and pay only for the seconds of usage you have. (Their "s1" core type is $0.04/hour: http://www.picloud.com/pricing/ ) They also have a system that lets you mount the same shared "drive" from multiple instances at the same time, so if you're doing file-based stuff, it's easy.
The thing I like about it is that running another instance of your script on a new machine (or 2k of them) is trivial. No need to wait to provision a new VPS. Starting and stopping jobs is fast, so scaling up/down is fast.
(I'm swear I'm not affiliated with them. I'm using it for a side project at the moment and have been so excited about it, I want to spread the word.)
The day Google announced they retired Reader was "missed opportunity day", over here.
Anyway, good luck to you.
* Feedly has a simple interface which is much more like Google Reader. But it depends on a plugin (or an app) rather than being an allwhere accessible website. Their Kindle Fire app seemed to load content extremely slowly, and its interface wasn't a simple list. And it's free, so I can't be a customer and I don't know how they earn money.
* NewsBlur is complicated and unresponsive. There are a lot of inscrutable unlabeled icons, and I don't understand why an arrow follows the mouse. It's a paid service - I can be a customer - but I don't know how sustainable NewsBlur will be with its current issues and with subscription as its only source of funding.
If NewsBlur seems solid by May, I'll gladly pay them. If Feedly seems less like a faceless data-miner by May, I'll use them.
Really? It is a feature for better reading, so you can can easily see which line you are.
I am working on making the UI more user friendly and adding OPML import/export.
It's going to be supported by advertising with a subscription service for premium features and no ads.
Still working on the UI to show the article content and links to attached files.
Overall I found Feedly to be better.
Feedly turned me off because it has no open web implementation, it runs through the Chrome web store. Aside from some developer related plugins I refuse to use a completely superfluous wrapper around a website if I am on my PC.
I've tried Newsblur but it's just so "busy" (I know where my mouse is on the vertical axis thank you very much). I may warm to it, but I'll probably end up with something a little simpler. But I wish all the success in the world to conesus, a variety of small, customer focused, customer capitalized independent web service vendors is a good thing for all of us.
Seriously, going to the feedly website, the first thing you're confronted with is a list of platform-specific apps you have to download, and "this website wants to install something, ok?"
There seems little reason for this with today's browsers, and it seems like a very good way to make a large proportion of potential users say "no thanks" before even trying it. I know it made me immediately toss feedly into the discard pile in my reader-replacement search...
Also while I've got you, two issues I've noticed.
1. When I click "All Stories" to refresh my feeds it defaults to listing everything, even though I have it set so I only view one item at a time. As soon as I click the button/arrow to go to the next unread item it shows the first item as it should and everything works fine after that.
2. I seem to run out of items after I've viewed ~20ish and I have to click "All Stories" again to fetch a new bunch
I'm using the dev site if that matters (Issues happen on both sites though). So far these are the only things annoying me and otherwise newsblur is pretty good.
2. That's an unfortunate bug, but one that will probably get sussed out when I work on scaling over the next month.
Dev is largely just a re-skin. They share a backend.
I can understand not doing that by default since people may want to see their unread feeds, but I always just go to the first unread one anyway.
The UI needs work but it is very simple.
Scalability will not be an issue. It was designed from the start to be scalable. Just add servers.
For me the interface of NewsBlur is so much better for getting through tons of feeds,
Likewise, what are some of the engineering things you'd really like some of the more dev heavy users to maybe help via pull requests with? :)
(I see that the github issues list is relatively short)
I have a couple of comments on how the UI works, but I'll consider those as time goes by and send you more specific feedback then (mostly I'm annoyed by things that work differently to google reader, so I'll hang back and see if they convince me...)
I would like to try Newsblur but it seems that to get immediate value I have to have been some reader user: I can migrate my googlereader account or upload an OPML file.
What I would like to see are different versions of OOTB "skins". I can sign up as an "HN-lover" and it seeds my account with 50 blogs I should follow. Then I can actually evaluate the product.
Also note that free accounts are not actually disabled. If you leave the session when you hit the paywall and login from the homepage, you now have a free account.
This is my new favorite typo.
Thanks!
Thank you urban dictionary...
And congrats on your increased traffic.
While I do provision new servers regularly, I don't auto-scale, so I don't have a need for pure automation. It would be nice, but that's something I look into when I learn about new deployment procedures.