How to build a news app that never goes down and costs you practically nothing
blog.apps.npr.org
blog.apps.npr.org
I agree with his preaching the power of flat-files. Not that flat-files should be used to do things that they inherently can't...but that too many projects (or hobby apps) don't consider them and then spend as much time figuring out how to keep their server from crashing. I find it pretty amazing that they have only one small EC2 instance for their news apps (this is separate from the NPR.org site overall) just to do cron jobs.
Flat files, of course, require good planning...not least of which involves an accurate gauging of how often an app's data needs to be refreshed. But I like that kind of planning and thinking more than I do the kind it takes to maintain a stable server.
I think the database-driven content comes primarily from kitchen sink packages like Wordpress. More often than not, it's better to get an overall look at your structure and decide what needs to be dynamic and what can be "static".
When talking about it internally, we refer to "flat-file" as "generated". Simply meaning that it's dynamically created by a task or user interaction.
Both databases and file systems use B-trees to implement fast read / writes, it's just that SQL databases enforce structure. OTOH, using flat files moves data checking into application space and losing certain ACID properties.
Similar trade offs are made between NoSQL and SQL, dynamic and static languages, but I digress...
1. grepping for conditions and extracting the query with sed
2. appending a char to a flat file (i.e. increasing query count by 1)
3. sort by largest files and get the top 10
Parsing log files[0]:
for ./*.log -type f -exec echo 1 >> $(grep cond_A {} | grep cond_B | sed -E "s:.*(query).*:./results/\1.txt:") \;
Finding top 10 results: ls ./results/*.txt -Sl | head
[0]: Untested code, and it only grabs one line per log file instead of grepping all matching lines in the log file. I'd have to move the grep to the outside and loop the `echo 1 >>` command, but you get the point. cat *.log | grep condA | grep condB | sed 'some regex to get rid of dates, etc' | sort | uniq -c | sort -nr | head -n 5
Or something like that... Not very efficient, but it would work in a pinch. I actually do something like this all the time for large datasets. In the time it would take me to write something better, this set of commands is already done.Also, I would assume that javascript becomes a requirement for the functioning of the site; I always considered ecommerce to be the one part of the web were you wanted to be able to operate without javascript, so as not to turn anyone away.
I found this blog post that says that in 2010, 2% in the US had JS disabled:
http://developer.yahoo.com/blogs/ydn/posts/2010/10/how-many-...
The closest analogy I can give is Varnish's Edge Side Includes. It's not exactly the same, but it's very similar.
I found the unexplained use of the word "Boyerism" in this article to be confusing and unprofessional. I have nothing against the use of slang, or the expectation that readers will need to Google terms: this is the reality of our culture in the age of the Internet and excellent search. However, the casual and unexplained use of a private group's inside joke reveals a lack of awareness of how search affects interacts with culture. Is it a deliberate attempt to confuse and snub the larger audience, or is in unintentional?
I built it as a portfolio piece and haven't finished it because I've been doing consulting jobs, but if you want to watch a particular movie online without going the pirate route, it's the start of a legal alternative to the old sites like sidereel.
The tech behind it is the same as this article. Flask and jinja render a static html page for each of about 92,000 movies.
I use a flask app for the search functionality that accesses an elastic search database.
I used mongodb because it was incredibly easy to create a local cache of the JSON data I was getting from the apis I was accessing.
It all took a lot longer than I ever would have though to even get it to this point. There were a lot of little annoying issues with the various APIs I had to access, and the annoyance of parsing XML amongst other schelps I had to deal with.
I have only ever mentioned it on hacker news one other time and the last time my elastic search server crashed from the traffic. It is all running on a $5 a month digital ocean vps.
The flask/jinja static page creation is rock solid and would never fail if I pushed it to s3, right now my elastic search server is the bottle neck. I haven't taken the time to throw hardware at it or set up clustering.
All in all it's a pretty cool service in my opinion, I built it for myself because I love movies and spend a lot of time watching them online and made a decision to never pirate a content creators work again. Also the experience of netflix, amazon and itunes is orders of magnitude better than the old megavideo/bittorrent trouble of finding the real deal and not being inflicted with spammy ads with voiceover.
I really like the flask/jinja/bootstrap/javascript/mongodb/elastic search stack. I've learned a lot of tips and tricks by building streamJoy and if people want I would be happy to share them with the community.
I know this sounds like self promotion of my app but I haven't even taken the time to implement affiliate tracking for any service besides amazon. Consulting is serious and real money right now and that takes priority over this little side project I did.
If you are using the data from the various APIs to render static pages, why do you need the local cache of the JSON data? And if you've got a local cache of all the JSON data, why render static pages rather than serve dynamic pages that reference the DB of cached data?
> If you are using the data from the various APIs to render static pages, why do you need the local cache of the JSON data?
Because the data that I am displaying for each movie is not available from any single source. streamJoy is the abstraction layer tying together the disparate data sources.
Mongo was the right choice because instead of having to map out postgres schema, I could just store the entire JSON response as a dict.
What I have learned from all of this is that API's are not always reliable. The information changes, you don't always know what you are going to get back.
Mongo just made this early process where I didn't know what I was going to get much more fault tolerant. And, no schema mapping.
> And if you've got a local cache of all the JSON data, why render static pages rather than serve dynamic pages that reference the DB of cached data?
Rendering static pages allow you to use S3 or nginx to serve your html. For a one man operation like I am running, S3 is manna from heaven for scaling the serving of html files.
I'm using nginx right now only because I haven't taken the time to use S3, but S3 is the smarter choice here.
My goal with this was to have the smallest possible dynamic server footprint as possible.
The other reason for caching json data is I run analysis on the db items, even though I haven't published any of those features yet.
By having my own copy of the data, I can run a process on 92000 items in 6 minutes instead of taking a day due to API rate limits.
Have you tried putting CloudFlare in front of it? Buttered jelly-roll manna.
In the event that we lost a geographic region, we planned to switch to a different S3 bucket.
Personally, I use Jekyll on my own blog in a similar manner (http://andrewmunsell.com/).
< ShamelessPlug >
I also wrote a tutorial (http://www.andrewmunsell.com/tutorials/jekyll-by-example/) about using Jekyll, in case you want to try something similar to what NPR did, but with a different platform.
< /ShamelessPlug >
While I'm sure you guys already do this, proper caching can have a similar effect to a completely static site in terms of performance.
Now, if you're serving large files and you really need more than 100Mbit sustained, S3 makes sense. But it's unquestionably a premium service for a premium price.
Now, when people say that S3 scales well, I totally agree. But why do they say that the prices are competitive, that's beyond me. Take Xirra's XS-12 storage (200€/month for 36 TB) and compare it to storing 18 TB (I assume RAID 1) on S3. 1 TB costs $95 on S3! How on Earth can you call it 'competitive'? Now, that's even without bandwidth costs. I totally agree it's a premium service for a premium price. There are plenty of cases when using S3 is just a big waste of money. (And Amazon's decision not to implement cost capping isn't helping either.)
I'm curious, what are the rules/requirements for initiating a new "NPR app"? An election app seems totally obvious, but what about other apps? Is it based on available data? Available funds? Pervasiveness of a certain story? An individual reporter's weight? (for example, if I was on the team and Nina Totenberg made an app request, I'd drop everything and do it for her - she's dreamy)
Also, how much lead time do you typically get with your apps? A few days, a few weeks, longer?
Its the difference between running a wordpress server and a jekyll blog.
The security and scaling benefits are immense. With javascript you can replicate a lot of the functionality of dynamic sites.
It takes a different way of thinking because you have to design your html + js to work for everyone that accesses it.
It is still possible to personalize the content a user sees because you can still authenticate them and use ajax GET/POST requests, but you have to do all that with javascript.
Then when they authenticate, the app only uses a GET request to fetch those certain feeds specified in their preferences?
Thanks for all the info, streamjoy looks pretty sweet btw
You still need a server, but instead of your server generating the full html for each request, it only needs to generate a json response for some percentage of total page views. This is just lighter weight, and the architectural approach we are heading to with backbone.js et al...
Basically you offload as much of the computational load to the client side AND the batch rendering process as possible.
One reason this was important for my purposes with streamjoy is that even though 92000 movies isn't Big Data, it still takes a long time when you have to run 10 kinds of algorithms on 92000 items AND deal with the restricted API limits of a big API like Amazon (about one request per 1.8 seconds unless you get special permission or are doing a lot of revenue)
So for streamJoy I had to do a lot of computation and network queries, and I had to respect the 24 hour period that Amazon wants for data freshness.
So the batch jobs I had to run had to be as fast as possible, and that means the fastest CPU + SSDs + ample RAM.
A dedicated i7 server with 32gb of ram and SSDs is $180 at the cheapest data center I could find. Thats way too much for a portfolio piece!
So I offloaded all the batch computation to my own local workstation and then just push the rendered results to the cloud.
Thats one of the big benefits of the flat file approach.
Thanks for the compliment!
As someone who remembers Perl CGI's ability to serve dynamic content as being revolutionary and awesome. It amazes me to hear someone who's only experience is dynamic content and needs http file serving explained to them.
I remember reading an article, I believe by Google, about using browser sniffing instead of encoding headers to determine if gzip should be used. I can't find it now but the conclusion was that it's almost entirely safe to use gzip 100% of the time.
500-series errors indicate that a request didn't succeed, but may be retried. Though infrequent, these errors are to be expected as part of normal interaction with the service and should be explicitly handled with an exponential backoff algorithm (ideally one that utilizes jitter).
However, I still think hosting on S3 is a great option. They are pretty reliable anyway.
Terrifying! Why do you need to specify your configuration in code? I would think configuration as data is simpler.
Data => Python
Data => JavaScript
With this setup they only have one: Python => JavaScript
Granted, storing it in YAML/JSON/whatever means that you could potentially have many different codebases / languages reading it without a Language A => Language B conversion. It just depends on what works best for your project / team.Code is data.
It's a Jekyll app mostly written in CoffeeScript, deployed to S3 with CloudFront CDN'ing.
Here's a lengthier introduction for anyone interested: http://nickmerwin.com/2012/02/25/npr-io/
Summary: Go ahead and toss off an unusual term or slang, but don't expect it to actually inform readers if they can't Google it. If you aren't trying to inform readers, exactly what are you trying to do?
I found the unexplained use of the word "Boyerism" in this article to be confusing and unprofessional. I have nothing against the use of slang, or the expectation that readers will need to Google terms: this is the reality of our culture in the age of the Internet and excellent search. However, the casual and unexplained use of a private group's inside joke reveals a lack of awareness of how search interacts with culture. Is it a deliberate attempt to confuse and snub the larger audience, or is in unintentional?
That was my guess as well. From this, I take it that too little thought was given potential readers. (And if it was meant only for insiders, why is it on the public Internet?)
If only there was this network kind of thing that people all over the world could use to find and read such text...
Between the swearing, and the grossly informal dialog when explaining what should be rational and defensible technical choices:
1) I can tell.
2) It doesn't reflect [well on] NPR.
NPR advertises itself as "news & analysis", and in my experience, excels at both. A key component of "analysis" is the rational study, explanation, and discussion around (often complex and nuanced) topics.
There are few topics as complex as that of software engineering, and it also warrants due consideration and analysis.
To see engineers describing their emotive appeals as "how they roll" does not lead me to believe that NPR's hiring in their engineering department is on par with their hiring in the editorial department.
As such, it reflects badly on their engineering department, and there is a strong implication that this is not somewhere that a studious engineer would choose to work.
Contrast with the posts from Netflix on their architecture. Those posts are not unduly formal, but they provide logical, reasoned arguments, sufficient background as to judge their claims and conjecture, and demonstrate not only their own technical and engineering capacity, but their respect for the technical capacity of their reader.
Fortunately for all involved, we don't need to rope aristocracy or philosophy into the argument to provide some level of understanding of the qualitative nature of engineering, be it physical or digital. Instead, we have algorithms, maths/logic, and real measurements to serve as the bedrock of our field.
Unfortunately, articles such as this one abandon that bedrock in favor of appeals to emotion, which leaves the article (and whatever conclusions it may ostensibly provide) unsupported by fact or logic.
This loosely grounded approach to discussion is perfectly suited when discussing one's television preferences, but provides a net negative value to the world of technical discourse by propagating a culture of unsubstantiated and emotionally driven opinion and pop culture ideals.
As such, it reflects badly on their engineering department, and there is a strong implication that this is not somewhere that a studious engineer would choose to work.
Clarity in writing is not a stuffy affectation, but rather is the entire mechanism by which one both expresses an opinion, and provides an understandable, rational justification for that opinion. Without this, the reader is left with nothing but unsupported conjecture, opinion, and emotive appeals.
This implicitly calls me a liar. As I stated above, the word choice and informality don't bother me. That they spent so little time thinking if the article would make sense to a random reader does. That goes against my expectations for NPR.
In particular, I am thinking of edge cases where a news story has a typo/other important correction and you want to update just that story. What is your strategy? How is it impacted by caching done by CloudFront? Thanks.
In fact, for the 2012 election, our fallback in case EC2 was unresponsive was to decamp to a coffee shop with a laptop and s3cmd.
That particular NPR blog always has greatly insightful posts.
Some of our more dynamic projects do require a server; the inauguration project that uploaded photos to Tumblr did, for example. https://github.com/nprapps/inauguration/
When we run servers, it's usually Flask and Nginx/uWSGI.
When a hard drive crashes or a truck runs into your data center (here's looking at you, Rackspace) or you need failover for any reason, that's when you wish you had virtualized in more than one machine.
Want something that's always available and never crashes? Look at freenet. Distributed computing model. If we can failover the DNS, you can have the same thing on the web.