636 karma · joined April 27, 2015
(1) It prescribes a stripped-down version of HTML and uses a JS loader to render fast and load as much resources as possible asynchronously
(2) It caches the website in the Google CDN and delivers it via HTTP/2
By applying best practices of structuring web applications and optimizing for the critical rendering path (1) can be achieved, too. The AMP loader is just an opinionated way to do this for simple pages. Everyone can get similar results without AMP by following web performance best practices like [1]:
- Reducing the critical resources needed
- Reducing the critical bytes which must be transferred
- Loading JS, CSS and HTML templates asynchronously
- Rendering the page progressively
- Minifying & Concatenating CSS, JS and images
And there is lots of good tooling for this (e.g. postcss, processhtml, cssmin, UglifyJS, imagemin, critical, gulp-rev-all, ...)
(2) is not only harder. It is what limits the broad applicability of AMP. Cached data in the Google CDN cannot be invalidated, neither in the CDN itself nor in ISP interception caches, corporate proxies or the browser cache [2]. The effective consequence of this is that you can only use AMP when the content is mostly static. This is the case for news websites or other publications that are only changed by human editors. It completely breaks down when you try to create a dynamic site, for example a social network or a shop.
That is why our startup Baqend [3] takes a different route. We say that developers are clever enough to use existing tooling to achieve (1): an efficient rendering and loading experience. And for (2) we add a caching scheme that also employs CDNs for delivery but keeps data consistent. This is made possible by a simple process:
1. When a browser connects to a Baqend-based website, it loads a Bloom filter containing all potentially stale cached URLs.
2. Every stale URL is requested using HTTP revalidation to refresh stale copies and update caches.
3. When an update operation changes a resource (e.g. an image, JSON object or even a complex query result) its URL is instantly invalidated in the CDN [4] and marked as stale in the Bloom filter.
4. When loading resource, a statistical estimation of the expected TTL (cache liftime) is made. Whenever that prediction fails, the Bloom filter and the automatic CDN invalidation compensate the difference between estimation and real lifetime.
Using this scheme (developed at the University of Hamburg in cooperation with Baqend), every kind of dynamic data can be treated as cachable data and the applicability of is not limited to data that seldomly changes and even works for write-heavy resources with rich consistency guarantees (Δ-Atomicity, Read-Your-Writes, Monotonic Reads, Monotonic Writes, Causal Consistency) [5].
Of course I'm biased but you'd like to see AMP-like acceleration coupled with fresh cached data plus tooling, layout and frameworks of your choice, have a look at our Backend-as-a-Service.
[1] https://developers.google.com/web/fundamentals/performance/c....
[2] https://github.com/ampproject/amphtml/issues/1901
[4] https://www.fastly.com/blog/building-fast-and-reliable-purgi...
[5] http://www.slideshare.net/felixgessert/talk-cache-sketches-u...
And since we are in the Backend-as-a-Service market, the name is not all that unfitting. Although it cannot be denied that from time to time some people think we are French an spelled "Baquend".
- Every one of our servers rate limits critical resources, i.e. the ones that cannot be cached. The servers autoscale when neccessary.
- As rate limiting is expensive (you have to remember every IP/resource pair across all servers) we keep that state in a locally approximated representation using a ring buffer of Bloom filters.
- Every cacheable resource is cached in our CDN (Fastly) with TTLs estimated via an exponential decay model over past reads and writes.
- When a user exceeds his rate limit the IP is temporarily banned at the CDN-level. This is achieved through custom Varnish VCLs deployed in Fastly. Essentially the logic relies on the bakend returning a 429 Too Many Requests for a particular URL that is then cached using the requester's ID as a hash key. Using the restart mechanism of Varnish's state machine, this can be done without any performance penalty for normal requests. The duration of the ban simply is the TTL.
TL;DR: Every abusive request is detected at the backend servers using approximations via Bloom filters and then a temporary ban is cached in the CDN for that IP.
In any case I will be happy to update the article and include it.
In any case, it's difficult in terms of presentation, since a "normal" master-slave-replicated Redis is a totally different system from Redis Cluster as the distribution models of Redis Cluster also affects the functional properties to a large extend, e.g. regarding atomic Multi/Lua blocks and all types of multi-key operations.
The failover solutions for relational database systems are a good point, they too should be included.
Regarding use cases, I totally agree: every system should be discussed in the light of the use cases it tries to solve. I personally think that Redis Cluster with a choice for trading consistency against latency would open up a whole new range of use cases that are currently not a good fit for Redis Cluster. For example, we have a Redis-Coordinator project (not open source, yet) which behaves similar to a scalable Zookeeper (BTW also neither CA nor CP [1]). However, it has weaker guarantees, due to lack of tuning knobs in Redis and Redis Cluster regarding consistency.
[1] https://martin.kleppmann.com/2015/05/11/please-stop-calling-...
But you are right, I think Redis Cluster could be added as a separate system.
[1] https://aphyr.com/posts/283-jepsen-redis [2] https://github.com/antirez/redis/issues/2672
There also is a paper on the NoSQL classification scheme used in the slides: http://www.baqend.com/files/nosql-survey.pdf
This announcement will have a huge impact for the mobile dev community that relies on Backend-as-a-Service systems. I personally think that Facebook had several reasons to shut down Parse:
1) The technology stack was really fragmented and often rewritten in large parts. I still have this statistic in mind how their 200 Rails API servers were only able to serve 15 requests/s each [1]. If you look at the database technology inside Facebook, there is much superior infrastructure that was never really integrated into Parse (Haystack@OSDI'10, Tao@ATC'13, F4@OSDI'14, Extended Apache Giraph@VLDB'15, RocksDB, etc.)
2) Parse did not have a core competitive advantage: it was just the first company to whole-heartedly pick up the BaaS-paradigm with sufficient man power and a good understanding of developers' needs. The technology itself was not particularly innovative in any way, just (mostly) solid engineering. However, there remained really basic limitations that were never addressed [2]. For instance the only (!) way to safely handle concurrency control was through counter data types.
3) The model of Parse promotes independent apps and websites outside the Facebook universe.
4) The pricing in increments of guaranteed 30 requests/s was okay for simple apps but absolutely useless for anything beyond that. In particular for websites which as of 2016 do an average of 100 requests per page load [2] a single user can leave a Parse app rate-limited or down.
The main asset of Parse were their great client SDKs and well-written documentation.
This is why we made a plan: we will fork the Parse SDKs to offer seamless continuation of apps relying on them, including Push and the other features dropped in the open-source Parse Server. We opted for this approach as the Parse Server implementation on Github looks really brittle and the convoluted Parse REST API is really not an option. By doing this we hope to provide a scalable and long-term solution for developers looking to continue their Parse-based apps.
Baqend [2] is a pre-seed startup founded out of the database research group at the University of Hamburg (Germany). Our product launching into production within the next months uses a new approach to consistent web-caching reducing latency in common web workloads by up to an order of magnitude [5]. It is due to this background that we very eager to not only provide great usability (which Parse also did) but also acknowledge the need for complex data processing: low latency access, partial updates, continuous queries and ACID transactions.
We'll post a detailed plan on our blog, soon.
[1] http://blog.parse.com/learn/how-we-moved-our-api-from-ruby-t... [2] http://profi.co/all-the-limits-of-parse/ [3] http://httparchive.org/trends.php/ [4] http://www.baqend.com/ [5] http://www.btw-2015.de/res/proceedings/Hauptband/Wiss/Gesser...
Percolator is the processing system Google uses to manage its search index incrementally. It's built on BigTable which provides cell-level linearizability, similar to what CockroachDB seems to achieve by using Raft.
What CockroachDB calls "intents" are per-value columns called "write" in Percolator and ordered by BigTables timestamp system. Percolators "master lock" is rebranded to "switch" and, tada, there you have CockroachDBs fancy lockless algorithm.
Edit: Amazon DynamoDB, too, uses a very similar approach to this in their client-driven transaction implementation https://github.com/awslabs/dynamodb-transactions/blob/master...
It would be interesting to have one of the CochroachDB guys comment on the novelty of their approach.