We reduced the cost of building Mastodon at Twitter-scale by 100x
blog.redplanetlabs.com
blog.redplanetlabs.com
I wish I didn't see this comparison, which is not fair at all. Everyone in their right mind understands that the number of features is much less, that's why you have 10k lines.
Add large-scale distributed live video support at the top of that, and you won't get any close to 10k lines. It's only one of many many examples. I really wish you compare Mastodon to Twitter 0.1 and don't do false advertising
> 100M bots posting 3,500 times per second... to demonstrate its scale
I'm wondering why 100M bots post only 3500 times per second? Is it 3500 per second for each bot? Seems like it's not, since https termination will consume the most of resources in this case. So I'm afraid it's just not enough.
When I worked in Statuspage, we had support of 50-100k requests per second, because this is how it works - you have spikes, and traffic which is not evenly distributed. TBH, if it's only 3500 per second total, then I have to admit it is not enough.
Mastodon actually has more features than the original Twitter consumer product like hashtag follows, global timelines, and more sophisticated filtering/muting capabilities.
Some people argue it's not so expensive to build a scalable Twitter with modern tools, which is why we also included the comparison against Threads. That's a very recent data point showing how ridiculously expensive it is to build applications like this, and they didn't even start from scratch as Instagram/Meta already had infrastructure powering similar products.
With things like twitter, the ui is not the hard part. Things like moderation are the secret sauce. All the corner cases and support for devopsy stuff likely account for a lot. Routing to specific instances for celebrities and such.
I originally read it as it's easy to clone twitter.
My response is it's easy enough to build a micro-blogging platform/service. it's all the other shit like moderation, regulatory/legal compliance, making a profit, keeping advertisers happy, etc that's hard.
So this is not meant to diminish the achievements of what you have built at all, it is more intellectually honest to say that "any high performance framework is most suitable for projects that are exact clones of pre-existing, mature things with battle-hardened specifications and end user behavior." While this might cover some greenfield projects, including the best capitalized ones that may matter to you, it does diminish the appeal of a framework for the vast majority of success stories from small & poorly capitalized teams. Those small & poor teams are very innovation and serendipity driven and hence rarely copying a pre-existing thing. And even if they try to become well-capitalized, they are almost always doing so by having worked on the thing they are copying already (i.e., already shipping version 1.0 for years).
It's too easy to study an existing system and, given the resources, create a perfect demo for how to dramatically improve it in certain ways. You can go on to ship the demo, but the demo wasn't made by the same kind of organization, with the same kinds of goals, as the original system builders. The extreme example of this is, of course, the demoscene and its hardware-bending tricks that achieve the impossible through a significant modification to the design of a "production" equivalent.
So it's better by performance metrics, better by codesize, but unknown on other metrics. Like, "do I know how to start building new things with this?"
Your comment makes it seem like their post is misleading, when it isn’t it just might be that what they do best isn’t useful to you (or possibly anyone).
I agree that copying an existing product will be easier and is usually cheaper and more performant because you can leverage your competitor’s R&D and lessons learned. I presume this is why some tech companies’ product lines are full of clones.
I don’t know if I’d suggest ECS for every team as you need the right tooling, culture, and leadership to pull it off. That said, I think the paradigm has unrecognized benefits when it comes to greenfield gameplay programming.
Or, your data structures become a way for sub-teams to communicate and share state without stepping on one another’s toes. A group handling pathfinding and a different one handling scenario/mission logic can both pay attention to the position data without needing to be highly coupled. It becomes easy to “just add a system” or “an extra component” to build on existing functionality.
The Fred Brooks quote about “showing me your tables” comes to mind.
What do you mean by normal approach? OO classes with inheritance?
When I was at IG, my team was 20-something ICs full-time on our project, bursted to maybe double that as necessary in part-time help from ICs in the wider org. We had a total of three backend engineers, of which I was one, as well as a backend intern for a few months, although we bursted the backend team for a couple of months to four for some ML help.
Your project sounds pretty cool! But I don't think the comparison to Threads productivity is quite right. The majority of IG engineering is focused on building polished native clients, not generally on backend infrastructure. It's true that Meta already has a lot of the backend infra built — if Red Planet Labs can come close to what you get at Meta, that's pretty amazing. But I don't think the numbers you're quoting are apples-to-apples, or mean quite what you think they do. I don't know if this version of Threads operated exactly like the one I worked on (a Snapchat competitor), but I'd be pretty surprised if there was a product team at IG that was majority backend.
(Edit: I also think it's worth keeping in mind the experience levels of who works on these projects — at IG it's usually a smaller number of E6/E5s guiding a larger number of E4/E3s. Person-years are not all equivalent! If you spent ten years building the Red Planet Labs infrastructure based on your time at Twitter, nine person-months of your team's time building a product on your infra might not be the same as nine person-months of someone else's.)
Anyway — I don't mean to downplay your product, and really, if it's anything like the backend productivity I experienced at Meta, that would be pretty groundbreaking. Curious to see what you launch :)
"But it's not a fully functioning 2023 Twitter!!!" I think some people miss the point. This is not about hey we built a Twitter clone. This is about a POC for a novel app architecture.
We need to be constantly examining and re-examining our thoughts about the best way to deal with distributed systems, scale, developer workflow. Even inventing new ones.
Then they shouldn’t have titled it “we built a Twitter clone”
Granted, a decentralized platform would eliminate some of those, just by being decentralized
Mastodon still requires security, compliance and moderation. And those requirements are going to keep getting more challenging by the year. It'll end up being another reason nobody will want to host content in a decentralized manner, the burden will become obnoxious.
Everyone (theoretically) would be complying to statues in its broadest sense, but jurisdiction, regulations, industry best practices, reporting requirements, and appetite for risk is going to be different from organization to organization. It’s not one-size-fits-all.
So some of these things are eliminated because the people hosting those are not put under the same kind of scrutiny as say, a Twitter.
My answer, born of moderating a modestly sized forum, was "Absolutely not under any circumstances."
>> 100M bots posting 3,500 times per second...
and
> We used the OpenAI API to generate 50,000 statuses for the bots to choose from at random.
I wonder: 100M OpenAI bots talking to each other continuously and with much vigor - how is this affecting OpenAI’s uhm… intellect?
But Twitter isn't, and was never, about live video support: this is pure feature creep and that's how you get headcount inflation and a company that can be run for 17 years without making profit (AKA terrible business).
> When I worked in Statuspage, we had support of 50-100k requests per second
Having served 150kqps in the past as part of a very small team (3 back-end eng.), this isn't necessarily as big of a deal as you make it sound: it mostly depends on your workload and whether or not you need consistency (or even persistence at all) in your data.
In practice, building scalable system is hard mostly because it's hard to get the management forgot their vanity ideas that go against your (their, actually) system's scalability.
I would love to see a blog post about your experience!
The background makes this a bit more interesting, because you can imagine how those early days impacted the arc of his work.
Serving one cached HTML with Nginx and serving dynamically generated content, updating a database is a different thing.
> Add large-scale distributed live video support at the top of that,
Why? For the love of all that is good and efficient, why? Why not have a separate platform for that? Or link to a different federated video service? Why does every platform need to do all the things?
StartHalfLife2LikeGame()
and you'd replace millions LOC with one line. If there is a perfect match between the framework and your app, there is no code. The more your app diverges from the ideal app the framework was written for, the more code you have.But this is the core [0] - a simple event location plattform my wife founded hat easily >200k LOC while the core (edit/search/detail page event location) had <10k LOC.
[0] If building Twitter I perhaps would try their setup, although I have been bitten with "magic" JVM frameworks in the past, e.G. we used one commercially and the license cost went up from 80k to 800k YoY. On top of scaling problems we could do nothing about because the framework was a black box.
I'm not saying there's nothing here, but I am adjacent to your core audience and I have no idea whether there is after reading your post. I think you are strongly assuming a shared basis where everybody has worked on the same kind of large scale web app before; I would find it much more useful to have an overview of, "This what you would usually do, here are the problems with it, here is what we do instead" with side by side code comparison of Rama vs what a newbie is likely to hack together with single instance postgres.
EDIT: Specifics
Here, the "views" are defined formally (the P-states), and incrementally, automatically updated when the underlying data changes.
Example problem:
Get a list of accounts that follow account 1306
"Classic architecture":
- Naive approach. Search through all accounts follow lists for "1306". Super slow, scales terribly with # of accounts.
- Normal approach. Create a "followed by" table, update it whenever an account follows / unfollows / is deleted / is blocked.
Normal sounds good, but add 10x features, or 1000x users, and it gets trickier. You need to make a new table for each feature, and add conditions to the update calls, and they start overlapping... Or you have to split the database up so it scales, but then you have to pay attention to consistency, and watch which order stuff gets updated in.
Their solution is separating the "true" data tables from the "view" tables, formally defining the relationship between the two, and creating the "view" tables magically behind the scenes.
It’s a ridiculous amount of fluff to describe that. Not to mention it’s proprietary and only supports the JVM and doesn’t integrate with the tons of tooling designed about RDBMS unless you stream everything to them, defeating the purpose.
What really irks me is that they go on and on bragging about the low LoC count and literally show nothing complete. They should’ve held on this post and released it simultaneously with the code.
Oracle has decent support for incrementally updated materialized views, redshift has some too. Materialize.com is an entire snowflake-like platform built around incrementally maintained materialized views.
Once SQL materialized views aren't enough, you might do this by replicating your database into Kafka, implementing logic in Flink or something, and reinserting into the same DB/Elasticsearch/etc. Very common architecture. (Writ small, could also use a queue processor like RabbitMQ.)
Their approach is to instead--apparently--make all of these first-class elements of the same ecosystem, not by "putting it all in the database", but by putting the database into the application code. Which seems wild, but colocates data, transformation, and view.
Seems like it would open up a lot of cans of worms, but if you solve those, sounds great.
Individually, none of these concepts are new. I’m sure you’ve seen them all before. You may be tempted to dismiss Rama’s programming model as just a combination of event sourcing and materialized views. But what Rama does is integrate and generalize these concepts to such an extent that you can build entire backends end-to-end without any of the impedance mismatches or complexity that characterize and overwhelm existing systems.
Indexes as arbitrary data structures that you shape to perfectly meet your use cases, a powerful computation API that's like a "distributed programming language", and everything being so integrated make a world of difference.I understand the desire to see all the code, and that's coming in two weeks. That said, the code in the post isn't trivial as it's showing almost the complete implementations of two major parts of Mastodon: the social graph and timeline fanout.
Next week you'll be able to play with Rama when we release a build of it, and the documentation will help with that.
Every time I hear this the reality turns out to be that building anything with this tech is like building something on top of SAP.
But I’m also just allergic to any post that says ‘look how amazing’ in general, so I’m a bit prejudiced.
Just looking at the first example tells me that there’s a million ways someone that doesn’t know what they’re doing can mess this up.
If the author of the platform implements some service on their own platform it’s always going to seem simple.
How about you go and implement a Mastodon server to their level of feature parity, and tell us how much effort and how many lines of code it takes?
I really don't appreciate this kind of fluffy, insubstantial, overly dismissive non-content on HN.
Obviously, with tons of implementation difficulties and details, and not actual graph structures, but as a top level analogy.
To be honest any parallel with frontend here is meaningless, reactivity and all the concepts at play have existed long before JS and browsers came along, it’s easier to explain from first principles.
[1] https://news.ycombinator.com/item?id=29615085
[2] https://martin.kleppmann.com/2015/03/04/turning-the-database...
Have you considered using that?
> To minimize memory usage and GC pressure, we use a ring buffer and Java primitives to represent each home timeline. The buffer contains pairs of author ID and status ID. The author ID is stored along with the status ID since it is static information that will never change, and materializing it means that information doesn’t need to be looked up at query time. The home timeline stores the most recent 600 statuses, so the buffer size is 1,200 to accommodate each author ID and status ID pair. The size is fixed since storing full timelines would require a prohibitive amount of memory (the number of statuses times the average number of followers).
> Each user utilizes about 10kb of memory to represent their home timeline. For a Twitter-scale deployment of 500M users, that requires about 4.7TB of memory total around the cluster, which is easily achievable.
Isn't this where the most difficult(expensive) part is and Rama has little to do with it? It appears that the other parts also do not have to be Rama.
> So instead of having to do network operations, serialization, and deserialization, the reads and writes to home timelines in our implementation are literally just in-memory operations on a hash map. This is dramatically simpler and more efficient than operating a separate in-memory database.
Updates per second to end users who follow the 7K tweets per second seems more realistic, it's the timelines and notifications that hurt, not the top of ingest tweets per second prior to the fan out... and then of course it's whether you can do that continuously so as not to back up on it.
Another important metric is "time to deliver to follower timelines", which is tricky due to how much variance there can be every second due to the extremely unbalanced social graph. When someone with 20M followers posts, that multiples the number of needed timeline writes by 15x. We went into depth in our post on how we handled that to provide fairness by preventing these big users from hogging all the resources all at once.
Now, the really hard part becomes selling. If companies start using your product to get ahead, that will be the real proof, otherwise its "just" tech that is good on paper.
On a side note, did you guys got any inspiration from clojure? I see lots of interesting projects propping up from clojure people...
Best of luck!
Interesting to finally see the announcement.
This write-up is very detailed but I couldn't find that explanation.
One famous example of this going to far: Mac Mail app used to play a whoosh sound when your email is actually sent. They changed it to whoosh instantly no matter what. Given how often an email might fail to send or get delayed, this meant an actually useful indication of "great, your thing was sent, you can close your laptop now" was rendered useless.
HN does this, and on slow days, about half of my upvotes don't go through.
Unfortunately the js used is very minimal. And if the upvote fails, there is no graphical indication that it failed to upvote.
The advantage of the overall architecture is that nearly all application functionality (for something like a social network) can tolerate much higher latency than an RDBMS, so you really want to have architectural building blocks that let you actually use this headroom.
You write the update directly to the cache closest to the user and into the eventually consistent queue.
We did this at reddit. When you make a comment the HTML is rendered and put straight into the cache, and the raw text is put into the queue to go into the database. Same with votes. I suspect they do this client side now, which is now the closest cache to the user, but back then it was the server cache.
I have to say in my ~12 years as an active Redditor I can’t recall a time where I saw any real state issues, even with rapidly changing votes, etc. Bravo!? Now that we’re beyond the days of molten servers, I have to say its overall reliability in the face of massive spiky traffic is quite a feat.
Or, maybe their pitch is that the streaming bits are so fast, you can just await the downstream commit of some write to a depot and it'll be as fast as a normal SQL UPDATE.
Reloading the page from scratch can be slow due to Soapbox doing a lot of stuff asynchronously from scratch (Soapbox is the open-source Mastodon interface that we're using to serve the frontend). https://soapbox.pub/
Rama's built-in telemetry provides the information you need to know when it's time to scale.
Within an ETL, when the computations you do on PStates are colocated with them, you always read your own writes.
Internally Rama does a lot of dynamic auto-batching for both depot appends and stream ETLs to amortize the cost of things like replication. So additional colocated stream topologies don't necessarily add much cost (though that depends on how complex the topology is, of course).
But I'm kind of getting the impression this works without any speed layer and is expected to be fast enough as-is.
Rama is not batch-based. That is, PStates are not materialized by recomputing from scratch. They're incrementally updated either with stream or microbatch processing. But PStates can be recomputed from the source data on depots if needed.
1. send event data to depot
2. trigger localized ETLs (or put it high-priority in queue) to recalculate just the impacted data into relevant PStates
3. await completion of aforementioned ETLs
4. run query from updated PStates
Maybe too heavy for an upvote, but very appropriate for a an important transaction like a purchase.
e.g. they mention "event sourcing" and "materialized views" in the post -- sounds good
But I thought I heard from a few people who were like "we ripped event sourcing" out of our codebase and so forth
And yeah your question is an obvious good one, and the Reddit answer of "write through cache" ... is less than satisfying to me
I FREQUENTLY have the problem where I reload the page and Reddit shows me stale data. It's SUPER buggy.
---
Anyway I definitely look forward to hearing people try this and what their longer term impressions are !
I basically want to know what the tradeoffs are -- it sounds good, but there are always tradeoffs
So is the tradeoff "eventual consistency" ? What are the other tradeoffs?
I was pissed off that I would have to type my comment again, but actually it did save it, and refreshing worked.
From what I understand Hacker News is architected more in-memory, on one big box ... Perhaps similar to the event sourcing model
(not knocking hacker news -- it's generally a very fast site, MUCH better than Reddit. Just that scaling beyond a single machine is difficult and full of gotchas )
Of course, specialist systems can often do much better.
The way applications are built, and have been built since before I was born, is by combining together potentially dozens of narrow tools together: databases, computation systems, caches, monitoring tools, etc. There has never been a cohesive model capable of expressing arbitrary backends end-to-end, and every application built has to be twisted to fit onto the existing narrow pieces.
Rama is a lot more than just "event sourcing" and "materialized views". Those are two concepts at its foundation, but the real breakthrough is being that cohesive model capable of expressing diverse backends in their entirety. It took me more than five years of dedicated research to discover this model, and it was extremely difficult.
But what are the tradeoffs? There's nothing that comes with 100x benefit with no tradeoffs
(side note: I worked on Google Code for a short while in 2008, concurrent with Github's founding ... I think Github moved a lot faster in a large part because they weren't dealing with distributed systems at first -- they had a Rails app, a database, and RAID disks, and grew it from there. We had BigTable and perf reviews :-P )
Eventual consistency is probably one?
Can I specify that comment editing is correct and ACID, while likes/upvotes are eventually consistent? (No is a fine answer, these problems are hard)
I read through much of the doc, and don't see a mention of the word "consistency" at all, which seems like an oversight for something that is unifying what would be in a database with computation.
You get read-after-write consistency for any PStates in a streaming ETL colocated with the depot you appended to. This is if you do the depot append with "full acking", which coordinates its response with the completion of colocated streaming ETLs. If you append at a lower level of acking, then you get eventual consistency on those PStates at the benefit of lower latency appends.
Microbatching is always eventually consistent as processing is asynchronous and uncoordinated to depot appends. Microbatching is higher thorughput than streaming and has simpler fault-tolerance semantics.
You'll be able to read a lot more about this when we release the docs next week.
I think many people are going to have problems programming with this consistency model, as they will with any that's different than a single machine. But that's basically "physics", so it's inevitable :)
But it seems like great work within the constraints -- look forward to learning more
I have indeed wondered why none of the cloud platforms have built more forward-looking tech like this -- instead it's copies of AWS and so forth
For the client, it's feasible to retain the timestamp of its most recent write. In this way, the system can ensure that the replica responsible for any reads related to that user incorporates updates at minimum up to that recorded timestamp. If a replica isn't adequately current, the read can either be managed by another replica or the query can wait until the replica catches up. The timestamp might take the form of a logical timestamp, signifying the order of writes (e.g., log sequence number), or it could be based on the actual system clock, where synchronized clocks become vital.
When your replicas are spread across multiple datacenters—whether for user proximity or enhanced availability—there's an added layer of complexity. Requests requiring the leader's involvement must be directed to the datacenter housing the leader.
https://avikdas.com/2020/04/13/scalability-concepts-read-aft...
https://en.wikipedia.org/wiki/Cache_%28computing%29#Writing_...
https://cloud.google.com/blog/products/databases/why-you-sho...
That does limit you to operations/queries you can describe in this dual format, but pretty often that's fine. Or if you can relax read-after-write you can just ignore the in-flight stuff and read from the main store and then there are no (added) limitations.
It’s hard enough trusting Google or Amazons cloud offerings won’t change.
It seems that’s what they’re proposing right? What am I missing?
Second, Rama has a pure Java API and is not a bespoke language. So no new language needs to be learned.
Isn't Mastodon a Ruby On Rails application?
And the Rama core will remain closed-source? That part seems like the toughest sell of all, at a time when the vast majority of developer tooling and backends are open source or at the very least source-available.
We're keeping it closed-source for now.
Rama sounds interesting to me for my 'next big project', but I'd not even consider building it on top of a closed core. I think this is a pretty common sentiment in these circles.
I understand building an OSS business is not easy either. But perhaps there is some middle of the road that you can walk?
- A contractual obligation to open source all (now current) code a couple of years in the future? - Or an almost-OSS license that makes life difficult for competing cloud providers, like https://www.mongodb.com/licensing/server-side-public-license... ?
Big if true (and if the opposite, of incrementally removing it also works). There have been similar platform efforts in past, such as https://news.ycombinator.com/item?id=20985429 . For that one, the "massive ask to give up every programming language and database they’ve ever used to depend on a startups closed source platform" seems like the biggest hindrance to adoption.
It’s hard to imagine it for a complex legacy application without having lots of added complexity. It wants to be the unifying programming model for the application. It would seem like running with two RDMS sources of truth simultaneously.
It’s like the xkcd “there are 12 ways of doing X, let’s create a standard to unify them” now there are 13 ways
9, which is 3^2, and 27, which is 3^3. Or 900 is Yoda's age, and 27 which is the 27 club of musicians who committed suicide.
Still super impressive. Reminds me of when I discovered Elixir while building a social-ish music discovery app. Switching the backend from Rails to Elixir felt like putting on clothes that actually fit after wearing old sweats. Rama looks like a similar jump, but another layer up, encompassing system architecture.
The numbers they got for Twitter likely include the time it took to build their infrastructure, common libraries (like finagle,…)
It’s also unsure what we would compare a tool like this to. I doubt you could just say “compare it to Rails” given how frameworks like rails are bound to specific data models, and most realistic applications. You’d have to compare it to some other opinion about how to wire together different data structures.
I think they have thought a lot about typical hard problems, such as having the timeline processing happen along side the pipeline, taking network / storage etc out of the picture. Nice work!
The way Ignite works overall is similar. You make a cluster of JVM processes, your data partitioned and replicated across the cluster, and you upload some JARs of business logic to the cluster to do things. Your business logic can specify locality so it runs on the same nodes as the relevant data, which ideally makes things a lot faster compared to systems where you need to pull all your data across the wire from a DB. Like Rama, Ignite uses a Java API for everything, including serializing and storing plain 'ol java objects.
Ignite's architecture isn't focused on "ETL" into "PStates". Instead it's more about distributed "caches" of data. It does have streaming for ingestion (https://ignite.apache.org/docs/latest/data-streaming), but you can transactionally update the datastore directly (https://ignite.apache.org/docs/latest/key-value-api/transact...). It also has a "continuous query" feature for those reactive queries to retrieve data (https://ignite.apache.org/docs/latest/key-value-api/continuo...).
Rama's data-structure oriented PState index seems easier to work with than building indexes yourself on top of Ignite's KV cache, but Ignite also offers an SQL language, so you can insert your data into the KV cache however, add some custom SQL functions, and then accept more flexible SQL querying of your data compared to the very purpose-built PCache things, but still be able to do lower-level or more performance-oriented logic with data locality.
Anyways, if you like some of this stuff but want to use an existing, already battle-tested open source project, you can look for these "in-memory data grid", "distributed cache", kind of projects. There's a few more out there that have similar JVM cluster computing models.
Also would love to hear folks’ thoughts on the sort of usecase where this data grid excels.
It's not so much that I think the comment is wrong or anything, but rather that it seems so similar to what I have heard in the past from power-lisp (or Clojure in this case) super-smart engineers.
I feel like we have reached a point in software development where "better" paradigms don't necessarily gain much adoption. But if Rama wins in the marketplace, that will be interesting. And I am quite excited to see what a smart tech leader and good team have been able to grind out given a years-long timeframe in this programming platform space . . .
This is meant to be hyped to sell your Rama platform/product/framework? That you have spent 10 years building in secret? During that time you have built a datastore and a Kafke competitor and ?
Should not those 10 years be factored into the time it took to develop this technical demo?
Is it 100x less code including every LOC in all of Rama?
I mean I am sure you picked a use cast that is well suited to creating a Twitterish architecture implementation.
If I went off and wrote a ThinkBeat platform for creating Twitterish systems and then created a Twitterish implementation on top if it, its real easy to reach low LOCs.
For the former, Rama has a class called "InProcessCluster" that works identically to a real cluster. It enables Rama applications to be tested and experimented with end-to-end. There's an example of this in the post and this is what we're releasing next week.
For the latter, Rama can be run on a single node with each daemon and module being a separate process. We made it really easy to launch single-node Rama instances with just a couple commands with the "rama" script that comes with the release. That said, we haven't spent much time yet optimizing small-scale Rama deployments and there's likely things we can do to make it more efficient (e.g. combine the Conductor and Supervisor daemons into a single process).
It's the ability to avoid the impedance mismatches which dominate existing tooling that makes such a difference. With existing databases, including RDBMS's, you have to twist your application to fit their data models. The existence of things like ORMs help, but they add their own layers of complexity.
With Rama, you mold your indexes to exactly match your application's needs. And you're always just working with objects represented however you want, whether appending data to depots, processing data in ETLs, or storing data in PStates.
That computation and storage are integrated and colocated is another way that Rama simplifies application development and deployment.
Their "variables" have names that you have to keep as Java strings and pass to random functions. If you want composable code, you don't declare a function, you call .macro(). For control flow and loops, you don't use if and for, but a weird abstraction of theirs.
I feel like this code could have been a lot simpler if it was written in a specialized language (or a mainstream language with a specialized transpiler and/or Macro capabilities.)
I'd quote the old adage about every big program containing a slow and buggy implementation of Common Lisp, but considering that this thing is written in Clojure, the authors have probably heard it before.
I've been digging around for a while and haven't found any posts with more than 20 faves. The accounts I've found with ~1 million followers have little to no engagement. I want to see how a post with a million faves holds up to the promises of "fast constant time".
I'm especially curious about these queries — fave-count and has-user-faved — since a couple years ago Twitter stopped checking has-user-faved when rendering posts more than a month or so old, so I imagine it was expensive at scale.
Tracking reactions is considerably easier than timeline fanout though, as a favorite does a small handful of things (updates set of users favoriting a status and sending a notification), while fanout has to do an operation on every follower (403 operations on average, sometimes up to 22M).
The code getting the favorite count for a status looks like:
.localSelect("$$statusIdToFavoriters", Path.key("*statusId").view(Ops.SIZE)).out("*numFavorites")
Because the nested set is subindexed, that's an extremely fast operation (looking at our telemetry, about 0.05ms).Determining "has-user-faved" looks like:
.localSelect("$$statusIdToFavoriters", Path.key("*statusId").view(Ops.CONTAINS, "*accountId")).out("\*hasFavorited")
The API server doesn't do these queries individually, which would be two roundtrips. It does them together in a query topology along with fetching other needed information (like number of boosts, number of replies, "has boosted", "has muted this status", etc.).So I guess, if you say "it's a Mastodon-clone", you cannot be accused of taking proprietary ideas from Twitter (this is just a guess, they know better).
But technically very interesting and refreshing to see. I really like their approach. It feels they are innovative.
And it's a very slim ActivityPub inplementation. For example, I don't think you can do basic things like get an individual post in ActivityPub. This should be easy simple json-ld to get but it's just 404. https://www.w3.org/TR/activitypub/#retrieving-objects
curl -L -H 'accept: application/activity+json' 'https://mastodon.social/users/Gargron/statuses/18614983'
It does have a bunch of stuff that isn't federated though, such as Like counts/collections. And of course it only implements the server-to-server (S2S) part of AP, not the client-to-server (C2S) part.And either way, I think the source code to their Mastodonlike will not be usable since it will be running on their Rama server framework.
Mastodon is the name of a piece of software as well as an API, a website, etc.
Naming this stuff is hard but calling it a Mastodon instance would be more confusing.
Except "Mastodon instance" means an instance of Mastodon, which is open source. Whether or not it was intended to be deceptive (I'd think a group of smart people would know better), this personally left a bad taste in my mouth.
They federated this brand new code in 9 months, and bluesky still hasn't released anything regarding federation. Don't keep your hopes up, it would kill their business model to let anyone run part of the network. People-driven networks are just not compatible with commercially driven ones, name one successful example.
So is Twitter not a Twitter instance? Like if it looks, walks and toots like a Mastodon, is it not a Mastodon instance?
[1] https://www.wired.com/2013/09/the-second-coming-of-java/
The basic operation Rama provides for evolving an application over time is "module update". This lets you update the code for an existing module, including adding new depots, PStates, and topologies.
FWIW, why hype at all? Why "We'll more in a week. Then more in two weeks." Show the code today!
I will make one minor suggestion that I hope is constructive. I found the post difficult to read, largely because you rapid fire introduce a bunch of completely new concepts and propose a solution to many problems at once. You make a passing comparison to "just event sourcing and materialized views", although this was the easiest way for me to understand what you are doing. Starting from event sourcing and materialized views puts the reader on a ground they already understand, and moving on from there to why rama is better/what it adds on top, would be an easier transition.
i mean if you think about this as public services not as a business, profit is secondary, and first is just to make the thing better and better for the users, no need for spying , no advertisement, no need for a rich piece of shit somewhere getting a piece of the money paid in your city for every taxi drive, food delivery or to give up privacy to a soulless/faceless entity just because you want to say something publicly or keep in touch with people. there is no disruption from their part, its just an old thing put on the internet, they are just in the middle of everyone's life, just sucking everything they can. is the actual state of affairs "efficient"?
there must be fed up engineers and tech people everywhere with the sad state of IT industry.
Saying this is 100 or even a million times cheaper is like saying taking a picture of Sistine chapel and printing out copies is a trillion times cheaper than making it originally.
Many of us on this site could make a number of products very efficiently and cheaply given a static and fixed set of requirements as well as an existing implementation for reference.
That being said it was a very detailed post, so kudos for that, but it’s far too vague to be actionable. Why not just release the code and post simultaneously instead of just bragging about how little code was required?
So no one will be able to run this except on the proprietary cloud
Any attempt to build a simplified version of the ecosystem will face the same fundamental distributed system tradeoffs like consistency/reliability/flexibility/... For example, one of the simplifications may be mixing storage/serving/ETL workloads on the same node. And the consequence is that without certain level of performance isolation, it could impact the serving latency during expensive ETL workload.
For Rama to be adopted successfully, I think it is important to identify areas where it has the most strengths, and low LOCs might not be the only thing that matters. For example, demonstrating why it is much better/easier than setting up Kafka/Spark and a database and build a Twitter clone on top of that while providing similar/better performance/reliability/extensibility/maintainability/... is a much stronger argument.
Mastodon has to send messages to each instance with a recipient. That server can then fan out to all it's subscribers. The way this point is worded makes me think all the bits are on just a single instance, meaning all the fan out can be dealt with internally without having to do any server-to-server at all.
That is a fair comparison to Twitter, which is single instance. But it sounds like a much reduced ambition versus the task Mastodon has to do.
Of course, for actual production use, there's probably a lot of things still, but this is a very nice works nonetheless
> The instance has 100M bots posting 3,500 times per second at 403 average fanout to demonstrate its scale.
If one clicks quick enough to jump to an actual post, it seems relatively static so it's hard to tell if the bots are deleting and recreating their posts or what. In true Xitter clone fashion, trying to view the Posts & replies from any one user is "sign in
Anyway, all of this is not to detract from your framework announcement as much as to have you consider that it's perfectly fine to label that instance as a load test, that's a fine thing, but calling it a legitimate instance seems to be a potential source of confusion
To get a better feeling of Rama's performance on your hardware, I suggest registering an account which will allow you to poke around the whole platform. It takes just a couple seconds to register and we don't send any emails.
One question, why Google Groups rather than something like Discord? Not sure I would trust Google Groups to be around long.
Are there any plans for exposing a Clojure API? Given that it's implemented in Clojure, seems like it would be a natural fit. Interop with Java is nice but can be cumbersome compared to the more natural calling conventions and idioms (threading macros instead of `..` builder patterns, etc).
Second whats the product/business angle on customer confidence, technical novelty, and your business core competency? A dated example but Im thinking of somewhere like basho with riak. Super cool tool, takes some mental adjustment to “get”, challlenges selling hosting vs software vs pro services.
We'll likely have a fully managed cloud version in the future.
Riak was good technology at the time, but it was really hard to distinguish from other K/V databases and didn't really move the needle on core business metrics (like development cost). Rama dramatically changes the economics of building large-scale software. It will take awhile for many to grasp that, as is obvious from many of the comments here, but I expect that as more and more users have massive success with Rama, that understanding will come.
Having a "Rama" for local-first, truly distributed (Solid Project style), self-sovereign-identity based apps would be differentially better probably.
However, I do not see Rama's initial market being startups, since they just want the simplest way possible to build UI + backend and want to iterate super fast with tech that their developers already know in the initial stages.
It takes strong conviction to work on something like this for a decade.
Is Twitters 7k tweets per second the average? If so, what’s the peak rate, and have you tested your system under this load?
That's a bit like starting an Oracle clone now and summing up what they spent on developer salaries in the last 40 years. You basically can't not "reduce costs".
And no "the original consumer product" is not a real cop-out, you probably still have tons of people building iterations.
That said: You need better advisors. Your investors and/or the board gave you bad advice on how to publish these accomplishments and talk about them.
I hope your go-to-market strategy works out a little better. Hyperbole is fine, but at least on hacker news, the audience is a bit careful with regards to grandiose statements.
What might work well on an investor presentation might backfire when you target engineers as audience.
I saw the Twitter post first and the blog next. The premise is compelling but it's been a promise made to the data and software world for decades together.
The architecture and the core primitives are something that we agree with a lot. Use cases and business value are a whole different ballgame.
We have invested the past 5 years at InfinyOn building Fluvio our open source rust implementation of core event streaming primitives which is implementing this architecture to orchestrate data as efficiently as computationally possible today. I am happy to see this project as an effort in the same direction.
If you need an example, Kotlin uses a nice internal DSL for HTML where you write things like
Div {
P {
+"Hello world"
}
}
This is valid Kotlin that happens to construct a dom tree. There is no magic going on here; just usage of a few syntax features of the language. Div and p here are normal functions that take a receiver block as the last parameter that receive an instance of the element they are creating. The + in front of the string is function with a header like this in the implementation of that element. operator fun String.unaryPlus()
The same code in Java would be a lot messier because of the lack of syntactical support for this. You basically end up with a lot of method chaining, semicolumns and their convoluted lamda syntax. The article has a few examples that would look a lot cleaner if you'd rewrite them in Kotlin.Could you please compare Rama with Kafka Streams, especially from the point of view, if I would try to reimplement Rama API on top of Kafka Streams? What fundamental difficulties I'd face?
What is it? build web-scale reactive backends with an expressive java dataflow API. Instead of a database you develop your own custom app-specific indexes which are reactive, distributed and durable. It's like event sourcing and materialized views but integrated in a linearly scalable way.
> I cannot emphasize enough how much interacting with indexes as regular data structures instead of magical “data models” liberates backend programming
> It allows for true incremental reactivity from the backend up through the frontend. ... enable UI frameworks to be fully incremental instead of doing expensive diffs to find out what changed.
Ok, so in my mind I am positioning this against Materialized / differential dataflow, whose key primitive is a efficient streaming incremental join that works across very large relational tables. Materialized makes SQL reactive, Rama gives you a java dataflow DSL for developing purpose-built reactive database indexes.
How it works? 4 concepts: Depot, ETLs, PState, Query
Depots: "distributed, durable, and replicated logs of data." [Event streams?] "like Kafka except integrated" "All data coming into Rama comes in through depot appends."
ETLs: data arrives via depots, and is ETLed to PStates via "a Java dataflow API for coding topologies that is extremely expressive". "Most of the time spent programming Rama is spent making ETLs."
PStates seem like reactive data structures that are also durable/replicated, these are meant to supersede your database and indexes, letting you build custom purpose-built indexes that contain 100M elements:
> “partitioned states” are how data is indexed in Rama ... Unlike existing databases, which have rigid indexing models (e.g. “key-value”, “relational”, “column-oriented”, “document”, “graph”, etc.), PStates have a flexible indexing model. In fact, they have an indexing model already familiar to every programmer: data structures. A PState is an arbitrary combination of data structures. ... nested data structures can efficiently contain hundreds of millions of elements. For example, a “map of maps” is equivalent to a “document database”, and a “map of subindexed sorted maps” is equivalent to a “column-oriented database”. Any [composition] is valid – e.g. you can have a “map of lists of subindexed maps of lists of subindexed sets”.
Query: once you develop PStates to aggregate relevant data into a custom index of the right ... shape?, query seems sorta like GraphQL selectors over your custom index:
> Queries in Rama take advantage of the data structure orientation of PStates with a “path-based” API that allows you to concisely fetch and aggregate data from a single partition
> “query topologies” ... real-time distributed querying and aggregation over an arbitrary collection of PStates. These are the analogue of “predefined queries” in traditional databases, except programmed via the same Java API as used to program ETLs and far more capable.
It's something else that maybe speaks the Mastodon API and/or ActivityPub, but we don't know since it doesn't really federate with anyone.
I commend the effort to try to make happen a non-open fediverse service, but appropriating the Mastodon name is just wrong. You should know better.
GC tuning on the JVM is much less of a topic these days than it used to be. The default garbage collector was changed at some point (G1). It has some configuration options but they come with sane defaults that mostly just work fine and adapt to your memory and cpus. You don't spend a lot (or any) time on tuning this typically. I know I haven't even looked at GC params in many years now. Never had to. And we run on modestly sized vms of 1 or 2 GB typically. This was different 10 or so years ago when G1 was still newish and not default. ZGC was introduced with Java 11 (I think), and aimed at very large heaps. It trades off additional overhead for guaranteeing very low latency. That tradeoff is why it is not default. For most users, G1 without any tuning whatsoever should be fine. Generally, if you are stressing your heap, you get more hardware. And if you are not, the GC should be keeping up just fine.
Anyway, like it or not, the JVM has been a work horse for big data for ages. Things like Hadoop, Kafka, Cassandra, Elasticsearch, etc. all run on it and scale fine (and typically without a lot of GC tuning). The only feasible alternative to the jvm used to be things like C++. Lately, Go and Rust are pretty credible in this space as well and both have had to do a bit of catchup in terms of maturity of tooling, libraries, and language features. Things like generics (Go), async (Rust), etc. are still fairly recent additions and both kind of relevant in a project of this type.
In any case, switching languages is hard for teams and these guys have been around for quite some time. When they started, Rust was a lot less mature than it is now and Go was still pretty new as well. Neither was an obvious choice for this stuff at the time.
I am actually really impressed. Well done! Good work!
There's lots of interesting lessons and knowledge in the design of this platform.
I also like how you've decided to use Java as your API rather than Clojure.
I hope you're not discouraged by HN's reaction to your hard work. Don't be discouraged!
ctrl+f "monetization"
ctrl+f "moderation"
ctrl+f "existing infrastructure"
ctrl+f "personalization"
etc etc
Yeah about what I expect from a "we rebuilt twitter for cheap" post. There's no point to the comparisons with the Twitter codebase size/cost. It completely distracts from what is probably a perfectly fine project.
It's a JVM application with all state duplicated N times, so at least on the memory side it's likely going to be a resource hog.
It’s a website builder with lots of themes similar in design.
I see “microbatching” in the diagram and, maybe this isn’t a fair take, but it feels more 2013 than 2023.
I guess most people can't accept things which is fundamentally harder in such architecture than normal ones.
Clever architecture can help as much if not more than clever coding especially when keeping it simple but scalable is needed.
But I won't ever consider investing in it unless it's some form of open-source. It's too much of a risk to have a closed-source core.
EDIT: Oh, I see in comments: "The customer API in Java, and the implementation of that API is in Clojure"
Is this the case — ie. would a TodoMVC app implemented in Rama also be much simpler than a traditional frontend/backend/database CRUD implementation?
X years from now "We reduced the cost of building _____ at Mastodon-scale by 1000x".
It's certainly interesting, certainly an accomplishment, but it's also the nature of the game. The present eating the past, to be eaten by the future. Rinse. Repeat.
Is there any rough infrastructure cost comparison?
Excluding the cost of engineering effort, which I understand is the major pitch.
is it baremetal?
vps?
how about doing a comparison on consumer grade vps like 1 vcpu/4GB ram setup comparison between your product and mastodon or pleroma for example?
i mean sure you can build a twitter scale product but federation means people can do that on their own and with your tech, they dont have to worry about scaling issues.
You built a Mastodon-compatible clone in Spring/Reactor.
Why Java?
It's a Java API so any JVM language can be used (Clojure, Scala, etc.).
Article: "building Mastodon at sub-Twitter-scale"
We've learned that when an article's original title generates complaints like https://news.ycombinator.com/item?id=37137317, the thread is likely to get derailed by shallow arguing about the title. It's in both the author's interest and the community's for us to nip that in the bud by (1) putting an accurate and neutral title at the top (preferably using representative language from the article itself), and (2) marking the title complaint offtopic since it no longer applies. These steps nudge the thread toward discussing the article's content rather than merely its title.
The point isn't the Mastodon instance, but rather that Rama enabled us to build it at scale with in a tiny amount of code and time.
Claiming to have enabled significant scaling of a Mastodon/ActivityPub-compatible instance is fine. Claiming to have replicated Twitter on the cheap is, from the post, not accurate.
All those use cases you listed absolutely can be implemented with Rama, and Rama's extreme cost benefits would apply to those as well.
https://www.washingtonpost.com/technology/2023/07/29/meta-th...
> That's why we're comparing it to the cost of Twitter's original consumer product.
Plus, you cloned a preexisting architecture. FB wrote theirs from scratch. Not apples to apples. This is much easier.
Nono, you can't say that when later on you say it's built on top of Rama. You literally spent 10 years building the framework to even make this.
And yes, you built this in 10k lines of code but how many lines of code is Rama? This seems disingenuous.
I see parallels though to Datomic, where they turned the database inside out, co-located the app logic and data and indexes, etc. There are a bunch of great videos on YT about Datomic by Rich Hickey & co, worth a watch and I think shine a light on the approach here, too.
They had to invent the computer first, and before that they had to create a universe capable of sustaining both life and computers.
And you can't take the "twitter engine" out of twitter and build other apps with it. A lot of it is custom built to fit the twitter data model.
Unlike - it seems - Rama.
The bright side is that it's easy to move to a new server.
> You can begin to understand this by starting with a simple observation: you can describe Mastodon (or Twitter, Reddit, Slack, Gmail, Uber, etc.) in total detail in a matter of hours. It has profiles, follows, timelines, statuses, replies, boosts, hashtags, search, follow suggestions, and so on. It doesn’t take that long to describe all the actions you can take on Mastodon and what those actions do. So the real question you should be asking is: given that software is entirely abstraction and automation, why does it take so long to build something you can describe in hours?
> At its core Rama is a coherent set of abstractions...
This conclusion is alarming to read from a company that's trying to sell a new platform. The vast majority of the work in building Twitter or Reddit is not about building a coherent set of abstractions, it's working with an often incoherent reality, dealing with a myriad of laws that describe, as if your web app were a human clerk at a post office, how to handle PII and credit cards and CSAM filters and audits and copyright claims and on and on...
I'm honestly shocked that the technical implementation of a simplified, coherent platform took a full 9 person-months. That shouldn't be the hard part. What I'd want to know as a prospective customer is how you handle exceptions to your beautiful, idealized architecture, when some foreign country requires that you only store comments posted by their citizens within their borders or something like that.
eta: they say that had it but removed it because apparently it's not something mastodon supports. so I guess it is a pretty good high level implementation.
To be fair they developed this whole new platform to build this app with. I guess that's where the effort went.
> Our implementation is built on top of a new platform called Rama that we at Red Planet Labs have developed over the past 10 years.
You shouldn’t be handling PII/raw CC’s anyways (assuming FinTech is not your core business)
Secretly scanning your customers private messages against an illegal and immoral hash table from a pseudo-government entity? Are you law enforcement? No? Then fucking stop.
Copyright claims? Fuck ‘em. Only do what you are absolutely, positively, no way-out legally bound to do. No more no less. Require formal, written requests and comply in the maximum amount of time allowed.
Audits? What kind of audit? If they’re non-financial you’re probably doing something wrong.
Corporate squares have ruined the tech scene, and it’s time to resist.
That said, as we described in the post our implementation of Mastodon is less code than Mastodon's official implementation. So not only is Rama orders of magnitude more efficient for building applications at scale, it's also much faster for building first versions of an application.
Comparisons to twitter are unfair, twitter is not really technical gem or is it? It's pretty impressive to build it with 3 ppl in 3 months, but hmm also seems feasible using other tech, given all blueprints are out there.
So I would not say it's remotely feasible to do this in less than one person-year with any other technology.
https://www.washingtonpost.com/technology/2023/07/29/meta-th...
https://joinmastodon.org/trademark
removed part about the mastodon subreddit since this is clearly not about the Mastodon software per se.
The top of the page reads "Red Planet Labs", the title of the article is "How we reduced the cost of building Twitter at Twitter-scale by 100x" and the first line of the article is "We built a Twitter-scale Mastodon instance from scratch in only 10k lines of code."
No reasonable person is going to think that this article has anything to do with the official Mastodon software, so there's no trademark issue here.
https://mastodon.redplanetlabs.com/timeline/local
and this from the gGmbH trademark policy:
> You may not use the Mastodon word mark, or any similar mark, in your domain name, unless you have written permission from Mastodon gGmbH.
Otherwise, people could just say you can’t use their trademark in any document that says something negative about them, and then successfully sue the press and angry customers for complaining about them.
Move along, nothing to see here.
+ the time spent creating Rama, the platform that enables it.
Very dishonest leaving that out.
Bold of you to come to HN with the breathless hyperbolic marketing fluff that may work on Twitter...
- "Depots" are event streams (for event sourced data repositories)
- ETL read one or more streams and project them to indexable read models...
- Which read models are called "PStates" and represent nested combinations of indices like hashtables, b-trees, linked lists and so on. The point of those being they have the data in fast to query way.
- And you have query engine which splits a query into 1+ index sub-queries and then aggregates.
Am I missing something, this seems relatively standard event-sourced / CQRS-like architecture, but streamlined to avoid redundancy and reimplementation of common abstractions.
It would've helped if the terms were less obscure than "depots" and "PStates".
The only systems that scale linearly are stateless systems. Mastodon is not stateless. And even stateless systems hit some bottlenecks eventually, as they exist and run in a scale-variant Universe.
So this claim by itself doesn't immediately impress me, just turns my red lights on, awaiting further investigation. But we can of course discuss why this claim is made and how is it supported. The article is long so I've not had the chance to read it entirely yet.
But we have X number of event streams mapped through Y number of ETLs to produce Z number of read model indices, in a shape that seems to form a highly interlinked DAG, which eventually loops back on itself in terms of message flow. Just the increased cross-chatter here as we introduce more features suggests non-linear scaling.
Individually, none of these concepts are new. I’m sure you’ve seen them all before. You may be tempted to dismiss Rama’s programming model as just a combination of event sourcing and materialized views. But what Rama does is integrate and generalize these concepts to such an extent that you can build entire backends end-to-end without any of the impedance mismatches or complexity that characterize and overwhelm existing systems.
You have the general model correct, but here are a few clarifications:
- PStates are partitioned, durable, replicated indexes that are represented as arbitrary combinations of data structures. A PState can be as simple an an integer per partition, or it can be complex like a map of lists of maps of sets. PStates allow you to shape your indexes to perfectly match your application's use cases.
- I wouldn't call Rama queries an "engine", as it's considerably more straightforward in how it works than something like SQL. The base query API is called "paths", which are an imperative way to concisely reach into one partition of one PState to fetch or aggregate values. There's also "query topologies" which are predefined, on-demand distributed computations that can fetch and aggregate data from many partitions of many PStates.
How do you ensure consistency here? How do you organize it in the data flow?
Say I update a user, because that user seems to still be there in the query result/indexes, but actually an event for this user being deleted has happened some time ago?
This can also happen I suppose of the depots run queries themselves on PState in order to determine if a certain event is valid at all or not, and how exactly to carry it out.
- You can finely tune your indexes to be exactly the optimal shape for your application (data structure). You can see this in our Mastodon implementation with the big variety of data structures we used for all the use cases. - You're generally just using regular Java objects everywhere: appending to depots, during ETL processing, and stored in indexes.
How you coordinate data creation with view updates is a deeper topic, so I'll just summarize one of the basic mechanisms Rama provides for coordinating this. Depot appends can have an "ack level" that determines the conditions before Rama tells you that depot append has completed. The default level is "full ack" which includes all streaming topologies colocated with that depot fully processing that record. With this level, when the depot append completes you know that all associated indexes (PStates) have been updated.
There's also "append ack", which only waits for the depot append to be replicated on the depot, and "no ack", which is fire and forget. These all have their uses depending the specific needs of an application.
I work in marketing automation, and I guess I have in one way or another my entire career. The clients who need to use the platform to communicate with their own clients over social networking may never touch our print delivery system, but that doesn't mean that print delivery doesn't exist or isn't important.
If you are unwilling to recreate the totality of the application in terms of functionality, then you are lying if you say that you have recreated it.