Startup CTO – Premature Scaling
medium.com
medium.com
This may not be obvious, since those startups are typically "agile" in the buzzword sense.
The cost comes when the universe of possible change is constrained to the point where the right questions aren't being asked because the answer would be "no we can't do that". So the startup enters a dark age in which scientific thought and rational problem solving isn't really relevant. Often at this point the team starts using Scrum to force the business to plan better, when the core problem is that the existing infrastructure is not able to support the needed experimentation fast enough, making it useless.
This can happen at various stages of funding, since sometimes a series A or B can occur without solid product market fit, or the company may be really good at growth in spite of having a promising but mediocre product.
So I view the CTO's job as preserving the ability of the business to behave scientifically and boldly until product market fit is found.
Sometimes you need to put a twist on the problem to make it a bit more interesting if you want to attract smart people.
Great engineers will not stick around if you call it a day and become a feature factory.
Unless your goal is to retain sub-par seat fillers, who show up for a paycheck. Then you can do whatever you want.
Part of that was when I told the team I wanted our splash page and everything on it to be just static HTML files that would be hand edited if and when it was needed. There was no need for fancy CMS tools for a landing page that gets updated maybe once a quarter.
So, that's part of why you need a CTO - to know when to use static HTML!
If that's a minuscule part of the job, then I agree with you here.
My point is, you are not a CTO if the technology behind the product is a static web site. What do your VP and Directors do all day? What about the teams under them?
There is a huge difference between a CTO and a front-end engineering.
Microservices do require more overhead. Additionally, they require you to be familiar enough with you business logic to know where you boundaries lie - changing abstraction boundaries is much harder with microservices than with a monolith.
I think that while they are great for scaling, they are a lot harder to use to successfully get off the ground (though I'm betting that there are places that have successfully done so)
If you're switching languages/stacks just for funsies, you're going to make services more difficult to maintain over time (and just generally slow down development)
As far as languages go, pick one. When you hit a problem that absolutely cannot be reasonably solved under the specifications you care about, consider a different one for that service. Apart from that, there's little value, and lots of overhead, in moving. Getting to write that one Haskell microservice is sure a romantic idea, until that employee leaves and nobody knows Haskell anymore, and suddenly the service is dysfunctional because dependencies have changed. (For non-mission-critical funsies, sure, use whatever, though non-mission-critical apps tend to have a knack for becoming very important for workflows, hah)
This growth journey makes more sense than starting out the door working on your microservices before business logic as this is like putting the cart before the horse.
Instead, people should start splitting up the monolith into sensible modules and libraries. Instead of a "user accounts microservice", start with a "user accounts library" with a well-documented API that stays in the main codebase and is deployed as part of the rest of the app.
Having nice hard boundaries that you can't just penetrate can aid and abet you in debugging.
It also probably helps you with respect to managing security.
As in most things, good judgement is key.
you can scale the base app by giving it more hardware, but when you have to batch process lots of stuff then think of services (most of the time you can just split that to another server).
Yeah. A starting monolith, some cronnish way to run batches, and some way to listen to queues goes a long way.
There's a new alternative: write it in Elixir/Phoenix. You'll get the same productivity benefits with none of the performance trade-offs. Elixir apps are reliable and concurrent by default, and the Phoenix framework has none of the "magic" that Rails has, making it very developer-friendly. :)
This isn't to say that Elixir/Phoenix are perfect. The ecosystem isn't as mature as that of Ruby/Rails, so finding answers to questions can take a bit longer. But that is changing quite quickly.
Fwiw, I've been pretty intrigued by the idea of dotnet on linux lately (integrates with postgresql and everything). dotnet isn't exactly hip, but C# is fast and it's mature (and has some pretty dope functionalish features. You can run F# if you want to, too). Planning to try how well this works in practice soon.
There are things I like a lot about Elixir, but also plenty that I'm not quite a fan of.
There is a case to made to switch ruby/rail for php/laravel. Just as rapid for development and easier scaling plus bigger ecosystem. There is also a case to be made for switching Django for vue or even simple jquery in some cases.
To my mind, services like Lambda are perfect for startups. In addition to scaling down, costing nothing when unused and allowing rapid iteration with zero operational concerns, they largely scale up and, in the unlikely event that the startup is wildly successful, they'll handle almost any load you throw at them until you've had time to re-engineer the backend to be more economical for the increased load.
Microservices sound ideal at first because it's easy to not see that the brunt of the effort is not writing the first version but monitoring, maintaining, and handling when some go down. By using microservices one introduces additional edge cases that don't exist in a "monolith". Quotations around monolith because I believe startups use it with some exaggeration. For example, a Django project of ~10k LoC might be composed of 15 separate Django apps behind one API and I don't see this as monolithic.
So I agree with the article keep it simple and easy to implement, worry about scaling later.
'Because it's cool'
Oy.
"Let's rebuild something based on what the bulk of us already know," I suggested. (Keep in mind what we had was a few weeks of prototype code that was not in production.)
"No, I really want to learn more node," was the only other input, and that 'won' the decision. 18 months later... "yeah, koa isn't really a goof fit for our use cases".
C# is a great language, JVM is great w/ Kotlin/Clojure, both of those have tons of web servers.
And yes, per another point in the replies, someone else should have stepped in to decide differently when there was a tie. I could possibly have pushed harder, but the 'bozo bit' basically was flipped on that project, and much of what I lobbied for was ignored, at least in part because I was the one saying it.
Otherwise, it will grow into a monster that nobody want to deal with.
So... about 2 years later, there's some traction from folks to rebuild it in something else. The latest shiny object is AWS lambda.
It usually only takes me a minute or two to ask them a future function question that's not reasonably attainable (and performant) with a NoSQL solution... so they wind up going back to whatever RDBMS is their standard.
Of course you could solve that by having an ETL process dump your stuff into SQL (we keep our stuff in Cassandra and dump it into SQL for analysis), but that's another ETL you need to keep track of and in our case we use a quite complex data model and the ETL only supports a subset of it.
Another one just off the top of my head is joins. Some data is just inherently better suited in normal form.
Curious if you have an example of this.
It's not really a very nice way to go long term, but can be convenient when prototyping.
I think CouchDB had some real potential, but Mongo is like Couch without the interesting parts. Couch was like indexes on crack. You can make an index that does arbitrary transforms. Sad part was, it didn't have any smarts to pre-cache those "views" in an intelligent way, which to me would be the whole point of that endeavor. I think something like it will come back eventually, but it's still a very research-y idea.
As an aside, what I really wanted from CouchDB was that I could create a view and as I inserted new documents into the DB, CouchDB would automatically generate the new view documents so when I query the view, everything is already precached.
At the time I needed to implement a fire hose listener that would would sip inserts and decide which views with which parameters needed to get called, and then just hit them with REST to do the precaching. It seemed cumbersome, but it sort of made sense why it would be hard for the DB to know what queries on the view were implicated by a new document.
Has CouchBase improved the story there at all? To me that was always the killer promise of CouchDB... Storing documents in a format that made sense at creation time and having arbitrarily complex views on that data... without a cache hiccup. But that "without a cache hiccup" never seemed to quite materialize as a first class feature.
I think this would be 10x more expensive to set up, manage, maintain and use than something like Riak.
Riak looks like it would be overpowered by my needs - but then again, I learn something new every day.
But then again, I could be super mistaken. I'm interested in seeing another side of it.
I am sure that for certain use cases, there are more tailored solutions, but "SQL + keystore" seems like a fairly universal, affordable and versatile stack.
Ignoring the difference for a minute. If you can fit data in Redis, chances are that they can fit in riak and a RDBMS is not needed at all.
Why Riak only then?
It is one of the very few databases that can be clustered. It is sharded, multi-master, highly available, to talk in buzzwords. It is very reliable and very robust, you can add/remove nodes as you wish, replication is highly configurable, etc...
Postgre/Mysql/MariaDB + Redis have no HA. They are both the disaster in disaster recovery.
1. Difficulty in transitioning a system with SQL-Gone-Wild to a system that can be implemented with other dbs without complexity overdrive 2. Escalating costs to get any additional availability guarantees out of SQL (availability & cost are key business issues)
In AWS & Google Cloud, you're going to need to pay outright for capacity as well as capacity devoted db recovery (a replica / set of replicas in another availability zone). On the other side with something like dynamo / cloud datastore, your payment model is based on requests / sec and the costs for replicated data. I'm just saying that there's nothing simple to deciding where price / availability / dev productivity balances will lead you to, and you can be an inexperienced developer who chooses either side of these technologies without really understanding the full story.
If you're inexperienced with relational databases, perhaps a Key-value store with a simple API is much simpler and easier to implement.
I tell every startup that I advise that "revenue solves many problems" because if you have cash flow, everything else becomes easier.. you can buy spot capacity in the short term. You can hire contractors to help refactor. You can rent coworking space for an office.
There are many "good enough" solutions for now. Take advantage of them until you know what your next steps should be.. and then commit accordingly.
E.g. one of your assumptions there are contractors who can help you. That can be true to a certain level, but even product vendors often lack the scale to support their own thing above the basic levels (this is true for post-IPO companies), now imagine this when you can't afford to incubate your product at an expensive tech hub location.
Also, when you have cashflow, your business is under pressure to provide better customer service and marketing, and this is where the things getting tricky: you need to decide whether you want to grow your business or pay back the technical debts. Because the investors priority is always business over tech (rightly), you need to look 1-2 years ahead and make your key technical decisions accordingly.
For instance, I had a client where a very large online retailer wanted a RESTful API into my client's data that they could call directly from their customer's browser. Before the Black Friday Freeze. This was a completely unforeseen requirement and the solution was NOT to scale the client's overall website to handle the 5-10K requests per second but create a micro service on a new stack that could.
Your startup will pivot, sign weird deals and change your target customer. Optimizing for scale early doesn't help, getting to the point where you know you need to pivot does.
That said, there is an important premise presented by the article. If I can reframe it slightly: don't over-complicate your system architecture, especially not needlessly. If, by selecting some fancy and trendy new technology, you will complicate your architecture or development process for unknown or modest performance gains, it's probably wise to postpone such adoption until the need is more palpable.
My advise, again showing my bias, is quite simple: just select foundation technologies that are mature and high-performance. Using high-performance foundation software gives your application a high performance ceiling, which in turn provides breathing room to build your application in a more "brute-force" fashion. With this approach, you can defer performance tuning and optimization for some time, and perhaps indefinitely.
I have seen plenty of small and small-medium businesses operate well with very modest technology platforms. But importantly, they are modest in architecture (few servers, few processes, low system complexity, simple deployments) because they used high-performance platforms (e.g., JVM, Go, so on). On the flip-side, I've seen many small businesses struggling and prematurely adopting complex system architecture because their software platform (e.g., Ruby, PHP) presents a bottleneck early in the business' life, which creates a challenging scale event before the company has even found their market sweet spot.
When I've written about this in the past, I've pointed out that many people call this situation a "good problem to have," which is of course valid in important ways. But it's also crude and demeaning to your development team. You're basically saying, "Fantastic, our business is scaling. Yes, our servers are on fire and my developers are panicking, but that's great!"
It's better, in my opinion, to have used a high-performance foundation and (potentially) avoid that unnecessary pain. All else being equal (e.g., importantly, your development team is comfortable with a JVM language, for example), I recommend keeping your application simple by using platforms that give you a high ceiling to grow.
You and I completely agree on this: huge scale is a irrational target out of the gate for most projects. The potential for huge scale preoccupies many an architect whose aspirations exceed their budget, to say nothing of their market reality. A premature preoccupation with huge scale tends to yield over-engineered systems that are needlessly complex and work against all of the goals you cite: simplicity and the ability to react quickly to changing business needs.
It seems my addition here is reduced to this: All else being equal—if your team is equally comfortable with two technology options—selecting the platform with a higher performance ceiling will further defer the architectural complexity typically associated with scale. Architectural complexity (e.g., multiple languages, lots of orchestrated or inter-dependent processes, many abstraction layers, complex deployments) tends to be the friction force against the simplicity, speed, and agility we both seek.
The GP points at the transition between the quick & dirty architecture and the architecture that handles Google-scale, but those aren't the only good stable points for a system.
[1] https://michaelfeathers.silvrback.com/the-myth-of-scaling
In a cloud environment, it's quite easy to be comfortable that you _could_ scale out to 10-100x of your current capacity, if you have to. But it's probably not worth proving that fact until you have some customers.
Or put another way, being in a very flexible environment makes it even more of a no-brainer to delay building for scale, since you have so many free or nearly-free options there.
Now of course, if you really do hit scale you'll have the fun problem of needing to optimize your spend, but you can probably delay thinking about that until well after you've proven product-market fit.
Honestly curious, I've never seen anyone scale just by "using puppet/chef".
Dumb, often limited scalability is easy now. Elastic Beanstalk approach is only as good as it's the weakest link, which you have no control over.
EC2 is having issues with recovering from standby? Good luck getting anywhere. You can sit on the chat with the useless AWS 'technicians' that are paid to waste your time until the issue is corrected.
A CTO should be able to foresee this and act before scaling becomes a behemoth project and no amount of money can buy time-to-market.
That simply is your opinion. You are more likely to lose customers for not focusing on the core product first.
except when you're selling an idea to a VC and scaling is all they think about.
I highly recommend Eric's Schmidt take on this: https://youtu.be/hcRxFRgNpns?t=838
Spending your own money provides good discipline in pragmatism, I've found.
It's not just the stack, the architecture matters more.
Put priority on paid users vs free, to cover server costs
It does not take more time to design properly from the get go.
That, of course, assumes you have a CTO and a team capable of their job :)
What usually happens is the first thing you build is wrong, and you need to iterate aggressively on not just the software but the entire business concept. Once you have product-market fit, then you are in a position to design an optimal architecture. If you go microservices before that, you are adding unjustifiable overhead, which will inarguably slow you down compared to a monolith when you are in the 1-3 engineer phase, and don't fool yourself about that.
It will all come down the the experience of the team, esp. the leadership - the managers, the leads.
A strong, experienced, mature and effective leadeship will enable a team to perform "magic" and set the right technology foundation that will prove its worth as the business grows over time
Building something quickly that can be broken up over time is the best way to an MVP, and you can deal with scalability issues when you HAVE scalability issues. I consider scalability to be a good problem to have, because by that time you have actual traction.
If you are a bigco with a big marketing budget in a field in which you already have experience and business then by all means scale to millions of users at day one. Don't take on the cost of it in a startup, because the problem is TRACTION not technology with regards to scaling.
Designing and building for scalability, flexibility, availability, etc does not mean going all out to boil the ocean and then find that you attempted to boil the wrong ocean.
If that did happen, you got the wrong team :)
You're making fantastical, hand-wavy claims that belie the hard tradeoffs that need to be made to get a startup off the ground. Microservices absolutely add complexity and require more effort to set up correctly and orchestrate, pretending that experienced leadership makes that problem go away is dangerously idealistic. What experienced leadership gives you is good judgement on whether the overhead is worth it given all the factors at play.
The way you speak about the team, and different roles makes me thing you've maybe worked at "startups" that are actually established businesses with traction which like to retain the mythology and allure of being a startup (perhaps to justify lack of profitability). Things are very different in a true startup where the tiny handful of employees all must wear many different hats.
Build a snappy prototype that doesn't scale then sell it to some clueless guy.