You don’t need to be “enterprise-ready” or “scalable”
gorelay.co
gorelay.co
I believe having a solid development environment where you can develop and test the whole stack is key to making reliable software. Is this common at other companies that you cannot run the whole app locally to develop and test against?
Yes. I work in B2B banking software, so having an "authentic" environment to test against is extraordinarily complex. Even the big vendors in the space seem unable to maintain accurate facsimiles of production with their own decades-old core software, so smaller vendors stand no chance unless they adjust their entire strategy.
What we ultimately had to do was negotiate a way to test in the production environment in such a way that all the activity could be tracked and backed out afterwards. A small group of important customer stakeholders were given access to this pre production environment for purposes of validating things.
This is something that has never sat well with me. But, the way the business works the alternatives are usually worse (if you want to survive long-term, that is). We wasted 18 months waiting for a customer to have a test environment installed before we learned our lesson about how bad those environments are in reality.
Running automated integration tests on the whole env is just an added step in the pipeline.
Another way is creating separate code libraries to abstract away external services. For instance, you might write ObjectStorage library that has real integration tests that use S3 but in your application you just stub out the ObjectStorage library. A lot of frameworks have verifying doubles or "magic mocks" or similar that will stub the method signatures so that can help make sure you're calling methods correctly (but wouldn't help catch data/logic issues).
Live upstreams are also prone to inconsistent state and race conditions--especially when you try to run more than 1 instance of the test suite at once.
For web services, I also like using "fixtures". You just manually run against a real service and copy the requests to a file (removing any sensitive data) than have your stub load the data from the file and return it. If you hit any bugs/regressions, you can just copy the results that caused them into a second file and add another test case.
I think that seems sensible, but how is that different to 10 containers ? Is the issue something about simpler network management inside a container?
They must have some special code that says that these accounts can have their fees zeroed and probably are subject to some special out of band audits. I don't really know the details of course but I was actually kind of impressed.
(no she is not actually a customer of her employer)
I wonder if other developers are familiar with this approach?
In the beginning (more than 20 years ago), I insisted that we implement test systems ourselves so we didn’t have to test in production; that was naive; it was expensive and there is so much complexity that this does not work well.
Heck, when I am in a place with bad internet, I spin up a Cloud 9 (Linux) or Amazon Workspace account (Windows) and do everything remotely.
I’m not wasting my time trying to use LocalStack or SAM local.
If you’re developing locally to none public endpoints, use a VPN or develop within an IDE hosted on AWS behind a VPC.
Why does the app you’re testing run differently in a container than outside a container?
Standard disclaimer to expose my biases not to do the whole “appeal to authority thing”: I work in ProServe at AWS (cloud app development consultant). But I did the same thing when I was working in the real world. My specific job looks just like a senior enterprise dev + a shit ton of yaml, HCL, draw io diagrams and PowerPoint slides.
The other bias I have is that I work for the only company where I never have to worry about an AWS bill and I can spin up ridiculous remote EC2 based development environments when needed.
But I also didn’t have many constraints when I was the de facto “cloud architect” at a 60 person startup.
That being said, how do you get on different versions if everyone uses the same package manager config file (requirements.txt/package.json/the wonderful system that Go uses/whatever Nuget/C# uses).
I’m also spoiled because I’ve never been in a position where I wasn’t the developer and leading the “DevOps” initiatives since working with the cloud - not bragging, I only opened the AWS console four years ago for the first time.
HN/Reddit is a good place to find out other peoples experience.
Docker also fixes the "I forgot to brew install a specific version" in this case.
Either you wind up an instance for each container, it takes too long to start the app with less resources or the image size is in dozens of GB, everything could be avoided.
Even if you do manage to get it all running, getting the multitude of database with useable data adds an odyssey to your day.
We have some parts automated, but ops cannot keep up with all the changes in the micro apps.
Back when I worked on a mono-rails app with a React front end, getting things up and running was as simple as running the JS server, rake db seed and then rails s. We used puppeteer to test FE, and rspec for the BE. Two test suites, but we knew if something was wrong.
That same place sometimes had a net outage, and I wouldn't even notice because I was running everything locally.
Now I can only get everything running locally with confidence, if I block out an hour or so.
Honestly... microservices are great in theory, but make dev a huge pain. Docker does not mitigate the pain.
Some of these when I took over there wasn’t even version control, and they were being developed by FTPing files to production server or even editing them live with vim or something. Certainly no reproducible containerization or even VMs to speak of.
I agree completely with you about having a solid development environment. I’ve spent the better part of a year creating such a thing for some of the brands, and it has increased our time to ship features and fixes probably ten fold.
Many editors / IDEs (e.g. Notepad++) can directly connect to the FTP server, so working on a FTP servers looks almost same as working on a local folder. You just make changes to the files, and then refresh the web browser to see the changes.
Version control at least is essential for something of any size with any number of people working on it. You must be able to "revert" a set of changes quickly and reliably if the application has any importance at all.
Don't want to be rude here, but version control is clearly not "required" in the strict sense. Plenty of software has been developed without it. Now I wouldn't want to work at a place without version control either, but it is as required for developing software as seat belts are required for driving a car. Less even, since at least for seat belts they are required by law.
I got hired on at megacorp to do pretty much that with a suite of green field apps. And so I did, across a few apps and a zillion different environments because all of a sudden spinning up a new integrated environment was easy. The devs ran a lot of stuff locally but still had cloud based playgrounds to use before they hit the staging environments. Perhaps the best part though was having the CI stuff hooked into appropriate integrations, so even if you couldn't run it locally you sure as shit tested it before it hit staging. SOC2 compliance was painful for the company but not so much for us.
If you're sensing a theme with "green field" you'd be on to something. Even in a largeish company it's not so much the complexity as it is the legacy stuff that'll get you. Some legacy stuff will just never be flexible (e.g. the 32-bit Windows crap we begged Amazon to let us keep running) and even the most flexible software will have enough inertia to make change painful.
Tangentially this is also why I get apprehensive when I see stuff like that linux-only redis competitor.
At my company we don't support local dev nor have any sort of staging environment. We have development linux servers connected to prod services and dbs. There are no pragmas or build flags to perform things differently in "dev". Things are gated with feature flags.
It scared me at first but now I think it makes sense: staging is always a pain and doing it this way we avoid parity problems between staging and prod. Local development would be impossible for a system at our scale but I think even a staging setup would result in more defects—not fewer.
Also your local dev should try and mirror prod as much as possible.
The only reason I don't run Linux is due to other MS tech - Teams, Outlook, and primarily office products.
Their solution was to have an anonymized copy of prod to a few (like 3 for 20 engineers) environment where features could be tested. Engineers usually paired in order use them concurrently for unrelated features and they were very expensive.
I did not live enough to see the on-demand environments arrive - devops was always busy with operations tickets to continue building those.
In my experience, recreating staging as "production junior with caveats" comes down to cost and living with bad architecture/design. I'm not defending the practice, but I do think it's incredibly common.
Usually someone develops a remote or mixed local/remote development environment which is easier to use than configuring the full stack to run locally, and then the ability to ever do so again atrophies until nobody can figure out how anymore.
In some companies I've worked at this is fine and not a major setback. In others, it's been a mess. I don't think that "can run the stack locally" is a good litmus test for developer experience - the overall development experience is a lot more nuanced than that. "How long does it take to get going again if you replace your computer" and "what's the turnaround time to alter a single line of code" are usually better questions.
This is very useful if the production environment is hard to replicate on each developer's local machine. For example, a team where some people prefer the MacBook Air and others prefer beefy Windows desktops might set up a cluster of Linux servers for shared development and testing. A common alternative is to tell everyone to run a bunch of VMs/containers locally, but this isn't always feasible depending on the size and complexity of the production cluster.
So basically: unit test fixture setup spanning multiple external services ("Chains?") in a scriptable build, with various common/comprehensive/programmable interaction flows, and then integrating with CI? I'm making some guesses here as I'm not very familiar with the space.
I think it's amazing that you have (had) made this into a career. I am a dev and spent a couple years as a QA developer. Do you think this is a space I could break into? (Decade plus of development experience plus the good fortune of growing up with a mother who was a software developer. I learned to code (badly) when I was 13. I'm more than double that age now :)
If so, I'm wondering if you'd be up for chatting some more like on IM (Google Chat or Whatsapp or whatever) to elaborate further if you think this is still a viable proposition for someone new. I'd be happy to pay you for your time if you think this could make me money. I know how to set up CI, I'm vaguely familiar with bitcoin (I used to own some), and I don't care for cryptocurrency but I'm happy to meet market needs with a smile to pay the bills and then some.
Do you think there's any space for low hanging fruit simple auditing of smart contracts? Maybe a linter?
EDIT: Here's what Ethereum suggests: https://consensys.github.io/smart-contract-best-practices/se...
In particular,
https://github.com/protofire/solhint/blob/master/docs/rules....
https://github.com/protofire/solhint/blob/master/docs/rules....
https://github.com/prettier-solidity/prettier-plugin-solidit...
The architecture was probably more complicated than it needed to be but it was also decades old.
Microservices based architecture with massive amounts of microservices and other service dependencies, its a pain to do more than local unit testing and mocks.
Are you using FaaS or a lot of other cloud features? Also difficult to fully test locally. Extract as much logic out of the cloud dependant code.
Often there will be sub-sections of large apps that can be run locally.
And if the only way to test is in production hopefully they practice some form of limited rollout so you can control who adopts the changes or similarly some form of optional feature flags that are used to enable the new behavior until it can become a new default. If their release model is "We are so good it goes live for everyone all at once" then that tends to make for a stressful release days.
For microservices if you use something like Consul for service discovery within those microservices then you can hit them with, eg, Postman and change your local Consul registrations so everything goes to a main environment except the handful of services you're working on - which you redirect to your local machine for debugging.
It depends what you mean by the whole app. To reflect locally 100% of the conditions the big microservice app I am working on is very hard as I have to replicate locally multiple Kubernetes environments, Kibana, Data Dog, Elasticsearch, Consul and several databases and data stores.
What I can do, is to run one or more microservices locally, connected to remote development databases and speaking to other microservices and apps from some development Kubernetes environments.
Extremely. We can't replicate our entire infrastructure in development - it's simply too complex and expensive. We use a combination of dev, integration, staging, and canary environments plus pilot environments to gauge quality before a full worldwide rollout.
Right now I’m trying to understand and deploy Kubernetes. Scalability and availability is good, but one of the reasons that I want to make it possible to roll out entire stack locally.
Basically it’s lack of devops. We don’t have separate role and developers do the minimal job of keeping production running well enough, they got other tasks to do. There should be someone working on improving those things full time.
This is not even sarcasm.
Often it's due to complexity. Systems accrete stuff (features, subsystems, dependencies, whatever) over time due to changing business requirements. The more this happens, the harder it becomes to integrate them properly without breaking other parts of the system. And the system is running the business, often 24 hours a day, which makes migrations hard to do. So you end up of something thats basically too complex to run locally.
And particularly in regulated industries like banking you have the security teams locking everything down hard. Problems accessing systems and getting realistic test data, etc.
What in the actual fsck? Check out this list of companies ranked by the number of employees they have - https://companiesmarketcap.com/largest-companies-by-number-o.... Okay, Amazon is #2 - but notice the top 100 doesn't include Apple. There's a lot more to business than FAANG.
If a company has VPs buying things surely they have some sort of internal vendor management that usually involves accounting/legal/compliance reviews. Being enterprise ready means you have the roles/documentation/process/procedure to answer those questions
A major bank employs between 50,000-10,000 software engineers and in the 250,000-50,000 employee range. Netflix has only 11,000 employees of all types.
I knew a guy who sold a $2M solution to track MSDS sheets in a prison system. Stupid simple application.
Now I don't know how much of the software on my TC26 actually comes from Zebra but, pretty sure, it's not zero.
ASSA ABLOY software is probably everywhere there's a door these days.
I don't know who the players in HVAC are but considering that a lot of HVAC/lighting control for like Target/Walmart are remotely controlled there's definitely somebody with tens of thousands of deployments.
Not to mention EdTech shit like BlackBoard, that gets deployed to every god damned school in the country
x/|x|
Ya know, like normalizing a vector. Of course this brief confusion doesn't matter, I sort it out quickly, and it is totally clear language to other database people -- it is just a funny quirk of overloaded lingo.
But I wonder if as big data and machine learning folks (since there's some overlap with Linear Algebra there) ever get their lingo mixed up.
"To predict where the ridesharing customers will want to go, we will apply the taxicab norm to our database."
I would also point out that normalization within mathematics does not have a single concrete definition either, and it depends on your domain and desired property to stabilize: https://en.wikipedia.org/wiki/Normalization_(statistics)
Nosql databases excel at simple lookups and gets by ID, rather than trying to sort and join pretty much ad hoc.
Ex Mongo sorts just fine, and can optimize sorting with indexes.
So decades of PHP+MySQL and Rails+PostgreSQL were simply misguided?
To the extent that transactions were the point, yes (though there may not have been many alternatvies for a datastore to use when Rails was getting started).
In the case of PHP+MySQL of course MySQL transactions were disabled by default in that era (up until 2010 in fact).
After all, what do you do when a transaction fails? Most webapps either retry or fail (if they even bother to handle that case at all), but you don't need transactions to do either of those. In an actual database application you can tell the user their transaction didn't commit and ask them to deal with it, but again you can't really do that over the HTTP request/response cycle (because what if the user submits the form and walks away? They expect their input to be saved, whatever happens on your backend).
But the above instance sounds like a variant of the "because, bwah, SQL is hard" line of thought that seems to have powered much of the (now historical) NoSQL boom.
We are having a lot more nervous takes around legacy LDAP/AD auth mechanisms these days.
It’s an API for easily adding SAML, SCIM, and more. Currently being used in production by Vercel, Planetscale, Webflow, and +200 other apps.
E.g I know nothing about Azure AD but I just emailed the WorkOS setup link to our customer’s IT team, they followed the instructions and it just worked. No back and forth.
Can’t recommend enough.
You probably have to prove that you have these on your roadmap and that you have engineering team that can implement these in not so distant future. That you know these things exist and have will to invest in it in future.
The alternative is that your product becomes a security exception, which infosec teams are more and more unwilling to grant, for pretty good reasons.
The good news for product teams is that there are more and more off-the-shelf options that can be quickly integrated vs. having to start from scratch.
But the main point of this comment is: I think you’re right that “near term roadmap” was good enough a few years ago, but less so now, and continuing to trend towards “not acceptable” if selling to enterprise customers.
(Observations as a recent/former auth PM at a SaaS/PaaS that sold primarily to enterprise customers).
History is littered with dead startups that designed for scale before they had enough usage to justify it. Within reason, having users knock your site over resulting in failures like the Twitter Fail Whale is a good problem to have.
With that said, you need to be prepared to scale up quickly once you have this problem. There's a reason Facebook counts its users in the billions, and Friendster is a footnote in internet history.
See I interpret that to mean that you should design for scale, but not implement or invest in scale until needed.
Instead, todays shops all start out with javascript front ends and 40 layers of backend complexity to accomplish the same thing.
However, we do very much care about our UI - and I think react at this point helps with that. But I don't think React complicates the stack too much. It may even make things easier, tbh - I don't have to test pages on backend tech, just endpoints.
I don't think that it's stupidly hard to get basic things set up for scaling at the start. These don't apply to everyone but:
- use a cloud provider, they let you grow
- write infrastructure as code (terraform, pulumi, cdk), it's not that hard and it makes building your 2nd, 3rd data center easier
- for the love of God think for 10 minutes about a scalable naming convention for data centers, infrastructure components, subdomains, IP ranges etc. This type of stuff causes amazing headaches if you get it wrong, and all you have to do is consider it at the start. Many AWS resources can't be renamed non-destructively, and are globally namespaced. Be aware of this. Don't paint yourself into a really stupid corner.
- even if you're not doing microservices, consider containers and an orchestrator that has scalability baked in (like Kubernetes). If you use a managed Kubernetes with managed node groups or fargate, you can basically forget about compute infrastructure from this point on
- deploy Prometheus, Grafana, Loki etc day one and build basic dashboards. With Kubernetes, you can get this installed quickly and within a day you'll at least be able to graph request/error count and parse logs.
- deploy an environment for developers to deploy to and use the product (even load test it)
The development cycle is now set up for building what you need to build. You can easily find developers at any level who have experience with containers, Kubernetes, Grafana so they can onboard quickly, be productive and troubleshoot their own stuff. They can refactor the monolith without having to invent service discovery, load balancing, ingress etc.
My personal experience with Kubernetes began at a start-up with a tiny team, none of whom had prior experience with it. We had started building a single monolithic app. But within a short time we figured out how to make a helm chart, add ingress and boom, we were "online". Every time someone was about to spend time and effort inventing something, we would quickly scan through the Kubernetes docs and CKAD material to see if we could avoid having to build it at all. We leverage cloud stuff too, where it was cheap and easy and scalable (like AWS SQS/SES/SNS/CodeArtifact/ECR/Parameter Store)
thats the whole point of engineering, picking/building the right solutions for the problems and context
some stuff just needs a vps, a php file and an sqlite db and it will scale
other stuff might need container orchestration or anywhere in-between (or beyond). it should depend on the context though and we should be prepared to change strategy as needed
I get that some people can just run a single go binary or whatever. The ones that don't, need to set themselves up to grow easily when they need to. Having some sort of basic plan or framework that doesn't take more than a week to set up is due diligence, and hiring (when your need to) will be much easier.
It's a little like telling a construction worker "You don't need a super-powerful 'jackhammer' to do your work. A simple claw hammer is all you actually need."
Well... sure, that's definitely true if you're building a house.
But you might have a problem ripping up a sidewalk...
There are plenty of companies building houses, and plenty who are working with concrete! Neither one of them are "the big kids."
They're not even necessarily different-size markets, either -- you can make a lot of money building houses or making sidewalks -- it's just that different problems require different solutions, and different software companies are going to have different challenges to the problem of "get bigger."
Granted all servers are C++ and talking to a local DB. No problems in processing thousands requests per second without breaking much sweat.
Upon startup all data (except couple of giant tables) is sucked from DB into RAM into highly efficient data structures hence most reading requests are handled in microseconds. Those 2 giant tables are loaded into the same structs but partially. This is usually enough as RAM keeps last few years worth of data and requests that need something older come just couple of times per month. Only writes touch the database and immediately reflected in RAM as well.
This is just an example. I write game like servers as well and those use different approach. Generally I treat each project / product individually.
BTW for others considering this approach, the DB will do similar things that FpUser is doing in the app layer - keeping hot/recent stuff in memory. But when tail latency and cost really matters this is a good technique.
Not really as the storage data layout and application data layout are different. If I were to use database only relying on it's caching abilities it would lead to complex queries that are definitely way less performant than getting data from in memory structures optimized for this particular application / usage patterns.
On top of that some requests require some relatively complex calculations that again are executed against optimized in memory layout and would suffer greatly if done against the database.
I did some experiments and the difference was very significant.
Sure it will not do Google scale and I do not loose my sleep over it. For most "normal" businesses what I have is way more than enough.
Besides, as I've already said I treat each product individually so my architecture reflects the actual needs and changes accordingly.
What is glaringly obvious from this article is that "enterprise" in this industry has lost much of its meaning. It's almost ubiquitous with "medium+", which I think is a self-defeating definition that starts to make sense only when you look at what features are commonly gatekept behind enterprise distributions of software.
Meta? This person wrote a whole article about it without knowing what enterprise software actually is.
Rearchitecting from unscalable to scalable data storage is notoriously difficult and expensive. Even the most famously competent companies and teams have struggled with that transition and invested millions and millions of dollars on it. Building on unscalable data storage when scalable data storage is readily available feels like planning for failure and making a huge bet that your product or system will never have widespread use.
Spanner and Vitess existed at Google in the early 2010s. It was mid-to-late 2010s before CockroachDB, TiDB, and Vitess were available as open source.
Battle hardened database administrators who've been working with traditional unscalable databases for decades often aren't up-to-date with what's possible today unless they frequently talk shop with peers at other companies or go to conferences.
Personally I hope old databases like MySQL and PostgreSQL die out within the next decade. They served us well, but they're 25+ years old and the weight of legacy code plus the need for backwards compatibility is a huge drag on their ability to evolve.
We didn’t have a “change email address” flow for two years. Instead we focused on the product. That’s what buyers really care about.
You need this stuff eventually - SAML included — but not before you launch and refine your product.
Also compared to the 00s, the 2010s were a decade of gold plating and planning for scales you don't need.
Instead, I found an argument about not building v1 with unnecessary optimizations.
Which I agree with but I didn't find anything new being said here.
What I've found _is_ worth investing in early on is measurement and metrics. No more blind changes, go look at the data and draw testable theories out of it. Huge quality of life improvement.
I wrote a cut down open sourced version which is far less advanced than the one I introduced to the DevOps team I was on.
HTTPS://GitHub.com/samsquire/platform-up
Paired with keepass2dir and secrets-login secrets-logout you can handle secrets too.
HTTPS://GitHub.com/samsquire/keepasscsv2dir
The expectation of the word scalability changes based on who you took ot: concurrent users, througput of data managed, data stored, processing time, system maintenance, feature extension, you name it. No system checks all the boxes 100% while still meeting the budget defined by a company.
> "A complex system that works is invariably found to have evolved from a simple system that works. The inverse proposition also appears to be true: A complex system designed from scratch never works and cannot be made to work."
--John Gall, Systemantics (1975)
Perhaps your application does need to worry about scale and then if you don't then you're a fool. Then you launch, succeed and collapse or get beaten by competition because the effort to scale was beyond you in the needed time.
The blanket statement that is true about software is "be wary of blanket statements about software development".
It's just insecurity and pure coping. They're not dealing with complex performance bottlenecks, so they say these efforts are wasted.
Living that nightmare right now.
Like Kubernetes cluster autoscaler and scaling individual deployments.
I want to be capable of running the same code on a 100KiB file as a 10000 petabyte datastore.
As it stands you need to write your application differently.
Usually where SAAS startups get into trouble is actually the UX and product market fit. There are a lot of successful SAAS companies with absolutely terrible UX. The reason they are successful is that the UX does not matter if what the software does is so valuable that companies want to buy it anyway. So, obsessing about a UX for customers that don't actually care too much may be a mistake. E.g. investing in native apps and tripling the number of frontend engineers you need for that is usually a mistake. Get the software in the hands of your customers quickly and learn whether it lives up to your sales pitch fast. Simple web UI and a simple backend with some conservative tech choices goes a long way. And that stuff is easy to scale.
A lot of SAAS software is bought by people who don't actually use the software themselves. These people buy based on a story and a promise; not the actual UX. There's this curious dynamic where you e.g. sell to the VP of sales who then makes his sales team use the software. The higher the price, the less the UX matters. That's why a lot of SAAS software is absolutely miserable to use: it's designed to make the boss happy, not users. This is also what represents a big opportunity: a lot of that software is not very well defended and can be disrupted with better products. And this happens all the time. But then those companies also end up selling to the same C-level executives and that software becomes less usable over time. Salesforce is the classic example. I've met sales people that hate that stuff with a passion but have to use it anyway. But they started out by disrupting that space.
Ironically, UX matters more if you don't have a remarkable product that is easy to copy. These products depend on easy purchase decisions by people lower in the org chart. Slack is the classic example that showed up in many companies because they had a freemium layer. And then those companies decided to pay up because they had a growing number of people using it.
So, that kind of product actually needs to scale early because it depends on a lot of users not paying anything before enough accounts start converting to paid accounts. Slack only managed that because they had lots of funding. Many companies would have run out of steam long before they could have succeeded.
Those users are going to be a lot more critical on the UX, and it's responsiveness and if the experience sucks, they'll be gone in no time and use something shinier. So, such startups will never break even unless they learn to scale before their runway runs out. For every cute little SAAS app that made it, there are hundreds of me too apps that never came even close to breaking even. Fail fast causes lots of companies to fail. Fast. Unless they can scale and grow. This is why investors are weary of B2C as well. Same scaling issues and it's a lot harder to get consumers to pay. And the UX is critical.
More traditional SAAS companies can afford to close a few big deals with a mediocre product before they have to scale. Mostly these companies keep on selling to bigger and bigger companies without ever needing to address their UX issues.