Throwing away 10 months of work after 2 months on the job
dancowell.com
dancowell.com
They were using Angular's dependency block them when I don't think they really needed to be blocked to ship a product.
> The relaunch of our ecommerce website had one goal: deliver blazing fast performance with Server-Side Rendering.
you’re suggesting they ship a rewrite that doesn’t do the one thing they’re rewriting for? at the time they also didn’t know when ssr support would be available.
they switched and shipped ssr support faster than it would have been available otherwise.
The reality is rewrites are often appropriate but they need to: 1. Have a clear tangible benefit (or ideally many benefits) 2. Be as tightly scoped as possible to achieve the above benefits. Rewriting everything at once is a horrible idea in all but the most egregious cases. 3. Be as incrementalized as possible.
Re-writes are also needed when the architecture doesn't match the problem space well. This can be incrementally done. Single threaded loops instead of using parallel-safe constructs like tasks and jobs.
It's one thing to rewrite a website, where the project timeline is measured in weeks.
It's another to rewrite large project where the -estimate- time is years.
We have mature desktop systems that have been in continuous development for 25 years. That might blow your mind, but its the same timescale as say Excel of Word. The question of rewriting, regardless of the tech stack, is moot. It would take years of effort (optimists think 5, more likely 10) and would cost a fortune. With negligible benefits beyond "now written in a language with a bigger developer pool".
I'm the first to admit that there are disadvantages to the stack. On the other hand we are where we are (doing well) thanks to that stack, so it's not all bad. We're still using code laid down in 1996, while being a modern desktop app. We're still the leaders in our field.
The time (aka cost) to rewrite is the major hindrance, and far outweighs any issues using an uncommon stack.
Systems that can be rewritten should not be rewritten. Systems that cannot be rewritten should be rewritten.
It's not quite true (not in all circumstances), but it's close. I think most enterprise / pile-of-business-rules software falls into the latter category, those systems cannot be rewritten, but they should be.
And there's many ways to rewrite and migration strategy other than full rewrite.
Too many people just want to remix code or rehome it, taking the papers out of the old book and putting them into a new cover. At the very least, take the time to make sure you fully understand the system before you replace it.
Actually read the source code before a rewrite.
That they gave the project another two months is another important part of the experience.
So the specific project was clearly a failure. They weren't talking about rewriting working code, let alone code that was making money.
To me, it looks like it took a lot to violate the never-start-over-prime-directive-for-progress in creative work.
“From the founders' perspective, they had a clear stop-loss in place if things didn't work out” with the “stop-loss” being that “[t]he React site would be functional and performant in 6 weeks, and ship in 10. If it didn't, I would step down and recruit a replacement.” It’s just such a horrible thing to put on people, and is in no way a real stop loss.
So things failed, and now we lost our chief engineer, and who knows if they will “recruit a replacement” when they are either gone or on their way out, and recruiting leadership can take so long it’s such a crapshoot. THAT is a “clear stop-loss” to this person? Big oof.
It’s just such bad framing and bad ultimatums and bad people all around that I feel like I’m watching trash tv — and I love it.
The statement is wrong-headed in so many ways it invalided anything the author could possibly say after.
that's pretty silly. you can be wrong in some areas while still having good insights in others.
Everything about this particular situation was exceptional. I focused on the decision to do a rewrite in the post because I thought it was the more interesting part of the story. In hindsight I might have gotten that wrong.
Or least, that's what I would think as a manger.
"Wow, you're serious. I believe (a) you're sincere in what you're telling me and (b) you are going to do everything this takes to make it a success."
I'll hook my wagon to that.
I think when most people say "don't do big bang re-writes", the reason behind it is the current development team unlikely has a complete understanding of the codebase. Once you have a million+ lines, no single person has a complete understanding of how everything works. People leave and new hires come up to speed, people forget how things work, etc., so the understanding of the system from an organizational perspective is rather fuzzy. Starting over in this situation is risky because the organization will inevitably have to re-learn forgotten things, re-implement and re-forget them to get to the same state that it was originally in. Much better to do one piece at a time in this situation.
We wouldn’t have been able to hit such an aggressive deadline under different circumstances.
I wonder how long it would have taken them to write with plain HTML served by simple views.
I know everyone is focusing on 168MB, but ...
The bigger issue that needs to be addressed for this leader is - why did the previous project take 10-months, when that same team was able to deliver an entirely new rewrite version in just 7-weeks (without all the code bloat as well).
1 minute is a long time, and I could see if you could get that loop under 10 seconds you could shave off months of dev time. Hell, I just recently changed a test timeout I was working with from 2 seconds to 1 second, and my morale improved in ways that surprised me even knowing all this.
A clean slate on the other hand is a lot easier. There's no vestigial pieces. No dependencies on hastily written functions you never got the time to rewrite the right way. You aren't locked into some version of a library from three years ago, you can use all the latest tooling.The shorter the codebase is in total, the easier it is to grasp the scope of how it all works together and move faster, knowing where other things should fit in a much simpler mental model.
That is an utterly moronic decision.
I find this comment by user "Dad" very endearing.
Maybe it is factually accurate that Angular was the wrong choice for e-commerce (at the time?). I don’t know. Without SSR, was there really nothing feasible they could do to reduce the page load time? Did Angular make tradeoffs that are better for a large, interactive web application; and you want a different set of tradeoffs for ecomm use cases? Or were they just stuck with the wrong tech at the wrong time: complicated enough that it’s too slow in Angular, and the project hadn’t caught up to them yet?
That JS bundle was doing _something_ (probably lots of somethings). How many features did the React rewrite drop? Could you have removed those features from the old codebase, or would that have been politically infeasible at the company? (“no, don’t delete that” vs “yeah, we can live without that, you don’t need to implement it” -> author gets to add a Loss Aversion sidebar). Are there still items in the backlog to re-implement missing features (that the team intends to complete)?
Development effort compared between stop-the-world rewrites (in order to hit their 7 week “really big shopping event”) vs rewrite following ongoing feature development and the painful merges mentioned in the article. Obviously business appetite changed for a stop-the-world rewrite, but could that have been identified earlier?
Changing requirements: did the Angular upgrade start with “we need SSR” as a top requirement? It seems more likely that there were a bunch of benefits / requirements for a technology upgrade, and over the previous 10 months they’d discovered SSR as a blocker. Erasing the other improvements makes the 10 months vs 7 weeks sound very impressive, but seems disingenuous.
Maybe my expectations are off for a blog post format.
It didn't read like that to me, I thought the point was they went for a rewrite on the new breaking changes major version, expecting specific features to eventually land, were waiting for them, and React had them already.
So it's more (to me) like 1) don't do that; 2) if you're going to do that maybe take a step back and there's something else out there already or that's a better fit.
It could have been any technology. The silver bullet is choosing the right tool for the job.
I don’t have an attachment to any particular tech. At the time React was what I knew, and I was coming off the back of building a server side rendered React site when I joined this company. I had a team of JavaScript-focused engineers to work with.
But it’s not a circle, it’s an arc. If it were a circle, running as fast as you can in one direction would bring you back around to the former. Instead we hit the end of the swing and backpedal until we remember why we were going this direction in the first place.
That about sums up everything wrong with our industry. Engagement over usefulness.
the point of an SPA is so that i can share code and data for different pages in the browser without having to get it from the server for each new page.
it does not at all prevent me from using links to the outside or even within the site. (i just built an SPA where all navigation is done via regular links, each view has a different url, they all can be bookmarked, and history works too so you can go back and forth using the standard browser buttons, and links to certain resources take you out of the app)
You still have to get it from the server for each new page, you just do it in the background.
It doesn't. Most users don't actually notice.
It was also the first SPA I made, in a career of decades. I also think I do not regularly use anything that benefits from being an SPA (I don't regularly use that diagram designer either).
Think data analysis tools, Intranet tools, e-commerce, and a host of B2B stuff?
I know, it all starts with a filterable list or a shopping basket, where page reloads or ajax "sprinkles" (regardless of what they're loading, HTML or JSON) are absolutely OK.
Especially the fully old-school approach with the page reloads - it can work and scale for simple and predictable requirements.
Until it doesn't - the first page with 5 interdependent autocomplete inputs takes weeks to develop and fix, the modals can't have routes without piles of jQuery hacks, and so on and ao forth
If a five input form takes week to fix, you have engineering problems beyond any choice of framework…
all the application and interface logic is in the browser. this is only possible thanks to SPA.
of course it is a tradeoff. if the user wanted to save their data outside the browser every few minutes then that would not be practical or if there is to much new data being generated so it doesn't fit into localstorage. i'd either need a server (even if running locally) that can save that data without user intervention or i should not use the browser as a development platform. in such cases a browser based SPA would have been the wrong choice. maybe electron, or a traditional desktop framework
It was fun to write. Relevant code is here:
https://github.com/ldyeax/jimm.horse/blob/master/j/j.php
https://github.com/ldyeax/jimm.horse/blob/cb6a1c03504cbe6b63...
I actually uncovered a niche bug across multiple browsers with this PHP code. In at least Firefox and Edge, if a page is cached with a fetch using accept:text/json, the response will be shown from view-source of the html page you have open, even if the content is different when you fetch the page with accept:text/html
The idea is the same code can work on front end and backend. On NextJS it will run your React on the backend but useEffects will not be called. Just the initial render is saved and piped to the browser.
These days you can slap together a very capable PHP based Ecom solution using Livewire for those sections that need interactivity.
People really need to realise just how detremental to the user experience it is to have mountains of JS being needed just for an ecommerce site.
Unless there is a legacy system driving the requirements in a different direction, my default application architecture is always going to be server-side React at this point.
Or maybe hire people who can code in more than one framework.
You need help if a lot of money to be able to make wagers like this. And I see no reason to risk my personal livelihood over bad corporate decisions.
I guess that's the difference between what I do and the crazy startup world.
I always remember the "let's throw away Windows ME codebase for NT" scenario. Sometimes giving up is better I suppose.
Well the Windows 9x codebase was always mean to be , eventually, thrown away. It existed only to be able to run the OS on very low hardware (eg Win95 run even on 4 MB of RAM).
You're not wrong, but the moment you speed up your site to being usable then the stakeholders above you are going to add 50 more (conflicting?) requirements until the site is slow as the old one.
I lost at lot of faith in the author's judgment at this point. I know this thing worked out and it was the right decision to port things over to React (and correct the bad decision to count on the Angular team to deliver SSR on their timelines). But you need to be able to explain things to the business without this sort of nonsense. I would be really concerned if I heard a manager saying something like that. It's not an argument, it's just a "trust me" that raises the stakes and pressure around the project.
Jumping from one extreme (completely broken development process) to another (threatening to voluntarily resign if arbitrary deadlines aren’t hit) just feels like a sequence of unhealthy extremes driven more by ideology than practicality.
It’s great that the process worked out, but I’d not be happy to work under either team to be honest. Plenty of teams manage to ship working code without either of these problems.
The key point is that the team in this story didn’t. They needed something extreme to make progress.
And remember that nearly all deadlines are arbitrary. Somebody needs to pick some date to see if something gets done by that date. There’s nothing wrong with an “arbitrary” deadline like the one described in the article.
> Somebody needs to pick some date to see if something gets done by that date.
Right. What you don’t need to do is say “I am going to hold my team to the arbitrary date I ballparked because I staked my reputation on hitting it no matter what”. That’s a miserable way to work.
And
“Don't do "big bang" rewrites.”
Cannot be oversold. Seen massive projects flounder under the delusion they could do a full rewrite and base the entire stack on some v0 core service/library/whatever the architect saw at the latest AWS reInvent or whatever.
It's been about eight months since, and the new app is really good. It will be a hundred times better than what we had.
Since we didn't do the MVP thing (which pretty much forces a "sunk cost"), we could do this.
But yeah, when I joined that software had been built for over 8 years, using technology and practices from the 2000s (think AJAX where X is XML, a PHP back-end that concatenaded XML tags, and a Dojo front-end), and my goal was to rebuild it.
I spent 2.5 years there, using Go as back-end so it's all integrated (one issue was that the PHP runtime they had at customers couldn't be updated because RHEL didn't have a newer version in the repos), React as front-end, etc etc. But there was just too much functionality to rebuild, and they didn't / couldn't find other people to hire (until I quit, because of course).
In hindsight, while PHP wasn't really my thing, I should've focused on improving the existing codebase and then only incrementally improve and replace things in-place. I mean at the same time, I knew a full rebuild of that was a Bad Idea, but that's what they wanted and I wanted to do stuff in Go as well so I didn't challenge it too hard.
But yeah, I could've easily made the back-end twice as fast with at worst a week or two of work; half the back-end's processing time and memory usage was because they took that concatenated XML, parsed it, and converted it into JSON. Because at some point the previous author realized JSON is the new cool kid on the block.
The backend is too complicated, because I originally wrote it as a general-purpose server, and it is now specific, but it works fine. No need to rewrite, although I’d love to.
The frontend had been “accreting” features, because it was really a high-quality prototype.
When we had nailed down the functionality, I knew this was a ball of mud that I didn’t want to maintain, so I factored the business logic into a standalone library, and most of the work has been on the “chrome” for the app. It’s not-simple, but much more straightforward than the previous version.
Also, this time, we brought in a professional designer, and have been applying style from the beginning.
One day I ripped out guava, handwrote the parts we were using and got our apk size safely within the method count limit. The project still failed horribly though but for other reasons :)
On the down side, the technologies the team did know were all outside of my expertise. I could sort of muddle through understanding but I was by no means experienced enough to be making any recommendations about it.
The system did need a total reset. If they wanted to stick with the technology they had chosen, they needed a different team, or at very least a different team structure with a couple of genuine experts setting direction and the inexperienced people given tasks appropriate to their abilities. If they wanted to keep the team, they'd need to start over with different technologies.
In the end I wrote up a document of recommendations for the leaders, with a list of problem areas and suggested direction. I never said explicitly to throw it out and start over, but left it to management to figure out the implications.
The custom UI/UX design is bad, would like to replace with default design based on OS. For a lot if the code I’ve written I also know better approaches now. Finally app is written using Xamarin Forms but I feel moving to MAUI will be a net benefit, even if I will encounter hurdles.
Now I know my client will not support me re-writing the app. So I might have to do this work in my own free time and once finished, show the client and then they can choose to buy the source code from me, or not.
I am sure the app will become much more stable and enjoyable for end users, so I believe I can convince the client to switch. There’s a long term task in backlog anyway to move to MAUI.
I think in “The Mythical Man-Month” [0] the author wrote a topic that discuses the idea of writing the first version of a project to throw away. As a learning exercise of sorts. There’s value in this idea.
Another idea presented in the book is the second system effect, meaning the first re-write risks being over-engineered, so that’s something to be aware of.
———
Personally, I think you should wait until there's buy in from the customer before doing the rewrite. Or at least that's what I'd do.
Maybe Avalonia would be a better idea, not sure how hard the switch would be.
cf. STEP 3 "Do it twice."
I'm glad he got to the real issue at the end. Doing a full big bang rewrite with no deliverables for 8 months. A classic waterfall project.
Maybe the react rewrite was the right call at that moment in time, but there were a serious of huge mistakes that led them there.
I've seen the same thing for agile projects too.
That sounds like a pretty decent concession. I wonder if it was ever even a real problem. I'll bet the real problem is that nobody was willing to call it done.
I think I agree.
If I have a choice between buying stuff on Amazon.com and Walmart.com -- and if one of the two gives me sluggish UX everytime I start my user journey. The other competitor site would eventually become my default go to and get all my money (everything lese being equal).
Brands are built at touchpoints -- according to my marketing prof -- and I believe it rings true.
SSR was a solved problem before SPA/huge frameworks became the often unnecessary de facto, standard. Why am I dealing with a backend app server that is deadlocked and exhausted its connection pool causing cascading failure because its (wrongly) making in-band calls to a micro service someone wrote which is a single threaded node app running a virtual dom and spitting out js that shouldn’t be required to serve up the app’s presentation layer, yet now is queuing up an ever increasing buffer of connections from the backend trying to get the page assembled to send to the user who does not care about any of the js so long as when they click shit, it clicks.
Js is a simple and small language. The runtime it executes in is designed for a very specific deployment target and even there it’s at its limits. Why did we let a tiny scripting language become this beast of complexity and obnoxiousness? Why can I put bundle optimization and tree shaking on my resume? Or that I inherited and maintained a not invented here fueled SSR system that would rehydrate Apollo stores and glamorous styles, for a site that would have looked the same and provided the same FE functionality using bootstrap css and js? Why am I adding extra headers to response bodies to debug if SSR is being cached by the backend properly, or if the caches are being busted properly and sweeping down the Russian doll layers at the react component level? Why, if it was known from the start that search crawlers wouldn’t render JavaScript, and as a result wouldn’t render the content you want indexed, did you not use a friggin backend templating engine that every MVC framework has a million options for, and include a damn partial that adds your js for a signup modal and some Apollo gql queries? All this complexity is for what? Bash is a tiny scripting language with very specific deployment targets and handles mostly presentation and in the what, 60 years it’s existed, shipping a shell script has been as simple as scp’ing a file to a server, which is how easily you used to be able to “build” and ship js. If the Linux OS devs didn’t need that kind of complexity, your hot or not site doesn’t either (: Is your job as a front end developer significantly improved from the days of using jquery or some other tiny no nonsense utility lib? I see the js being written and I don’t see a huge quality of life improvement to justify the chaos these systems introduced to the entire stack, build system, CI system, deployment orchestration, bundle optimization, library bloat and security upkeep, dependency hell for package upgrades, CVE shit show and supply chain security. What is your app doing that all of this is needed, for presentation? Modern web apps are enormous, render slow as hell, and often break expectations set in stone 20+ years ago, like the back button doing what it says on the tin, or inspecting a dom element and being able to actually see the content, not off in some store somewhere and bound to some property event that injects it and removes it as needed.
People have just accepted that this is how FE dev is done now, with jr devs knowing only this insanity and not anything else. I try to delude myself to think that this is fine, it’s better, somehow, and yes I realize that modern frameworks are in some ways better, react components are modular, reusable, transpilers are able to do chunking, lazy loading, tree shaking, gzipping, and compile time performance optimization as it writes the file, but why is this needed for a single threaded language primarily doing network IO and string replaces? Nobody on the FE team at that place with this insanity could grok how it worked, and didn’t understand when things rendered with SSR would be different, or how to debug any of at all. They didn’t run SSR locally to confirm, they didn’t use any tests, and trying to force them to model the request flow between backend micro services, the added complexity of any caching solution, to constantly be mindful of not having access to the window element when running in node server side without special consideration for shims or null objects, is not fair to them, but it’s not fair for people familiar with deploying complex backend SOA systems to be force fed a maintenance and stability nightmare either. There has to be a common ground.
I know my rant is exaggerated and overly critical, but god damn y’all, reel it in a bit, will you? I’m hopeful that things get slimmed down and less complex as the JS community corrects for the over correction that caused all of this and got us here, and I see things that tell me it is happening, so hopefully soon my job won’t have to mentally model and reason about these moving pieces that result in static files on a cdn.
If I can ever finish it, my solution to end the need for SSR at the lib level, which isn’t original by any means, was a background queue, and a pool of headless chrome docker containers that handled it all. Send it a url, the page is rendered in real chrome, guaranteeing runtime compatibility, dev tools api calls will execute some js in chrome console to inline styles and prepare the js, and finally spit out a full html page complete with bundled inline styles and JavaScript, along with any rehydration on load events. No need to have special js packages or servers. No need to think about it at all as a FE dev. If it works locally in chrome for you, it’ll be the same thing from SSR. Some middleware in the backend app to route crawler/logged out/whatever you want requests back to the SSR system if not in cache, serve from cache if in cache, or pass down the call chain and let the backend serve up html with client rendered js if it’s a request you don’t want SSR’d. Still complex but the core of it works for any and all websites, not just one with a special SSR js package and config. For shops with many apps and frontends, it could all run through the same chrome ssr render pool, because it’s just a bunch of sandboxes chrome instances.
Note: it’s been a minute since I’ve been in engineering at all (out of work) so things may be much better now. I also am fully aware that the SSR system of insanity I described above was overly complex from the day it was built, and nowadays many frameworks ship with bundled SSR functionality.
I noted this earlier in the thread, but the fact is that any user-facing application with a UI is going to include JS/TS and CSS, and therefore is going to include a frontend build of some kind. Any build included in an application comes with the requisite build orchestration for all environments, so Docker config, build specific config, CI config, etc. If you add a second language to an application, you add a second build, and the cost of setup and maintenance of your application basically doubles. Not only that, you now have to have engineers proficient in multiple languages on your build. Most engineers are really haughty about the language they write and the part of the codebase they prefer to work on, so finding someone who is equally comfortable AND proficient writing python applications and frontend code is extremely hard. Most likely you'll get a Python engineer who will reluctantly write some awful React and complain about the entire time, or vice versa.
So if you come to me and try to add some Python or Ruby or god help you PHP to my single build Node/React application because you think its somehow going to make things simpler, I will simply eject you into the sun.
I'll write an API in any language that makes sense for the performance requirements. But if the app has a UI, absent any legacy requirements that force you to do otherwise, it 100% absolutely has to be entirely in JS or TS or you are shooting yourself in the foot.
https://www.joelonsoftware.com/2000/04/06/things-you-should-...
1. I think there is lot of good evidence in the past two decades where companies got killed by competitors because the competitors had the luxury of starting from scratch with better, fresh ideas. Think all of the innovations in Google Chrome when it first came on the scene, e.g. V8 and it's process model.
2. There have been examples of companies that have essentially "ripped out their guts and gone with something new" over the years - think Apple and all their chip architecture migrations over the years.
3. I've come to the conclusions that some old code bases are really just like quick sand, and they need to be scrapped. The hard question is how to do it. I think the biggest mistake is planning a major "big bang" release. While in some cases that is inevitable, the cases where a big bang is truly required are a lot rarer than one might think. From the description in the article I think there is probably a lot that could have been done beyond "we need to do a complete architectural change to get SSR". The primary problem is page load time. I've seen companies have success by putting up landing pages, which still have valuable info, but also start the download of the large bundle (the whole "it's not the wait time, it's the perceived wait time" that's the problem). Point being I think there are ways to do some gradual improvements while a larger rewrite takes place simultaneously behind the scenes. The worst thing to do is tell the business (or even worse, the public) that our current product is dead-ended but a great new thing is coming out a year later, fingers crossed, which in fairness was a big part of Joel's point.
Oh you want to rework on the routing or tooling or anything like that impacts other parts of the app? Good luck getting 5 teams to sit down and agree.
There were some rough edges, but it only took me a few days of hacking to get it working. Thankfully, the ecosystem is much better today
The other thing I can think of is that small teams or individuals can often outpace larger teams, at least when it comes to pure software development - in practice, a team handles things like 3rd party integrations, customer feedback, stakeholder management, demos, etc as well.
I’m glad you had an easier time than we did!
I've been suggesting codebases cleanups but we only do it when the clients complains about severe bugs (which, hear hear, are due to our known core issues)
Trying to fight one of these off now. We don't need a fancy shit ass UI. We just need stuff that works and is simple. But plenty of one-trick pony developers and architects who build complicated mechanisms out of piles of nodejs and microservices which merely consume time, energy and attention and delivery sweet fuck all ROI.
This is so funny. I feel like in the old days it would be called out without remorse, like an old hand putting down a maimed horse. Nowadays you have to tapdance around the issue to such a degree that you look insane to try, people will just spout jargon longer than a ducks dick and your manager will sucker down.
Also I mock myself regularly :)
kidsthesedaysgetoffmylawn
I tend to just use whatever Debian ships with package-wise when I'm doing something for my own purposes. People laugh at me. Their funeral.
remember the pdp-11 was a 16-bit machine, and there's also c64 lunix, etc.
That has the ability to spiral out of control completely if you don't keep a really sharp eye on what includes what.
Here's buildroot in the browser: https://copy.sh/v86/?profile=buildroot
It's 5MB
If we built the app with the stable branch the bundle size was orders of magnitude smaller: less than 200kb. Still a bit of a chonker, but more reasonable than the ridiculousness the experimental SSR branch spat out.