My $work had been in that situation a couple years before I'd started. I've been working on revitalizing/removing some of the C we had for several years, while working alongside people similarly managing the php.
I don't have silver bullets for you, but hopefully you can benefit from my experiences.
> - this code generates more than 20 million dollars a year of revenue
Priority 0: don't fuck this up. Proceed cautiously, with intention. Focus on observability before you make changes. Get some sort of datadog type product, or run something in house.
Start building the culture of understanding risk, mitigating risk by having monitors. Get the other developers on a pager duty rotation, work to get them personally invested in operational excellence.
Get management on board with investing time in it: it's risk mitigation for their business. Get any incidents in front of them. Explain how and why it happened, what lead up to it, and things you're considering doing to remediate. Track how much time winds up getting spent there, and use that as an argument to proactively fix things. Most management will understand that if you're getting randomized, you're not being able to make progress on any single issue.
Work on getting a docker compose setup going so you can easily create a dev environment that looks exactly like production.
Use that to start creating black box tests. Consider things like selenium or postman. Your goal is to test as though you were your user and have no clue about the internals of the program. You do this so that when you make changes, you're not having to update tests as well. Write the tests first. Think in terms of TDD given/when/then. As you add new code, write unit (and integration) tests. Don't try to unit test existing code unless it's very simple.
> it runs on PHP
I feel the pain. Part of the engineering challenge here is accepting unfortunate initial conditions. Your goal is to raise the bar to sustainable.
> it has been developed for 12 years directly on production with no source control ( hello index-new_2021-test-john_v2.php )
Priority 1: get this in source control. If you can't get the other devs immediately onboard, copy what's in production to your local machine and start a git repository. If you need to copy down files after they've changed them in production, that's ok, just start building the repository and tracking some history.
After rsyncing changes down, you'll be able to diff with your latest checkout to see what changes have been made.
This is another mentoring opportunity. Show the other devs how using git is making your life easier. Show them how it's helping you manage the risk that they're afraid of.
Ideally, get a gitlab or github account for it, and start getting a CI pipeline going. Proceed slowly here and make sure you build consent from everyone. Maybe start with a private gitlab account and again, show the other devs how it's saving you time.
The first iterations may just be starting a Dockerfile to recreate the production environment.
> - it doesn't use composer or any dependency management. It's all require_once.
> - it doesn't use any framework
The silver lining here is that it means you don't have any external dependencies :) My biggest concern here would be:
- is it using pdo/mysqli, or is it on the legacy mysql extension?
- is it using parameterized queries, or are you going to need to audit for sql injections.
Chip away at this over time. It's not urgent. I'm sure many HN heads may explode at that thought -- but until you've triaged everything, everything seems urgent. You've got unmet basic prerequisites here. Say you start fixing sql injections before observability -- how do you know you haven't accidentally broken some page?
> - no caching ( but there is memcached but only used for sessions ...)
Nothing to fix! wonderful! You can figure out a good caching strategy after everything else is under control
> - the routing is managed exclusively as rewrites in NGInX ( the NGInX config is around 10,000 lines )
Having it centralized is actually a bit of a blessing. It means you're not having to scour the application for where it's being routed.
Start collecting nginx access logs, and getting metrics on what the top K endpoints are. Focus on those. Configure it to have a slow request log as well as an error log.
Do yourself a favor and setup the access log to use tabs to delimit fields. It'll make awking it, or pulling it into a database for querying much easier.
> - no code has ever been deleted. Things are just added . I gather the reason for that is because it was developed on production directly and deleting things is too risky.
The silver lining is that the unused code is inert. This is another "chip away with time" type task. start finding paths that haven't had requests in N months. Use analysis tools to show that something isn't ever called. When someone starts with the "well, what if...", remind them that it's in the repository, and isn't gone forever. It's just a revert away.
A bit theme here is fear. You need to start instilling confidence and resiliency in the team.
> - the database structure is the same mess, no migrations, etc... When adding a column, because of the volume of data, they add a new table with a join.
This isn't the worst thing in the world. It's also not urgent. Start putting together ERD diagrams, get the schema in source control, get a docker image going so that you can easily stand up a test database in a known state, nuke it, and start over.
> - JS and CSS is the same. Multiple versions of jQuery fighting each other depending on which page you are or even on the same page.
Slowly work on normalizing the jquery version. Identify all the different versions used, where they are, and make a list. Chip away at the list.
> no MVC pattern of course, or whatever pattern. No templating library. It's PHP 2003 style.
Not the end of the world -- this is pretty low on the priority list. Both are luxuries, and you're in Sparta. Start identifying the domain models, define POD classes for them, start moving the CRUD functions near by. The crud functions can just take the POD classes and a database connection.
> In many places I see controllers like files making curl requests to its own rest API (via domain name, not localhost) doing oauth authorizations, etc... Just to get the menu items or list of products...
Same as with the domain model and jquery, make a list, chip away over time. Be sure the curl calls have timeouts. Slowly replace the self-http-requests with library calls. Explain how if you only have N request workers, if all of those N requests are then making subrequests that there wont be any workers available to serve them, and they'll fail.
> - team is 3 people, quite junior. One backend, one front, one iOS/android. Resistance to change is huge.
This is a bit of a social problem. 4 people can be very effective though, if you're all working together well. Get to the root of why they're resistant and fix that. Are they just set in their ways? Afraid?
Work with them to rank their top 3 challenges, and work through what solutions may be.
> - productivity is abysmal which is understandable. The mess is just too huge to be able to build anything.
> This business unit has a pretty aggressive roadmap as management and HQ has no real understanding of these blockers. And post COVID, budget is really tight.
Measure, so management gets some visibility. Push back on work if you don't understand it. Be clear about what a "definition of ready" and "definition of done" is.
Don't stop the world to fix things. Consider having one person working on a fixup project while everyone else gives them cover by taking on the management ask.
> I know a full rewrite is necessary, but how to balance it?
I can't emphasize it enough, do not rewrite -- resist the urge, if you can't, find a different job. One of the challenges here is integrating with respect to time.
If you have no observability, and no tests, how are you supposed to even show that your rewrite behaves correctly? And if your team is so afraid of breaking something to the point of never deleting code, how do you expect them to handle deleting all the code?
That's about all I've got in me. I hope you're able to implement some meaningful change. Take it one day at a time, and just try to make it better than it was the day before. Good luck :)