As with software, you cannot start with an enterprise-grade, architecture astronaut-level system. Start really, really small, and scale your engineering process slowly over time with a focus on adapting to new situations and addressing the problems that you're having now/today, rather than problems that you anticipate having down the road. The biggest problem that I see in organizations is that they start with something really, really huge like SAFe or whatever when individual teams aren't able to deliver reliably in the first place. Start with the fundamentals, master them, and then build on it.
Here's what has worked for me in my current position and in the past:
* Finalize and clear all work in progress. This needs to include anything that's in QA and/or not yet deployed. Get everything that's currently in progress out to production.
* Give your team a day or two off. Try to make this a Friday and/or a Monday. Don't bother them for anything over the weekend. Let them recharge after coming off of what sounds like a pretty bursty work schedule + getting all of the work in progress resolved.
* Set your sprint length to two days. Yes, just two. It's temporary. Read on.
* Clear the schedules for everyone on your team. And not just "I'll try to rearrange these meetings": as the person responsible for this teams performance, you need to set expectations with the rest of the business that the developers will be unavailable for direct support, meetings, etc until further notice and that all queries should come to you instead. If one of those queries really requires developer input, create a ticket and plan accordingly during the next sprint planning. The only thing that should be on the calendar for any developer on your team for the next month is a 30 minute sprint planning on sprint day 1 first thing in the morning, a 10 minute standup on day 2 first thing in the morning, and a 60 minute retro on day 2 at the end of the day. In longer sprints, only the standup should be on the calendar.
* Set an explicit "definition of done" for any given task. e.g. Code is written, tests pass and cover the new functionality, deployed and works in xyz environments, etc.
* Tell everyone that it is completely okay to mute slack notifications for a few hours at a time so they can focus. Also let them know that if anyone outside of the engineering team is reaching out, they should redirect things to you with no exceptions. Also set the expectation that you expect people to not be a hero. If a sprint isn't on track, it should not be up to any one person or the entire team to work overtime to try to make it happen. Time spent not programming is as important as time spent programming. It is okay for a sprint to fail during this adjustment period and when you're making other changes to scoping. A sprint failing is not a reflection on the quality of the team or anyone's work. It simply means that the sprint was inappropriately scoped for the time allotted.
* Pull two small (~.5 day) tasks into the sprint per developer. You're shooting for 50% utilization for now. Each of these tasks (especially the ones that are assigned to the junior developers) should be very clearly defined with specific steps to reproduce if it's a bug or specific acceptance criteria. This adjustment period is not the time to commit to large chunks of work unless you can figure out how to break them down into small, shippable, bite sized tasks that can be done in less than a day and in isolation. This is as much a learning exercise for the developers on the team as it is for you. If it will take more than half a day, it's too big. Defer it or break it into small, shippable increments. A task is not eligible for a sprint if it doesn't have clearly defined criteria written in the ticket.
* Standup agenda: 1. What did you do yesterday? 2. What still needs done today? 3. Can those things realistically be accomplished before the retro this afternoon? 4. Do you need support to get those items resolved before the retro this afternoon? 5. Who can commit to providing that support? --- For question 5, make sure that you pay attention. If the person that commits to providing that support still has a ton of stuff to do before retro, push back. The answer to question 3 might not be "Yes". Decide which task should slip: either the task that belongs to the person that needs support or the task that is intending to give support.
* The first sprint is probably going to fail. At the end of the second day, hold a retrospective and preface the discussion with something to the effect of "I'd like to understand what either the business or the engineering team can change to help the next sprint operate more smoothly.", give out notecards, and have people write their feedback anonymously. Junior developers may not be comfortable giving direct feedback about how things work, especially if you are causing problems. Same goes for more senior engineers or other company leadership. I've definitely been in that boat before and it took months to figure out wth I was doing wrong so I could fix it. Anonymous feedback opened that channel quickly and after seeing my in-person reaction to that feedback, they were much more comfortable with giving it verbally. I know it sounds cheesy and ineffective, but it works and junior engineers can have really valuable insights since they're coming at an engineering role with fresh eyes. It's really easy for more senior staff to get tunnel vision when they're in a familiar environment and not question why we're doing x.
* If the current sprint failed, do one of three things (decided by the team): 1) Lower the utilization percentage that you're shooting for in sprint planning by 5% 2) Change one or two things about your organizational process based on the feedback that you got during the sprint retro. 3) Remove a day from the next sprint.
* If three sprints in a row are successful, do zero or one of two things: 1) Add a day to the next sprint and keep the utilization the same or 2) Keep the sprint length the same and increase the utilization by 10%. Do not do both of these things in the same sprint no matter how good you feel about how things went, and don't make this adjustment before three sprints. This will ensure that you can _reliably_ deliver the contents of a sprint before you start changing stuff. If you're moving at a good clip, things are being delivered reliably, sprints are at a comfortable size, etc, it's completely okay to not make any scope changes here.
* Do not add new work to a sprint if the team runs out of things to do. What you committed to in sprint planning is what you get for these two days. If one or two people are done early, those people should go help with any outstanding tasks. This is a really great opportunity for pair programming, which would be very valuable to your more junior engineers. If everyone is done early (and again: you already have a well defined definition of done: things need to be either in production or ready to deploy), then let the team figure out what to do with the time. A few examples that I've seen: work through some relevant online learning resources/generally focus on professional development; identify potential quality of life improvements in the codebase (e.g. tests that could be written, components that could be refactored, QA process improvements, CI improvements, etc) but do not work on them directly -- save the ideas for sprint planning. the point right now is to identify improvements, but not take on potentially disruptive/sprint-breaking stuff in an ad-hoc way; listen in on a customer call to hear about real, day-to-day pain points in the product; take a walk/go get coffee/etc with another team member and talk about the work that has been done, what's coming up, what they've struggled with in building the product, etc; contribute to an open source project that you rely on;
* If you add a new team member: 1) their first day should be the first day of a sprint. 2) assign them nothing for the first sprint except for onboarding, project setup, etc. 3) lower your utilization by 10% to account for the time that your team will need to get the new person up to speed. You can increase it after three successful sprints per the above process.
Repeat this process until you're at a comfortable sprint length. I've found 5 days to be pretty comfortable (and it avoids having mega tasks that balloon over time and never get across the finish line after weeks and week of work), but that will obviously vary by team. This is going to feel like a really weird process in the beginning for pretty much everyone involved. People will either be really stressed or have nothing to do. It self corrects over time if you trust the process. It's gonna feel like you aren't getting much done for a little while, but this is really a point where you need to slow down for a bit in order to speed up.
One contentious point in this process is only aiming for 50% utilization. This is often questioned/outright ridiculed by engineers and company leadership. As the person responsible for this team, it's your job to remind them that writing code is only part of the job. "Utilization" doesn't refer to how much any given team member is being utilized. It's how much of their time they're utilizing for directly working on a ticket. Supporting other members of the team to finish the sprint successfully, helping out with QA, doing a deployment to production, internal meetings (maybe), etc are all parts of the job that take time, but they're often neglected when trying to figure out how much stuff to load into a sprint. Having empirical data on how much time these things take is not important. What is important is that they happen, and if your sprints are packed to 90-95% of the available time, you're just setting yourself up for failure.
Hope this helps! Good luck!