Most of our test pipeline was broken, due to a long series of weird events, in which the uninstall phase of the previous version of one of our apps, on our CI computer was stuck, and on another CI computer somehow the CI user gitconfig file was somehow corrupted and this failed another test which tested our git interface. We're not sure why all this happened. Also we have started using nugets (I don't like them but the guy who's in charge of the CI wants to use them). And something was completely broken with the permissions on the internal nuget server we use. We're still not sure what. But IT took 2 weeks to sort it out.
Our team is very small. It's only one guy who deals with the CI and he does it part time. I used to do it, and then we had another excellent new guy who took it upon itself and really built amazing and mostly stable pipelines for all our products all by himself. But he left several months ago and management didn't hire a replacement. So now our CI pipeline is sometimes broken for days without anyone noticing.
Speaking specifically to test automation I find that most companies simply prefer to make the calculation that they'll take the risk of an issue occurring over investing in test automation (and testing in general) until the problems become so large and financially catastrophic, or a regulatory issue, that it gets embarrassing for a VP/Exec and then the decision is made to properly invest. That's not an ideal approach because building this out as you go is more efficient.
The typical other problems are how best to shape the testing pyramid. Having large number of unit tests > medium number of integration tests and a small number of end-to-end tests is always proposed but I find in most cases the pyramid becomes inverted due to the natural issues associated with problems really being only of concern if they actually impact an end user and the end-to-end issues cast a wide net even though they are very challenging to maintain and do require appropriate investment in resources.
I want to mention Tesults (https://www.tesults.com) with which I am involved to tackle another issue. Clarity around test results and what is actually being tested due to bad reporting and trouble accessing relevant results data. Understanding testing across the systems for a team is essential to knowing what's actually running, how often and what the output is. This helps understand what additional testing may be needed and identifying what's failing.
Overall, most of the issues related to test automation stem from the calculation made too often mentioned above.
I want to understand how universal these pain points are VS unique to the team I am working with at the moment.