Whilst trying to avoid the "No True Scotsman" fallacy, I'd argue that this system is FP only in name, but not in spirit. Even in Haskell, you can spend all your time in the IO monad and use IORefs as shared mutable cells, but you'd have a hard time arguing that such code is "functional".
> It's hidden, and, when it's being done in the context of a generally functional idiom, it's downright pernicious.
I think what we're seeing is a distinction between _syntactically FP_ and _semantically FP_ qualities. It's easy to apply _syntactic_ idioms obtained from FP, as it allows you to avoid and reduce state wherever possible. However, in a language where mutable state is assumed, and it's your responsibility to not use it, you don't get the _semantic_ guarantees about the behavior of your program.
I don't like having to exercise discipline, because no matter how good I am at it, I'm only a temporary part of any software system. IMHO, the fundamental goal of software architecture is to institute bias directly into a codebase to support the problem domain. The way in which you work with a codebase is informed by how that codebase wants you to work with it: you'll naturally avoid things that are made difficult to do, and prefer things that are made easier to do.
Programming languages are essentially the basement level of any given architecture, because it is nearly impossible to override the decisions your language makes for you. It is almost always going to be easier to use what the language provides you, and if the language provides global mutable state, it will always be tempting to couple two otherwise separate regions of your codebase by a mutable cell. Some languages especially make FP idioms difficult (hi, Java), so you end up fighting an uphill battle -- unwinnable if you're not extraordinarily careful.
> There will always be some business constraint that prompts people to take shortcuts.
To borrow a phrase, I don't think FP can "win" until we deal with the forces that make mutable cells such an attractive choice. There are multiple facets to the problem; it's not enough to just pick languages that make FP the easier option (or mutable shared state the harder option). IMO, we need to have an industrial expectation of domain modeling, and architect our systems specifically with our problem domain in mind, so that problems in that domain -- and expected evolution in that domain -- can be handled not only easily, but intuitively within the set architecture. (Lispers go wild over defining their own language within Lisp for exactly this reason -- but all things in moderation.)