(Yet another in the long list of Very Popular But Also Just Wrong ideas surrounding this not-that-complicated concept.)
What the monadic interface to the IO type (very carefully phrased) in Haskell allows is building up a value to be interpreted at run time with a convenient API that looks very like it is an imperative programming language. But it actually constructs a large, complicated value that (per the linked article) is really just a big data structure, which at runtime Haskell lazily walks through (laziness being key here since the vast majority of programs have effectively infinite loops which mean the data structure representing the program would be infinite in size if strictly manifested in memory) to run the program.
In practice, you don't need to think about this distinction and you program with the IO type in the top-level IO context as if it is "really" an imperative language that is really doing what you tell it to do.
But "monad" has nothing to do enabling side effects. Monad is just an interface. Interfaces lack the power to fundamentally rewrite the nature of a language. If you just hacked out the place in the Haskell base library where IO declares that it is an instance of Monad, you could still do everything Haskell does today. Nothing about IO itself would change. What would change is that all the generic functions that have a Monad constraint would no longer work on IO, and you'd have to clone all of https://hackage.haskell.org/package/base-4.16.3.0/docs/Contr... into a new specialized set of functions that work on IO. But that's all that would happen, and a whole lot of code would have to be mechanically translated to use these new non-generic functions. Haskell would not break. Haskell would not suddenly become unable to speak on the network or read files. It would merely become less convenient to use for no benefit.
Edit: You may be thinking of pseq, but I don't think that's in the Haskell 98 report.
WARNING: the below contains math. Ignore if math makes you mad.
If "side effects" means the availability of certain operations such as mutating state, throwing errors etc, then monads certainly do have something to do with side effects in that these operations can often be described by an algebraic theory and algebraic theories have a direct correspondence to certain kinds of monads (the monad of free algebras of the theory).
You'll be stuck with a wrong understanding of them as long as you think that's what they are. That is one of the things they can be used for, but that's all. It isn't what they are. There are plenty of "monads" that are not being used for "side effects" and are perfectly pure no matter how you look at them, such as the list instance of monad. Rather than arguing with me about how if you bend the list a bit and sort of redefine "side effect" and if you squint a bit, the list instance really is related to "side effects", just don't look at it that way.
Haskell then decided to use this DSL for arbitrary types, again just to build data structures. Then, people noticed that monads were really cool.