"Monad" is an abstraction. What kind of abstraction? An algebraic abstraction.
That means really understanding the concept is much like understanding other algebraic abstractions.
The most basic concept in classic 19th century abstract algebra is "group." Just like the monad concept, this concept involves a set of values that can be combined with a few carefully chosen operations.
Just like with the monad concept, the group concept doesn't lend itself to immediate grokking. So a lot of people get frustrated by abstract algebra. They feel like someone just isn't telling them what a group "actually is."
But group is an abstraction over concrete "implementations". It is a common base for many algebraic topics, like integer arithmetic, modular arithmetic, matrix arithmetic, polynomial arithmetic, and even more complex structures. If you are familiar with the theory that applies to groups in general, you have access to proofs and formulations that can be applied to many different topics. Sometimes that generality is useless, sometimes it is very productive and succinct.
What the groups have in common are a binary operator that's associative (like plus or times) an identity element (like zero or one) and a way of taking inverses. This is all codified as group axioms. If you just look at those axioms you might say "so what?" but the concept is born from actual mathematical practice and is significantly useful and interesting.
Monad is an abstraction over different computational topics: I/O computations, randomized computations, failing computations, and so on. It captures in an elegant and abstract way the operations and elements required to express these topics. General functions can be written polymorphically over all monads, just like theorems and computations can be written to work for all groups.
So for a monad, you need return, which lifts a base element into the monadic class of values. (This notion of having a base element and a lifted set, for example Int and Maybe Int, is itself a basic abstraction that monad builds on, namely the functor abstraction, whose only operation is fmap, an abstraction of list mapping.) And you need bind, which is some way of combining one monadic value with a function producing another monadic value.
Those operations need to work together in reasonable ways specified formally by the monad axioms or monad laws.
Again, you can look at all that definition stuff and say "So what?" But again, the concept makes sense, it is useful, and it is born from abstracting over concrete topics. (Moggi wrote the first paper about the usefulness of the monad concept in computer science; the concept originally came from category theory, which is kind of like abstract algebra.)
The do notation is a good example of the usefulness of having an abstract type class for monads. It gives you syntactic sugar that works in a well defined way across many many topics.
Abstract algebra doesn't make any sense if you don't know how to work with plus and minus. Monads don't make sense if you're not comfortable with implementing pure combinators for simple computational structures.
So you should find some way to practice some of those basics, and then the abstraction "monad" will have meaning and not just look like a random assemblage of made-up rules.
Look at an example of using do notation with Maybe types to express failure. It's not that amazing, but it's useful enough, and makes sense. Now learn to implement the same thing from scratch without the syntactic sugar and without the monad functions. You will first write the whole thing with explicit pattern matching, tediously. Then you will implement the crucial combinators that lets you "bind" one Maybe value to another computation returning a new Maybe value. Then you can look at the source code for the Maybe instance of Monad and see that it is just the combinator you have written.
Then you can study the State monad in a similar way.
And then you can notice how the general monad functions are useful for both of those topics, failure and statefulness.
The concept monad in the context of category theory is even more abstract, but there's no need to worry about that level of abstraction merely to learn Haskell programming.
The reason monads are a big deal in Haskell is that the computational structures they conveniently express happen to be those which are otherwise described with "imperative" language features: mutation, jumps, and side effects.
So if you're interested in expressing those computations in a pure way, you should be a bit curious about the monad concept. If you're not interested, that's fine, but it's close to Haskell's reason for existing, so that's a more basic question: is it interesting to write programs in a pure way? If you say no, you're right to give up on Haskell.