In mathematics monads usually arise as adjunctions between two functors, for example beginning with a set of elements, you can consider the free monoid generated by it and forget the group structure, this gives you a much larger set. If you did this operation on a set of characters, you would get the set of all strings of those characters, eta would in this case be the operation that given a character in the character set gives you the corresponding string of length one and mu would concatenate two strings.
Imagine if you wanted to learn about monads, but every article went "Monads are a simple and powerful idea that, interestingly enough, can be very, very well expressed in Latin. Therefore, I will switch to Latin for the remainder of this article. Oh... You haven't studied Latin? You really should! It's really very useful. Moving on... Cogitus sin extricatus..."
If you want your audience to understand, you need to explain it in C.
There is a translation from any typed language into an untyped language. Writing code in that untyped language is not going to be type safe, while the code generated (correctly) in that untyped language from the typed language is still guaranteed to be correct.
It is entirely possible that the only way to get anything safe out of some Haskell code is to rely on checks the Haskell compiler gives you at compile time, which the C compiler cannot give you.
That said, people often underestimate the kinds of guarantees you can bang out of a C compiler, at the cost of a bit of verbosity.
Following on this thought: Every compiler is essentially an assembler programmer! And we all know how error prone it is to code in assembler. So how can the compiler ever produce error-free binaries?
The advantage of mathematics is that it cleanly separates the external view and internal view of a concept. The axioms are easy to state abstractly and they are the most important part, haskell only allows to abstractly define the type of the operations, but can't abstractly enforce the laws. Rather any instance of the typeclass is assumed to satisfy the laws.
Those happen to also hold for certain constructions in functional programming languages, like lists (the list monad) and several others, not by coincidence, but because those languages have a close connection cartesian closed categories.
I firmly believe it is not helpful to explain something by analogy, because an analogy only goes so far. The mathematical notion of a monad is not complicated at all and is only obscured by writing page after page about them in the syntax of some arbitrary programming language.
A monad is some parametric type, T, along with two functions called (in Haskell anyway) "return" and "bind". Return "injects" values into the type taking values of type A to values of type T(A). Bind transforms values of type T(A) into values of type T(B) using a function like A -> T(B).
Then these two functions must follow a few rules.
That's a monad. The Haskell fragment above describes the signatures of those functions, notes how they relate to the parametric type, and also produces a facility for overloading `return` and `(>>=)` ("bind") and even working with them when lacking a concrete choice of type `T`.
All of those questions change their answer depending, terrifyingly, on how you informally define the word "monad".
Given a type `m` and a type `a` a monad provides functions with types:
bind : m a → (a → m b) → m b
wrap : a → m a