"Stats for scientists" was a one semester course, mostly involving plugging numbers into formulas. There was lots of hypothesis testing.
"Math stats" was for math majors. It was two semesters, and focused mainly on proofs. I took math stats. I also ran a tutoring session for the scientists.
A problem is that the basic stats course is taken by a lot of students who wouldn't have gotten through calculus. So, building stats on top of calculus would have created two forbidding layers of abstraction instead of one.
On the other hand, I graduated from college in 1986, and we did all of our calculus by hand. I wonder if a potential compromise today would be to teach stats by exploration using random numbers. You're still doing integration, albeit numerically, but maybe it wouldn't seem so forbidding. And by playing with random numbers, you can learn the hard way what erroneous conclusions you can draw from them.
I’m not sure there is a single right way to solve any given problem isn a Bayesian way, but it does force you to think more about the problem at hand and make your assumptions explicit.
If you're interested, there's a great 100-level course online from Joe Blitzstein at Harvard: https://projects.iq.harvard.edu/stat110/home
The course was about distributions, how product distributions interact, the central limit theorem, etc. That is, it was about turning stochastic models into predictions. Later we had a course "Statistics" that is about matching seen results to stochastic models.
How can you study these without knowing the Lebesgue integral?
> how product distributions interact
How can you do this without the Fubini-Tonelli theorem?
etc. etc. etc.
Starting with measure theory would be what leaves people feeling that mathematics is useless formal bickering.
The main difference between the two is that a Riemann integral count every area of space the same ([0,1] counts just as much towards your integral as [1,2]), but Lebesgue integrals use what is called a measurable function, which maps f(set) -> positive number. You can use this function to weight different parts of your integral differently. Now what you can do is make measurable functions that map f(set) -> how likely that set happens (which is exactly the probability)