All that's fine, and if you keep clicking through some of the links I posted you'll find that the sorts of analyses you're asking for have been done, because indeed, it's not enough to merely see smoke. You need to find the fire too.
I can't sum up every possible argument in a HN post, only show that there are problematic signs worthy of further and deeper inspection. That is, skepticism isn't some weird cranky idea. At least some of it is based, at heart, in the question of whether there are mistakes in algorithms or software. That's a totally normal question to ask about any data analysis.
For instance, let's take your audit point. You audit the data going back. In the USA this is somewhat possible. For the HadCRUT dataset it's not, because they lost/destroyed the original raw data, and their computations to clean it up are not reproducible, as the notes about "Harry's" attempts showed quite clearly. That is already a major problem; in effect we're being asked to take the correctness of CRU's data adjustments on faith.
So let's focus on the NOAA dataset. A large amount of it is synthetic, more and more over time. You say we can't just scream "fake data" and remove it, and I agree we can't because that would at this point erase most of the official US temperature record. But ... seriously? You don't find it at all concerning that in the past humanity was apparently capable of measuring temperature with thousands of weather stations across the USA, and now we've apparently lost that ability just when we need it most - forced to rely on 'estimates' generated by computer simulation? The foundation of science is fitting theories to match the data, but here, the data is being generated by the theories. I think anyone who cares about science should care at least about that.
Now, which adjustments are questionable and why - that's a great question and is pretty much where I'm up to in my own research. There are two main strands of alteration: TOB adjustment, and the PHA.
TOB is "time of observation bias". PHA is the pair-wise homogenisation algorithm. There are other adjustments made too: more and more algorithms have been added over time, but those seem to be the two big ones. PHA seems to be involved in the 'estimation' process to fill out or adjust the reported weather station readings.
This graph shows the affect of them split out:
https://realclimatescience.com/wp-content/uploads/2019/08/US...
As you can see the adjustments have little/no impact today. But they rewrite the past significantly: raw data pre-adjustment doesn't show any warming in the 20th century, the 1940s in particular were as hot as today. TOB adjustment was developed first and pushes down the temperatures in the past, then PHA+extras (what this graph labels "final") was developed later and pushes it down further.
This can be interpreted in a few different ways. One way is that climatologists found that in the past people were consistently bad at measuring temperature accurately, so they process the data to remove the failures and the data is now accurate. Another less charitable interpretation is that the algorithms have the effect of fitting data to match the theory, a form of scientific malpractice that is not unknown in history. To figure out what's going on requires much more analysis.
Note that nothing here requires an evil conspiracy. Climatologists were faced with a huge problem in the 1970s: basic chemical theory said CO2 is a greenhouse gas, we're burning lots of fossil fuels, so the world should be getting warmer. At least if you assume the climate is quite simple. But the raw data didn't show that. It was also messy, filled with glitches and could obviously benefit from some cleaning up before usage. So scientists started trying out various analyses and adjustments, but ultimately they're not comparing against a ground truth. They don't know what the data "should" be except through climate theory, so it would be very easy to end up trying lots of algorithms and ending up selecting the ones that result in "obviously correct" outcomes. Meanwhile the number of people involved in this is very small, they're all colleagues, their income depends on the belief their understanding of the climate is always growing, and it's unclear how you could ever resolve the apparent contradiction between chemical theory and raw thermometer data in a field like climatology. It's not like you can experiment on the Earth.
In other words, there's a big grey area between cleaning up the data and adjusting the data to fit the theory. Did they go into that zone, and if so, did they cross the invisible line? It's subjective, but not obviously impossible.
Anyway back to the analysis. So far I only found an analysis of TOB adjustment. It can be found in the second half of this rather long blog post:
https://realclimatescience.com/erasing-americas-hot-past/
Start reading from "I have new tools". The quick summary is: some weather stations used to be read in the morning and others the afternoon, and this was suspected to need fixing. The author analyses Minnesota because it has a balanced mix of AM/PM stations, and shows the PM stations are indeed consistently reporting about 0.5 C of difference to the AM stations. But then he analyses their coordinates and finds that this is because the PM stations are on average a degree of latitude further south, and that average temperature increases about 1C per degree of latitude, so you would expect such a difference based on location alone. In other words, correcting for location eliminates any difference that could be caused by TOB.
I'm still in the process of checking these claims. I didn't reproduce the graphs yet like I did for the claim of synthetic data. The claim 1 degree of latitude = 1 degree C of warmth appears backed up by this source:
https://www.physics.byu.edu/faculty/christensen/Physics%2013...
So it sounds plausible. If no TOB is observable in Minnesota, why did NOAA scientists conclude it was a necessary adjustment to make everywhere? It's also kind of suspect that this would matter, since random noise should get washed out by the averaging and long term trend analysis.
BTW, you ask do I accept that data has to be adjusted because stations moved. I certainly did at first, when I started looking into this. The deeper I look the more I'm starting to question this. Station moves and upgrades can certainly introduce artifacts into the reading of a single station, but, scientists only really care about global averages. For this to require correction at all would require the effect of moves to be highly correlated.
Final notes on the science.
what you very probably should never do unless your intent is to mislead, is simply present 100% raw data on a graph
That's not the science I was taught! One of the pillars of science is providing your raw data. If there is uncertainty in the measurement that's what error bars are for, and the various statistical techniques for handling uncertainty. Picking data points and excluding others that don't fit the theory is literally "bad science 101".
Also, if climatologists have such severe concerns over the integrity of their raw data that they need to interpolate or synthesise nearly half of all readings, they should be screaming at the top of their voice that nobody should do anything until the global weather station network is in a better shape and that should be their top funding priority. Are they doing this? If so I've never seen any reference to it. We should be seeing more measurements and better data over time, not a decline in the number of trustworthy weather stations (which is what's happening, hence the rise in synthetic data).