import pandas as pd
import matplotlib.pyplot as plt
import matplotlib.dates as mdates
import seaborn as sn
df = (pd.read_csv("/tmp/overview_2022-03-21.csv") # i just used curl beforehand
.assign(date=lambda x: pd.to_datetime(x["date"]))
.set_index("date")
.melt(value_vars=[
"newCasesBySpecimenDate",
"covidOccupiedMVBeds",
"newAdmissions",
"newDeaths28DaysByDeathDate"],
var_name="Data", ignore_index=False)
.assign(Data=lambda x: x["Data"].replace({
"newCasesBySpecimenDate": "New Cases",
"newAdmissions": "Admissions",
"newDeaths28DaysByDeathDate": "Deaths",
"covidOccupiedMVBeds": "Ventilated"
}))
)
ax = sn.scatterplot(data=df, x=df.index, y=df["value"], hue="Data")
ax.set(xlabel="Date", ylabel="Daily rate", yscale="log")
ax.xaxis.set_major_formatter(mdates.DateFormatter("%b"))
plt.show()
I spend 2 minutes on the pandas part and 20 minutes on the plotting part, which really says it all. Seaborn's support for smoothing is really bad and doesn't play nicely with datetimes for some reason, so if I wanted smoothing I'd need to do it myself. And the other stuff I left out requires going into matplotlib's documentation which I don't want to spend time on.pandas is as good or better than R's dataframe manipulation, but R's plotting tools are best in class. I hate all the python plotting libraries.