* old attitude: why does pandas have to make things so hard
* new attitude: pandas has a crazy difficult job
I think this is most apparent in the functions that decide what "[d]type" a Block--the most basic thing that stores data in pandas--should be.https://github.com/pandas-dev/pandas/blob/4edcc5541ff3f6470f...
And then, for the ubiquitous Object dtype, often figure out which of the many possible more specific types to cast it to.
If you think that is easy, ask yourself what this outputs:
import numpy as np
np.array([np.nan, 'a'])
Lo and behold--it produces an array where the np.nan has been converted to the string "nan".And yet
import pandas as pd
pd.Series([np.nan, "a"])
Knows this, has your back, and does not stringify it.It also has a pathological fixation on when it tries to convert dtypes, since avoiding all the bad conversion outcomes is a relatively time intensive process (compared to e.g. creating a numpy array).
I realize things could be much easier in pandas user facing interface, but really appreciate the sheer amount of effort that has gone into its dtype wrangling.