22 karma · joined November 10, 2017
https://www.dunderdata.com/master-data-analysis-with-python
* Redness 5.8 vs 1.1
* Swelling 6.9 vs 1
* Pain 86 vs 23
* Fever 10 vs 1
* Fatigue 60 vs 46
* Headache 55 vs 35
* Chills 28 vs 10
* Vomiting 3 vs 1
* Use of pain med 37 vs 10
Only 1100 kids in the study and only 660 of them were followed up at least 2 months after the 2nd dose. These numbers seem unreasonably low.
This means that hundreds of kids had worse reactions than the placebo. All of this to prevent 11 symptomatic cases. I couldn't find the prognosis of these 11 symptomatic cases, but its highly likely that all of them were mild. To me, it seems like the overall health and well being of the vaccinated group was less than the placebo.
Seychelles just reported 1800 cases for the last week or 1.8% of its total population. This equates to 6 million cases in the US or more than 800k per day, a number never approached.
> It looks like the people testing positive in the Seychelles are mostly unvaccinated
It stated that 1/3 were in fully vaccinated and the rest were in partially vaccinated or unvaccinated. There are now 60k that are fully vaccinated, leaving just 8.5k first-dose only. The article was written before the 1800 cases dropped a couple days ago.
More data is needed for hospitalizations and deaths, but as it stands now, it does not look convincing for the vaccine. It's likely more cases will continue to come over the next few weeks, even more than the record breaking prior week.
Seychelles and other island nations that had no major covid outbreaks before vaccine launch (only 500 cases prior and 7.5k since) are the best real life settings we have for vaccine efficacy. For much of the rest of the world, a huge number of people were already infected before vaccine launch, making it difficult to separate out the effectiveness of the vaccine.
Edit: Here is an animation that shows the usual correction - https://twitter.com/jburnmurdoch/status/1350079943863115777
Scapegoating the orthodox jews? More than 92% of 70+ in Israel have received both doses of the vaccine. Are you blaming the high number of cases on this very small segment?
What happens when elderly all cause mortality is elevated substantially in 2021? This is the ultimate test for the vaccine.
This is even after a down week in cases. Next door Lebanon is seeing cases fall at the same or faster pace and they peaked at the same time as Israel. Lebanon has yet to start vaccinating.
Case fatality rate has not fallen either. You would expect a huge drop in CFR if the elderly were protected, but this hasn't been the case.
This is despite over 80% of its 70+ age group having both doses of the vaccine. At least half of the 70+ were fully vaccinated by January 23rd.
It is the elderly all cause mortality that ultimately matters to me. If elderly are dying of other causes, then there are problems.
Regardless, 2021 elderly mortality in Israel will serve as a good test for the vaccine as the vast majority of 70+ will be fully vaccinated for nearly the entire year.
* 99.96% of the placebo group did not have a serious case of covid
* 99.25% of the placebo group did not get covid
* No one in the trial was deliberately inoculated - participants just lived their lives as normal.
* Only symptomatic participants were tested - this seems unbelievable to me - we have no idea what the actual incidence rate is because not everyone was tested
* Vaccines might just mask symptoms - since not everyone is getting tested, vaccine makers just have to make sure there are no symptoms. No symptoms equals no test.
* No trials done with two placebos - we need trials where both groups are in a placebo group. One gets a shot that gives a mild side effect and the other gives no side effect
* No trials done with unrelated immune boosters - we need to see how well this vaccine performs against other immune boosters. This could be a drugs or even supplements (vitamin D & C and exercise).
* Two shots were given in a trial lasting just 4 months - 4 months is an incredibly short amount of time to know whether it will be effective long term. It also gives Pfizer two chances to boost immune
* Long term health consequences of vaccine - we only have 4 months of data, which is way too short of time to see any longer term consequences
* There is a huge incentive to provide something that masks symptoms - billions of dollars are at stake. Big pharma is one of the very last companies I would trust with a novel vaccine
* No coronavirus vaccine in history - many coronaviruses currently circulate, but no vaccine has even been produced for them. It seems quite coincidental that humans finally put the pieces together for our current novel strand
We now know that this apparent good result was not because of masks, but because of geography. Czechia is ranked #1 in the world in cases / million in the last week and #3 in deaths / million (and its numbers are still climbing). Every single eastern european nation had low cases/deaths in the spring and now all of them are having their first large surge.
It's abundantly clear that cloth masks, as a whole, do not play a major role in how the virus spreads in a country.
There should be a library to do exploratory data analysis quickly, without having to touch matplotlib, numpy, or pandas, and without installing something like pandas-profiling to make reports.
This is where the apps will come in to allow users to quickly generate reports on things like missing values, duplicate rows/columns, outliers/bad data, view different colors, etc...
I'm definitely open to looking at alternative backends in the future and will check out PyQtGraph, but am sticking to matplotlib for now.
You cannot control axes plot figure size from seaborn directly. You have to access the figure from the axes (which most people don't know how to do) or create the figure first by importing matplotlib. Really annoying for those that just want to analyze data quickly. Grid plots have the ability to adjust figure size, but return a seaborn object and not a matplotlib figure.
Agreed, docs need to get better. Better datasets, a gallery, etc... I've only spent a week on this, so there will be a lot of improvements in the future.
• Not allowed to set figure size
• No wrapping of tick labels
• No strings for pandas aggregation functions
• No automatic ordering of x/y labels (dexplot provides several options)
• Having to use separate grid functions (catplot, lmplot) for multiple subplots
• Something like 5 different functions for scatterplots. Dexplot has one
• No relative frequency bar charts, which are a fantastic way to explore data. Dexplot provides normalization over any set of variables
• No stacked bar charts
• Seaborn docs have distribution plots (box, violin) in the "categorical" section. A major distinction needs to be made between plots that aggregate, those that show distributions, and those that plot raw data (like scatterplots)
• Returning of matplotlib axes or seaborn grid objects. Dexplot always returns the matplotlib figure
• Seaborn is essentially dead as far as I can tell with few changes in the last 2-3 years. There are even parameters that continue to be non-functional
In the future, Dexplot will add:
• Many more plotting functions
• Several apps (built from ipywidgets) to explore data. Currently, there is one for viewing colors
• Better automatic figure sizing (it exists now, but will be improved)
• Automatic DPI detection so that matplotlib inches correspond to actual screen inches
Dexplot aims to be very intuitive, easy to use, consistent, and allow easy exploration (the name is a smashing together of data exploration plotting).
Here is one example comparison between dexplot and seaborn. https://twitter.com/TedPetrou/status/1271436948721328129
Examples such as these are what drove me to create the library.
I'd love to get feedback and happy to take detailed criticism.
• Not allowed to set figure size
• No wrapping of tick labels
• No strings for pandas aggregation functions
• No automatic ordering of x/y labels (dexplot provides several options)
• Having to use separate grid functions (catplot, lmplot) for multiple subplots
• Something like 5 different functions for scatterplots. Dexplot has one
• No relative frequency bar charts, which are a fantastic way to explore data. Dexplot provides normalization over any set of variable
• No stacked bar charts
• Seaborn docs have distribution plots (box, violin) in the "categorical" section. A major distinction needs to be made between plots that aggregate, show distributions, and those that plot raw data (like scatterplots)
• Returning of matplotlib axes or seaborn grid objects. Dexplot always returns the matplotlib figure
• Seaborn is essentially dead as far as I can tell with few changes in the last 2-3 years. There are even parameters that continue to be non-functional
In the future, Dexplot will add:
• Many more plotting functions
• Several apps (built from ipywidgets) to explore data. Currently, there is one for viewing colors
• Better automatic figure sizing (it exists now, but will be improved)
• Automatic DPI detection so that matplotlib inches correspond to actual screen inches
Dexplot aims to be very intuitive, easy to use, consistent, and allow easy exploration (the name is a smashing together of data exploration plotting).
Stack Exchange is not really any different than visiting a stranger's home. SE has its own set of rules and customs. It takes effort to sit back, listen, and understand them without questioning them. Just accepting them for what they are without any judgment.
Stack Exchange has high standards for what it considers to be an acceptable question as well as an acceptable answer. These standards are quite different than nearly all other forums that came before it. It was their site and they created the rules. It's our job to play their game.
If I don't like the rules of the game that they've laid out, then I can play a different game. There are literally hundreds, if not thousands of other forums where you can ask questions and get answers with less standards and barriers to entry than Stack Exchange.
One of the reasons SE emerged as the leading Q&A site is that they imposed these strict set of rules. They provided a very structured way to both ask a question (minimally sufficient reproducible example) and to produce an answer. Another goal was to limit discussion.
Packaging all these ideas together along with points and badges lead to its tremendous success. Getting answers to programming questions before SE wasn't nearly as easy. But their gamification with explicit rules worked better than that came before it.
If the first Q&A volunteers from SE did not adhere to these rules of the game and if they were not strictly enforced, then it wouldn't exist today. It exists precisely because these rules exist and were adhered to.
As a human being who uses emotional-based reasoning, it can hurt whenever an interaction on SE is met without the expected conclusion. When someone posts a question on SE, the expectation is that it is answered. Getting a question downvoted or closed does hurt, but it's very likely that these actions are just the rules of the game playing itself out.
This really isn't different than playing Monopoly and landing on the 'Go to Jail' square. It doesn't help meeting your goal for the game and can sting in that moment. Eliminating the Go to Jail square from Monopoly fundamentally changes the game. You can imagine more 'negative' rules from Monopoly being eliminated such that the game wouldn't even resemble itself anymore. And if such a game were invented, it likely would never have been successful. If I don't like Monopoly's rules, I can play another game.
SE is successful because there are a strict set of rules that restrict the set of content that can be posted. It is these rules that have lead to its success.
This doesn't absolve SE from criticism just like it doesn't absolve Monopoly from criticism. You are free to dislike any of its set of rules. But at the end of the day, you still have to play by the rules or find a different site.
This also doesn't absolve any of the people moderating, asking or answering questions, or commenting. It's possible to break the rules or code of conduct. SE has guidelines for this as well. As humans, there is going to be some failure to maintain civilized discussion. But, enforcing a rule correctly as a moderator that someone else might perceive as negative (going to jail) is just part of the SE game.
This process is becoming much more robust and standardized thanks to the new ColumnTransformer which allows for applying transformations separately (in parallel) to different subsets of the data. It is built to accommodate Pandas DataFrames, so you can give it column names. The OneHotEncoder has been upgraded to handle string columns.
I am very excited about this release as handling string columns was easily the worst part of Scikit-Learn and there was no canonical way of going from a Pandas DataFrame to a Scikit-Learn estimator. I also cover KBinsDiscretizer which bins numeric columns and will replace Pandas cut and qcut functions in your workflows.
Appreciate any feedback on the article.
The reason 10% was set as a threshold is that many people actually believe it to be higher than that. 10% is an absurdly high number of toxic posts and the real number is going to be far lower. I value authentic statements with accurate data to the highest degree.
I hope that we can get some real research on this topic with more accurate data which can only help to improve outcomes.
The comments from SO were, unfortunately, strawmen arguments. They apparently indulge in self-flagellation to appeal to the loudest.
I think it would be better if the SO team said something assertive along the lines of "No, we are not a toxic wasteland, but we have problems ..."
I am not denying anyone's personal experience or that there isn't a problem. I am rejecting the claim that SO, as a whole, is a toxic wasteland.