Stan is one of those technologies I keep finding is actually powering the more 'friendly' interfaces I run to one off jobs - especially in the mcmc world. Every so often I think I'll spend some time to learn stan proper but it's such an all-encompassing project that I get intimidated and stick to the derivatives. My loss!
Bravo to the team behind it and for making and supporting such a powerful tool!
Transport Tycoon and Railroad Tycoon are not related despite both having trains and the word tycoon in them. Both great games though. I play OpenTTD to this day.
Data scientists (in my opinion as one) should be spending most of our time listening and talking to our colleagues, our clients, our peers, teachers, and - last of all- to the data. Last because we dive into the data searching for things, answers to questions, people asked for and that other people inform our journey and our methods.
Counter-intuitive in regards to the phrase but good ds is people work first and data work like 9th.
Talyllyn on its eponymous railway in Wales has been in regular service since 1864. The big boys are actually some if the youngest steam engines, built right at the very end of the age of steam in the US.
Two thoughts:
1) Much of this is the bias of this community towards new and more more technical companies. At many of our work places Data Science is just part of the landscape and is now baked in. But there is a wide gap between that and legacy companies where the idea of using these techniques this way isn't new at all. In other words we've been through the initial wave of the hype curve and are now into mainstream growth of the idea.
2) Branding matters. In many companies that I get to interact with (fortune 100s) run-of-the mill analysts exist for ass-covering the non-data-based decisions of higher ups. Post-hoc decision justifying. But DS teams (currently) actually get a seat at the table and get listened to. If for nothing else that makes DS way more effective then just BI or whatever. And as long as we get things right that should turn into a virtuous cycle of being heard, being seen to contribute, and being asked to contribute more.
The value is a "Data Scientist" isn't that they know how to use a tool - it's that tell know why to use _that_ tool (technique) and not this other one.
It's not at all unusual to have runs of stories. Stories are not independent things. Stories are the public face of a background process (investigative reporting) that surfaces many interrelated aspects of an issue. Waiting for all the pieces to fall into place is waterfall development, you'd never ship. Writing many stories as the facts become known/interpretable is agile. Stories build over time.
I did exactly this; went to Hunter College CUNY and paid my way through with a part time job. Mom and Dad payed all of freshman year tuition and I did live at home. End result: no student debt. But I still needed that safety net to make that possible.
If you are doing things right you never say the equivalent of "they're all white" - you give a distribution. Explaining that and what it means is a communications issue, not a data one.
It should also be noted that we have anti fraud tech baked into our system that fires before our geodata stuff runs. Fraud gets cleared out for being fraud not for being bad location data.
Not really fraud, phones have different levels of horizontal accuracy based off of privacy settings, GPS & cell signal strength, battery life, CPU load and that's before you get to the obfuscation that exchanges do. At Dstillery we built a geodata classifier to tell us what was and what wasn't good data. We throw out 60% to 75% a day as not useful for learning anything from. But we can be picky since we combine web and location data we aren't beholden to needing the unreasonable amounts of location data you need working with just location data. Location should be holistic part of the data, not the be-all, end-all.
And even before that, the steamboat was invented here, the Manhattan project was called that because it was originally in Manhattan. Edison was over in Jersey and Tesla on Long Island. Republic had it's plant out on the island too where they built the LEM. IBM's Watson lab is just north of the city and Bell labs a bit west. New York has always been a tech hub but it wasn't as obvious since there was - and is - so much else going on.t
My family has had the same apartment in Queens since 1916. It'll go to my brother at some point in the future. An outlier for sure but plenty of the people I grew up with were in multi generational apartments and expected to be the next generation to take it over.
Yes, in a toy space ML replicates simplistic measures, there's only so much information to be had in such a space to start with. The power of these techniques is when we apply them to high-complexity or even unbounded-complexity areas.