The xeus-cling work is awesome and has made it possible to do data science prototyping in C++. There are lots of other C++ notebook examples in the examples repository: https://github.com/mlpack/examples/
291 karma · joined February 15, 2018
The xeus-cling work is awesome and has made it possible to do data science prototyping in C++. There are lots of other C++ notebook examples in the examples repository: https://github.com/mlpack/examples/
"swim-floaty"? I get what they are referencing, but can we really not come up with a better term than... "swim-floaty"?
Oh, I'm no Luddite, don't get me wrong---I'm a machine learning researcher. I have no problem with data science. The difference between all of the complex technology you pointed out inside the baby monitor and what we're talking about here is that all of that complex technology is robust and, for the most part, well-designed! The data science work here has tons of issues, and I think the vast majority of reactions are reacting to that: machine learning and data science really can be useful... but not if you apply it really badly. Someone else commented elsewhere in the thread that a simple thresholding algorithm would be just as effective---and not suffer from the myriad potential problems present with blindly applying TensorFlow because it's cool.
I'm with you there, that's a big downside too, but that's not the "real" downside---there are like seven different downsides present with the data science going on here and it's hard for me to say which is the biggest issue, because they're all issues. Data science is not trivial!
However it is really important to consider why baby monitors are so primitive: because the cost of a false negative is huge. I didn't see any mention of this in the author's experiments (only a '>98% accuracy' note). So let's talk about this a little bit: is "accuracy" what we want? Probably not---I don't care if I get accidentally notified, but I care very much if I don't get notified when the baby is crying. So you want to weight your classifier's predictions heavily against false negatives (at the price of false positives). It would be good to make an ROC curve to characterize this behavior. More importantly though, any predictive model assumes a stationary distribution; i.e., training conditions accurately reflect test conditions. But will they in real life? What about when your neighbor's house is under construction? Can interference from chainsaws cause the model to fail to detect the baby crying? What about the dude down the street with his super loud motorcycle? What happens then? I bet the training set doesn't have situations like this.
I really, really don't want to come off like a wet blanket here. But I feel obligated to, because this is a model that directly impacts the welfare of a human, and so we should at least talk about or discuss potential drawbacks. (Again, cool weekend project, just, we need to be clear about the implications of outsourcing the decision of whether the baby is crying to a black-box model where we can't interpret what it's doing.)
In Georgia, the ballots are printed from a terminal that the voter uses and then scanned, leaving both an electronic count and a paper trail (the ballot itself is ejected from the bottom of the scanner into a sealed ballot box).
However, our scanners jammed 40 minutes into the day; after a couple hours, a technician managed to come to our precinct and opened the ballot box and revealed that a lackluster design in the ballot box caused the ballots coming out of the scanner to sometimes curl up and jam. Without any realistic solution, we just had to open the ballot box every time it jammed (supervised every time to ensure no monkey business) and push any stuck ballots out of the way of the scanner so that more could be scanned. Amusingly we had good success regularly giving the machine a good shove to dislodge any stuck ballots.
We also had problems printing receipts---in our case, we only need to print 3 from each scanner, but we ran out of scanner receipt paper. Since another precinct called us during the day looking for extra receipt paper... it wasn't available. But, we dodged a bullet, since there was just enough paper to print 2 of the 3 receipts. (1 gets posted on the door of the polling place; 2 go to the county. We wrote an apology on the receipt and only sent 1 to the county. I verified on the Secretary of State website that the votes tabulated for our precinct matched what our receipts printed. Cool to be able to double-check like that!)
I spent a while thinking about what a pollworker would need to do to illegally cast ballots. It would be a tall order indeed and would require cooperation and secrecy from everyone there, since the only way to cast a ballot is to scan it, and everyone can see the scanners at all times. I can't see it realistically happening in any precinct.
Of course, I'm not actually internal to Microsoft or Github, so I have no idea and it's all opaque to me.
It seems like the historical uptime page paints a far rosier picture than I am actually experiencing.
...is it time to move away from Github?
1. We played Minecraft with the specific intent of making visually appealing buildings. So at some point, when you can't see, that's not going to be fun no matter what you do...
2. Minecraft really doesn't have any accessibility whatsoever. You can scale the UI, but... high contrast mode? If you could even get that working with the base game, it's definitely not going to work with the mods we were playing with. As my friend went blind, it got harder and harder for him to deal with any zombies or anything that was moving, since it took him so long to slowly scan the screen and understand where he was. We considered trying to make mods to make things a little bit easier, but struggled with coming up with any mod that would actually improve things. :)
3. My quick glance at the API suggests that we'd end up with a situation where we were playing the game and he was... programming. That may work for some people, but I think for this group that borders too close on after-hours work...
For a while he was DMing and he would use a screenreader to access his notes plus the rolls. We don't typically use maps or boards, but instead try to do it all with descriptions of places. It does mean that the rooms we enter all tend to be fairly simply shaped, and it's possible that each of us has pictured a slightly different room, but it all works out in the end.
Don't want to move to California? Can't pass the interview? Salary too low in the offer? There could be lots of reasons.
I agree with your other points, but what I'm saying is not exactly that. What I'm saying is that TensorFlow is not a community-led project in the way that other open-source projects are. It's Google-led, and if you don't like it, you can go away or fork it (...and if you fork it, it's unlikely you'll get many users given the Google marketing machine).
(The same complaints have been and are being regularly made against Red Hat, and also Fedora, given that Fedora is basically "experimental" Red Hat.)
You think they are publishing all these papers and open-sourcing TensorFlow for the good of the community solely?
No... if the entire ML community publishes papers with TensorFlow as code, then now Google can sell that many more TPUs and that much more access to their cloud systems to run the models...
But with Google-run (or other company-run) open-source projects, it's a completely different story. Want to be a TensorFlow maintainer? Great, sure, you can contribute, but you'll never get to make any decisions about the direction of the project itself unless you're inside Google.
To me it's a disappointing semi-abuse of the term "open source"... it doesn't mean what I think it should...