200 karma · joined September 20, 2013
Note that it's very tied down to our use case right now: only compatible with Logistic Regression, and currently it assumes fixed hyperparameters (will change this in future though), assumes a production pipeline of min-max scaling, imputation, then classification.
We're using logistic regression not because it performs the best, but because it's the most understandable. When cases get flagged for manual review people need to know exactly what seems dodgy about the account, and with Logistic Regression you can read the exact contribution from each feature to the final fraud probability. Seen as the features mean something real and tangible (unlike in neural nets), this means a manual reviewer immediately knows which aspects of someone's behaviour are out of the ordinary when they get presented with a new case (we have a really nice internal UI for presenting this). This saves several minutes per case which really adds up.
Performance-wise Logistic Regression is good, but it can't automatically learn non-linearities in a feature value and its propensity for fraud, and it can't learn about two features that together should indicate a probability of fraud greater than the sum of its parts* . If this becomes a problem for us we'll start looking into nonlinear models where the inner workings are somewhat communicable to the manual review team.
* You can alter feature definitions manually to capture nonlinearities (e.g. a feature which is "user_has_done_x_and_has_done_y_too", but this is very very manual, and needs to be potentially rewritten/manually re-optimised on every retrain. We don't do this.
{
"intercept": 1.0,
"features": {
"feature_1": {
"coefficient": 1.0,
"range": [0.1, 10.0],
"mean_feature_score": 1.0,
"imputation_value": 1.0
},
{
....
}
}
}-------
Training/retraining
- train on an ad-hoc basis, every few months right now moving to more frequently and regularly as we streamline the process
- training done locally, in memory (we're "medium data" so no need for distributed process), using a version-controlled ipython notebook
- we extract model and preprocessing parameters from the model and preprocessors that were fit in the retraining process, dump to a json config file in the production classification repo
-------
Production classification
- we classify activity in our system on a nightly cron*
- as part of the cron we instantiate the model using config dumped from the retraining process. This means the config is fully readable in the git history (amazing for debugging by wider eng. team if something goes wrong)
- classifications and P(fraud) gets sent to the GoCardless core payments service which then decides whether to escalate cases for manual review
-------
* We're a payments company, processing Direct Debit bank-to-bank transfers. Inter-bank Direct Debit payments happen slowly (typically >2 days) so we don't need a live service for classifications.
Quite simple as production ML services go, but it's currently only 2 people working on this (we're hiring!).
How is an ad buyer ever supposed to make an informed decision about how susceptible their chosen ad vendor is to fraud?
They have to defend autopilot not only to protect the brand but to protect the public's perception of autonomous vehicles in general.
Self driving tech is poised to save many many lives. So from a utilitarian perspective, it's probably justified to take extraordinary measures to make sure reactionary media and public whim doesn't kill it off, however uncomfortable that might seem in the short term.
Whilst this case is incredibly sad (and I don't want to downplay that in any way), if you're trying to minimise the overall amount of fatal crashes, exonerating the tech is the priority (if it is truly not at fault).
US: 11 l/100km, UK: 7 l/100km
UK is pretty representative of Europe. (Also the UK improved by 20% since the 2008 figures). US and Canada are unusual in having such widespread domestic "truck" ownership.
http://www.autonews.com/article/20140403/OEM05/140409928?tem...
https://www.gov.uk/government/statistical-data-sets/env01-fu...
TBH I don't rally like Wired's reporting on this kind of stuff. To get an overall view of whether this is useful or not, we need to know:
- lifetime of the pollutant
- air changes per day
- whether this type of pollutant is an important one to tackle
...without that we can't know whether this is just an art project or something practically useful.
GoCardless (Direct Debit, EU only at the moment) is one company that doesn't charge a fee for chargebacks. They take on all the risk themselves.
- day bus service ends and night buses begin
- your phone signal dies
Does it do it at other times too?
There's probably a more user friendly way to say that. Cool idea though.
Supercritical steam is a special form of steam that can not be described as a gas or a liquid. It's somewhere between the two: molecules aren't bunched together in dense clusters that settle at the bottom of a container (as they are in a liquid), but they also aren't flying all over the place individually in a low density vapour (as they are in a gas).
How's that possible? Water molecules have relatively strong intermolecular attractive forces between neighbouring molecules. They like to stick together, even though there's no permanent connection between them. They are like mini-magnetised marbles. This explains why water has a much higher boiling point than most tri-atomic molecules.
When you increase the temperature of liquid water, the molecules in the liquid vibrate and move around within the liquid, and as you cross the boiling point, the vibration and movement of the molecules is so great that they are able to escape the pull of their attractive interactions with their neighbours en masse. When this happens, the molecules shoot off into the vapour, where there is an (almost) unlimited amount of space for them to shoot around in.
Now consider what happens when you do this at high pressure. High pressure essentially means that there are lots of molecules in the gas phase moving around really quickly. Now, when the temperature gets high enough that molecules have enough energy to overcome their attractive interactions with neighbouring molecules, they leave the pack: but this time with nowhere to go to. The pressure is so high in the 'gas' phase (i.e. there are so many other molecules up there) that they are forced to just bump around where the liquid was but at extremely high speeds. This type of behaviour is pretty difficult to distinguish from the behaviour in the high pressure 'gas' -- in fact, after the system has time to equilibriate, they are exactly the same.
Clearly then, the transition from 'liquid' to 'gas' at this point is pretty much indistinguishable. The liquid may begin to display the molecular kinetic behaviour of a gas, but the density stays the same.
The end result is: When the pressure and temperature is high enough, to onlookers it appears as if the entirety of the fluid is half way between a liquid and a gas, and is stable in that state. That's called a supercritical fluid.
I've spent a lot of time studying/writing/playing music and through that have got to personally know many of the most talented and versatile musicians I've ever come across. These people are skilled like Douglas Crockford, John Resig, you name it. But they have to make the assumption that the music they want to do - their own music - will never make any money in a recorded format, forcing them to do wedding gigs during the day instead.
I'm interested to know what people think about this. Do people think that the end (taking power away from big music industry) justifies the loss for those small-time players, or is it something that simply hasn't been considered at all?
Do you have a justification for saying that all music should be free, or is it just that it would be nice if all music was free?
1) The Seene app (iOS app store, free), which creates a depth map and a pseudo-3d model of an environment from a "sweep" of images similar to the image acquisition in the article
2) Google Maps Photo Tours feature (available in areas where lots of touristy photos are taken). This does basically the same as the above but using crowdsourced images from the public.
IMO the latter is the most impressive depth-mapping feat I've seen: the source images are amateur photography from the general public, so they are randomly oriented (and without any gyroscope orientation data!), and uncalibrated for things like exposure, white balance, etc. Seems pretty amazing that Google have managed to make depth maps from that image set.
Aside: up until now I had imagined that if a skydiver were to ever drop a rock of that size (or a dense piece of equipment like a DSLR) during freefall they would never be able to catch it, but as the terminal velocity of a skydiver in "dart" position is about 320km/h that's not the case. Pretty cool.