HNHacker News
TopNewBestAskShowJobs

angusb

200 karma · joined September 20, 2013

submissionscomments
angusb··on Ask HN: What does your production machine learning pipeline look like?
Ah interesting! Blind spot in my knowledge right there, thanks for pointing it out
angusb··on Ask HN: What does your production machine learning pipeline look like?
Can't really talk about features on here :(
angusb··on Ask HN: What does your production machine learning pipeline look like?
Sadly not. I'd be totally up for open sourcing if there's clear demand. If you can find it, send me an email at angus@{company_I_work_at}.com

Note that it's very tied down to our use case right now: only compatible with Logistic Regression, and currently it assumes fixed hyperparameters (will change this in future though), assumes a production pipeline of min-max scaling, imputation, then classification.

angusb··on Ask HN: What does your production machine learning pipeline look like?
Can't really talk about features on here. Any smart fraudster should be watching every single thing I say :)

We're using logistic regression not because it performs the best, but because it's the most understandable. When cases get flagged for manual review people need to know exactly what seems dodgy about the account, and with Logistic Regression you can read the exact contribution from each feature to the final fraud probability. Seen as the features mean something real and tangible (unlike in neural nets), this means a manual reviewer immediately knows which aspects of someone's behaviour are out of the ordinary when they get presented with a new case (we have a really nice internal UI for presenting this). This saves several minutes per case which really adds up.

Performance-wise Logistic Regression is good, but it can't automatically learn non-linearities in a feature value and its propensity for fraud, and it can't learn about two features that together should indicate a probability of fraud greater than the sum of its parts* . If this becomes a problem for us we'll start looking into nonlinear models where the inner workings are somewhat communicable to the manual review team.

* You can alter feature definitions manually to capture nonlinearities (e.g. a feature which is "user_has_done_x_and_has_done_y_too", but this is very very manual, and needs to be potentially rewritten/manually re-optimised on every retrain. We don't do this.

angusb··on Ask HN: What does your production machine learning pipeline look like?
For Logistic Regression we find human readable config makes a lot of sense. It's pretty intuitive if there aren't too many features - if the model starts behaving weirdly, we can sometimes track it down to a change in a single feature using this (especially when viewing recent git diffs).
angusb··on Ask HN: What does your production machine learning pipeline look like?
Yep, we made our own. I haven't heard of PMML before - quite cool! What we've made is a bit more readable for what we're using it for though, IMO. Looks like this:

    {
        "intercept": 1.0,
    
        "features": {
            "feature_1": {
                "coefficient": 1.0,
                "range": [0.1, 10.0],
                "mean_feature_score": 1.0,
                "imputation_value": 1.0
            },
            {
                ....
            }
        }
    }
angusb··on Ask HN: What does your production machine learning pipeline look like?
Fraud detection at GoCardless (YC 11). We use the same tech for training and production classification: logistic regression (sklearn).

-------

Training/retraining

- train on an ad-hoc basis, every few months right now moving to more frequently and regularly as we streamline the process

- training done locally, in memory (we're "medium data" so no need for distributed process), using a version-controlled ipython notebook

- we extract model and preprocessing parameters from the model and preprocessors that were fit in the retraining process, dump to a json config file in the production classification repo

-------

Production classification

- we classify activity in our system on a nightly cron*

- as part of the cron we instantiate the model using config dumped from the retraining process. This means the config is fully readable in the git history (amazing for debugging by wider eng. team if something goes wrong)

- classifications and P(fraud) gets sent to the GoCardless core payments service which then decides whether to escalate cases for manual review

-------

* We're a payments company, processing Direct Debit bank-to-bank transfers. Inter-bank Direct Debit payments happen slowly (typically >2 days) so we don't need a live service for classifications.

Quite simple as production ML services go, but it's currently only 2 people working on this (we're hiring!).

angusb··on Hackers Make $5M a Day by Faking 300M Video Views
Can anyone explain why it's the ad buyers that lose out in this case, not the ad network? Surely it makes way more sense for the ad networks to bare the financial responsibility of preventing ad fraud and not the ad buyer? (The ad equivalent of a money-back-guarantee)

How is an ad buyer ever supposed to make an informed decision about how susceptible their chosen ad vendor is to fraud?

angusb··on A Tragic Loss
Are you saying that just based on the fact that they've been developing it for less time than Google, or is there a more in depth case for that somewhere? I would be very interested to read about that if there is.
angusb··on A Tragic Loss
It's not just about Tesla though.

They have to defend autopilot not only to protect the brand but to protect the public's perception of autonomous vehicles in general.

Self driving tech is poised to save many many lives. So from a utilitarian perspective, it's probably justified to take extraordinary measures to make sure reactionary media and public whim doesn't kill it off, however uncomfortable that might seem in the short term.

Whilst this case is incredibly sad (and I don't want to downplay that in any way), if you're trying to minimise the overall amount of fatal crashes, exonerating the tech is the priority (if it is truly not at fault).

angusb··on Terms and conditions word by word
But that's exactly what it should say!
angusb··on Tesla Lining Up Attack on Auto Industry Lobbying
Yes. New light vehicles average fuel consumption 2008:

US: 11 l/100km, UK: 7 l/100km

UK is pretty representative of Europe. (Also the UK improved by 20% since the 2008 figures). US and Canada are unusual in having such widespread domestic "truck" ownership.

http://www.autonews.com/article/20140403/OEM05/140409928?tem...

https://www.gov.uk/government/statistical-data-sets/env01-fu...

angusb··on Tower purifies a million cubic feet of air per hour
This really doesn't seem like a lot. 1m cubic feet is a 30x30x30m cube. In one day it processes 24 of those. Even if there was no such thing as wind or diffusion and you only needed to treat the 30m of air next to the ground, this "neighbourhood" would only be 150m long and 150m wide for it to be cleaned in a day (as they claim).

TBH I don't rally like Wired's reporting on this kind of stuff. To get an overall view of whether this is useful or not, we need to know:

- lifetime of the pollutant

- air changes per day

- whether this type of pollutant is an important one to tackle

...without that we can't know whether this is just an art project or something practically useful.

angusb··on Candy Japan hit with credit card fraud
Yeah, I agree. I guess their argument is that it's difficult to catch every single fraud attempt, and in this case the behaviour was just not picked up by the processor's inbuilt fraud detection systems. Still, it should be the processor's responsibility.

GoCardless (Direct Debit, EU only at the moment) is one company that doesn't charge a fee for chargebacks. They take on all the risk themselves.

angusb··on Tesla Announces $500M Common Stock Offering
There's also the capital cost of installing superchargers to take into account (which only supports your argument further).
angusb··on Under Pressure
That there was a continuous stream of small bubbles (as opposed to one gigantic one) suggests that the failure mode was a leak, not an explosion. Still, I would have been scared.
angusb··on UK unveils plans for huge lagoon power plants stretching miles into the sea
Here's a very readable quantitative discussion about tidal power generation in the UK:

http://www.withouthotair.com/c14/page_81.shtml

angusb··on Citymapper is what happens when you understand user experience
I know it switches from live timetable to "every 10 minutes" style estimates when:

- day bus service ends and night buses begin

- your phone signal dies

Does it do it at other times too?

angusb··on DiscoverTracks – Musicians you love listen to what you like
Did anyone else find "Musicians you love listen to what you like" pretty difficult to digest?

There's probably a more user friendly way to say that. Cool idea though.

angusb··on Solar plant has generated “supercritical” steam
Replying because I don't think the explanations you've got so far are easy enough to read || accurate. Here's my understanding:

Supercritical steam is a special form of steam that can not be described as a gas or a liquid. It's somewhere between the two: molecules aren't bunched together in dense clusters that settle at the bottom of a container (as they are in a liquid), but they also aren't flying all over the place individually in a low density vapour (as they are in a gas).

How's that possible? Water molecules have relatively strong intermolecular attractive forces between neighbouring molecules. They like to stick together, even though there's no permanent connection between them. They are like mini-magnetised marbles. This explains why water has a much higher boiling point than most tri-atomic molecules.

When you increase the temperature of liquid water, the molecules in the liquid vibrate and move around within the liquid, and as you cross the boiling point, the vibration and movement of the molecules is so great that they are able to escape the pull of their attractive interactions with their neighbours en masse. When this happens, the molecules shoot off into the vapour, where there is an (almost) unlimited amount of space for them to shoot around in.

Now consider what happens when you do this at high pressure. High pressure essentially means that there are lots of molecules in the gas phase moving around really quickly. Now, when the temperature gets high enough that molecules have enough energy to overcome their attractive interactions with neighbouring molecules, they leave the pack: but this time with nowhere to go to. The pressure is so high in the 'gas' phase (i.e. there are so many other molecules up there) that they are forced to just bump around where the liquid was but at extremely high speeds. This type of behaviour is pretty difficult to distinguish from the behaviour in the high pressure 'gas' -- in fact, after the system has time to equilibriate, they are exactly the same.

Clearly then, the transition from 'liquid' to 'gas' at this point is pretty much indistinguishable. The liquid may begin to display the molecular kinetic behaviour of a gas, but the density stays the same.

The end result is: When the pressure and temperature is high enough, to onlookers it appears as if the entirety of the fluid is half way between a liquid and a gas, and is stable in that state. That's called a supercritical fluid.

angusb··on HipHop: A "Popcorn Time" for music
I know that there's a kick to be got out of circumventing the draconian rules big music/film industry lobby into law, but what do the writers think about independents that they effectively take down in the same blow? This is a genuine question, not an attack.

I've spent a lot of time studying/writing/playing music and through that have got to personally know many of the most talented and versatile musicians I've ever come across. These people are skilled like Douglas Crockford, John Resig, you name it. But they have to make the assumption that the music they want to do - their own music - will never make any money in a recorded format, forcing them to do wedding gigs during the day instead.

I'm interested to know what people think about this. Do people think that the end (taking power away from big music industry) justifies the loss for those small-time players, or is it something that simply hasn't been considered at all?

Do you have a justification for saying that all music should be free, or is it just that it would be nice if all music was free?

angusb··on Lens Blur in the new Google Camera app
Ah cool, so not a problem. Out of interest how long does the camera-moving step take?
angusb··on Lens Blur in the new Google Camera app
One limitation of this that nobody has mentioned yet is that if you have to pan your cameraphone as you are taking the photos to generate the depth map you will have a harder time composing your photo than you would with a traditional photo. Usually I like to spend a few seconds getting into the best position and framing my photo carefully before taking it. Getting the photo I wanted would therefore be much harder if I had to pan the camera around as I was taking it. I don't have Android so can't test it out... anyone using the app got any views on this?
angusb··on Lens Blur in the new Google Camera app
A couple of other really cool depth-map implementations:

1) The Seene app (iOS app store, free), which creates a depth map and a pseudo-3d model of an environment from a "sweep" of images similar to the image acquisition in the article

2) Google Maps Photo Tours feature (available in areas where lots of touristy photos are taken). This does basically the same as the above but using crowdsourced images from the public.

IMO the latter is the most impressive depth-mapping feat I've seen: the source images are amateur photography from the general public, so they are randomly oriented (and without any gyroscope orientation data!), and uncalibrated for things like exposure, white balance, etc. Seems pretty amazing that Google have managed to make depth maps from that image set.

angusb··on Norwegian skydiver nearly struck by meteorite
I just spent a while looking up the maths on this and was really surprised to find that the terminal velocity of a spherical rock (diameter 10cm, density 2.5g/cm^3) at this altitude is remarkably close to 300 km/h (I got 340).

Aside: up until now I had imagined that if a skydiver were to ever drop a rock of that size (or a dense piece of equipment like a DSLR) during freefall they would never be able to catch it, but as the terminal velocity of a skydiver in "dart" position is about 320km/h that's not the case. Pretty cool.

angusb··on Startup Idea: Robot Cars
Sadly you're right. But my point grapples with a larger issue too: should a startup launch if they aren't confident about the safety of their product?
angusb··on Startup Idea: Robot Cars
I liked this article, and this is only a small point in the context of how interesting the rest of it is, but it would be reckless to say that you need to exceed only 1M driving hours accident free to be better than humans. Of course this would lead to a lower empirical accidents:mile ratio for your new tech, but you still wouldn't have enough data to be confident that your accidents:mile ratio fairly represents the chances of the new tech causing crashes. I'm not well read enough on p-values/confidence intervals/chi-squared tests to explain why, so maybe someone who is can explain this if there's enough interest. Basically someone needs to get all Evan Miller on this (e.g. http://www.evanmiller.org/tesla-fires.html )
angusb··on Too many electric cars, not enough workplace chargers
I think the idea was that instead of simultaneously charging several cars, it charges them one by one (round robin). You just have multiple connectors to stop employees having to leave the office to change the connectors over when one car finishes charging. But I guess it would require smart-ifying the charger base so it could decide when to send power to each of the cars.
angusb··on Nissan Sells 100,000 LEAFs, Captures 48% Of Worldwide Electric Vehicle Market
Useful point of reference: the VW golf (Europe's most popular car) sells around 500k/yr, so the Leaf's sales are about 7% of that. Seems significant.
← PreviousPage 2 of 2