State found $117k in double payments through data analytics
routefifty.com
routefifty.com
I'm curious of it's a net win, though, once all the costs (and benefits) are considered.
This was 117K was found in transactions during 8 out of 12 months in the year. The article doesn't specify the amount of money transacted during that time. It doesn't give us the total amount of transactions either.
It gives an estimate of 18 million in the total year so the closest estimate is 117k out of 12 million. Which is ~1% and seems like a win win. State employees and folks should feel good that they aren't fucking up very often and the analysts should feel good because functionally this number should go to 0. Operationally no idea what it would cost to maintain though. cough
However the pilot program only analyzed transactions from Jan to Sept. For all state agencies (which seems impressive from an aggregation and governance stand point). With in that time period it found 117k in duplicates, so assuming that dollar/transactions are at a constant rate through out the year (they aren't).
117,000/(18,000,000 * (8/12)) ~ .975 % loss in state to vendor(?) transactions.
Put another way, with in the time period they studied assuming 12 million dollars in transactions were made, they can potentially recuperate roughly 1% of the total value transacted.
For comparison there's a Kaggle dataset of 2 days of European credit card transactions from 2013. It contains 284k transactions of which 0.147% are fraudulent.
This 1000x difference can either be good since it means the state is mostly paying appropriately or bad since the detection is not as good as the credit card company's.
god damnit you're right. whoops
To be fair, I wouldn't be surprised if there are additional benefits beyond the $117k. For example:
- Reducing the attack surface of fraudulent invoicing.
- Retaining staff/skill that are useful in other projects not discussed in the article.
- Seeding awareness/interest for similar projects in other departments of the state government.
I've met certified accountants whose daily job was to take an excel spread sheet, add 1% to every cell with a number, then return to the client. It's not automated but not because it can't be automated.
As a disclaimer, I have recently become enchanted with automating accounting work due to hearing about the problems accounting functions face. I'm a tad biased.
Large banks, automated trading shops and assimilated have automation, because it's their jobs. Other things don't.
Recruiting chartered accountant is easy and cheap. Recruiting developers that can deliver is not as easy or as cheap. Bear in mind that it's not a tech environment so no source control in place, no servers to deploy to, won't be easy to execute.
Give me a lever long enough.
[1] https://a16z.com/2011/08/20/why-software-is-eating-the-world...
[2] https://techcrunch.com/2016/06/07/software-is-eating-the-wor...
The whole business was based low cost labor taking whole days to evaluate stuff manually and write risk reports. Crazy thing is they were hiring loads of developers and data scientists, but the core of the business is all manual.
There were multiple people, of varying skill, willing and able to automate significant chunks of this with introductory R/Python/VBA skills. It’s not a cool shiny front end app, but a few scripts here and there could save an FTE easily. But politics prevents these things from being adopted.
Cost and scarcity of tech talent is not a real issue for light officework. Servers, Source Control, etc. are missing the picture. You just need people with light programming capability and a little bit more agency and voila they’ll be 10x as productive as everyone else.
Well, that's just dumb. We can do better. Those accountants you met can do better!
I'll take 10% of savings
I'll take 9% );
8%?
Back to 10%
>Other double payments made by mistake included times that the state received multiple invoices.
I would assume each paid invoice was a new transaction ID. The real problem seems like there are two invoices being paid. Not that there are two transactions (which seems like just a symptom). It's possible that each invoice even has it's own ID.
What seemed like a very small and simple problem initially revealed itself to be massive problem that even a team of 35 struggled to maintain. The purchase history for any given item spanned multiple systems with completely different topologies of data glued together by, literally, tens of thousands of lines of SQL.
The project had been around for ~17 years when I worked on it and while data was landed in a final format where a query like the above could have been done, I wouldn't bet any serious money that the calculations were correct.
Some five or six multi-million dollar rewrites had been attempted, but could never be done.
Not to say Ohio's system is that this level, but I doubt there's a giant table sitting somewhere that such a simple query could be applied to.
I like to think that somewhere in all that SQL was an Office Space-esque line of SQL syphoning off a few cents per transaction, but was never able to find one :)
"Don't be snarky."