How We reduced our Google Maps API cost
blog.cityflo.com
blog.cityflo.com
Their strategy was to have a pool of API keys attached to new accounts that would take advantage of the Google Maps API free tier, and monitor its usage. As the free tier usage would run out, the system would roll over to a new API key automatically.
Wrote that one up in big red marker in my report...
Going with a pool of free keys will be much more dependable, even if somehow more complicated to manage and easier to break.
You can then keep fooling them by creating new accounts or changing IPs (assuming your usage doesn't have clear patterns they could look at).
But such events would be clearly disruptive for the business. Works for a POC, but if your business has actual customers, it's a terrible solution.
Let's hope there will be alternatives to Google provided traffic data. For now they seemed to monopolized it by offering it for free while losing money to discourage competition.
What happened to actually trying to solve problems with programming?
Interpolation is one solution. Caching is another. Temporal analysis. Put everything together.
You don't need to query the magic Google box for every small update you make (and they might get that info from the transit providers, which given my experience are not that great sometimes).
For example, if you ascertain that a bus is more than 2 minutes late, switch to polling that bus more often until it makes it to its next stop. And then switch back to linear interpolation once it gets to that stop. But you'll pay a little bit more for the added accuracy.
Morale of the story: if you want to get high-resolution real-time data, you (and your customers) have to pay for it, as that shit ain't easy.
Hardly "low accuracy".
The key change is modelling the problem as one of routes rather than journeys. Since a route can be "polled" at a certain resolution & used regardless of the number of journeys on it.
> This approach made the API calls independent of the number of vehicles and dependent only on our stops, which helped us in scaling up our fleet with no additional cost.
Unless I'm misunderstanding, they aren't polling these inter-stop coordinates
Won't Google just ban the server's IP?
It didn't work out, and it still baffles me how anyone thought it could.
Last time I asked this question [0] on a different story [1], the responses I got were that it definitely is wire fraud, but this is so mind-blowing that I would like to ask again to confirm.
> 3.3 Restrictions.
> Customer will not, and will not allow third parties under its control to: ... (d) create multiple Applications, Accounts, or Projects to simulate or act as a single Application, Account, or Project (respectively) or otherwise access the Services in a manner intended to avoid incurring Fees or exceed usage limits or quotas;
A snippet from the article: A federal court in Washington, DC, has ruled that violating a website's terms of service isn't a crime under the Computer Fraud and Abuse Act, America's primary anti-hacking law. The lawsuit was initiated by a group of academics and journalists with the support of the American Civil Liberties Union.
Did you think my use of the word services was somehow related to terms thereof, because no. So to quote wikipedia because it was the first that came up when I googled "Theft of services is the legal term for a crime which is committed when a person obtains valuable services — as opposed to goods — by deception, force, threat or other unlawful means, i.e., without lawfully compensating the provider for these services", do you see how someone might argue that changing out the api key could be seen as a form of deception?
I am not saying that I would think it right that someone bring this to criminal court (I figured I better put that out there as even the simplest of comments can be misunderstood, so who knows what several paragraphs together might lead to), I am not saying that they would even win, I am not saying anyone who did it would be doing so for the purest of motives. But I am saying it does seem something like theft of services by using deception.
on edit: the theft and new api keys refers several ancestors back to this anecdote "Their strategy was to have a pool of API keys attached to new accounts that would take advantage of the Google Maps API free tier, and monitor its usage. As the free tier usage would run out, the system would roll over to a new API key automatically."
I.e. Google knew there was rampant abuse by people like this (this example is not the only thing like this I have heard of...) so Google fixed the glitch and in the process ruined it for all the people genuinely using the service's free tier.
This is why we can't have nice things :) I guess we are lucky that Google didn't decide to just cut their losses and close the whole shebang down - that would be a shame as Google maps is really useful IMO.
We were going to not display suburb data because of cost. In the end I found a creative commons placename database (geonames.org). For placenames with >500 people it's ~10MB of data and that covers the entire planet (surprisingly small). I then wrote a KD-Tree based library to look it up the nearest point in this table extremely efficiently (log(N) time).
I'll admit i haven't updated or maintained it. The server running it has been chugging along well though >5years later. https://github.com/AReallyGoodName/OfflineReverseGeocode
At a previous employer I tried to convince my managers to let me do this for months. They always balked. Their loss.
The above solution is a copy paste-able set of classes with zero dependencies that will output the nearest place. Sometimes a single purpose solution is perfect and i'm really not kidding when i say i haven't maintained or even looked at it in over 5 years yet it's still running fine as part of a larger application.
I am not sure if I'm missing some big, but it really seems like there wouldn't be an ongoing need for querying google's APIs in this case. Finding the boundaries of the neighborhoods and especially labeling stops should be easy enough to do if it would save so much money.
Of course, this is for a use case where you have similar routes every day, this allows you to really tune the Kalman filters.
0: https://github.com/Project-OSRM/osrm-backend/wiki/Traffic
The real problem with this approach is that you will never know when real time is conflicting with historical unless you're calling the API constantly anyway
I mean, I'm not disputing it, I just hadn't realized we were already there.
> (c) No Creating Content From Google Maps Content. Customer will not create content based on Google Maps Content.
> (d) No Re-Creating Google Products or Features. Customer will not use the Services to create a product or service with features that are substantially similar to or that re-create the features of another Google product or service.
What would be a surprise is if Google forced on your GPS to gather your location in other contexts.
https://www.theverge.com/2017/11/21/16684818/google-location...
It doesn't take too much tin foil to think that android may be phoning home when it detects an SSID, the physical location of which is already known.
It is unclear to me whether their current practice is in harmony with the Google Maps TOS. For some places there are also open traffic data sources: https://github.com/graphhopper/open-traffic-collection
But I'm not sure if it applies to this use case.
If you don't want features like real-time traffic awareness, it's worth investigating the open source tooling. It can save a LOT of money.
Until now there wasn't a need since Google provided the data for free.
But everything changed and I bet there are lots of people who need the data but can't afford paying Google.
The only question is how you gather the data since you don't have Google maps and Waze?
You can try to provide a Waze alternative, but how would you convince people to use it?
Thanks for the excellent work, though!
https://github.com/Project-OSRM/osrm-backend/
(Reboot discussion at https://github.com/Project-OSRM/osrm-backend/issues/5209)
The principal downside is that the routing graph takes a lot of time and memory to prepare; runtime RAM usage is also high, though not so much. I think there's some potential for reducing its memory footprint.
Valhalla builds on the older A* algorithm, so it's not so fast (or memory-hungry), though it does make some use of hierarchies to shorten query time. Graphhopper is another, featureful routing engine designed for use with OSM data (written in Java).
https://github.com/Project-OSRM/osrm-backend/wiki/Traffic
It works quite well.
[0]: https://github.com/opentripplanner/OpenTripPlanner
[1]: https://github.com/opentripplanner/OpenTripPlanner/pull/2077
[2]: https://github.com/opentripplanner/OpenTripPlanner/pull/2698
https://cloud.google.com/maps-platform/terms
(a) No Scraping. Customer will not export, extract, or otherwise scrape Google Maps Content for use outside the Services. For example, Customer will not: (i) pre-fetch, index, store, reshare, or rehost Google Maps Content outside the services; (ii) bulk download Google Maps tiles, Street View images, geocodes, directions, distance matrix results, roads information, places information, elevation values, and time zone details; (iii) copy and save business names, addresses, or user reviews; or (iv) use Google Maps Content with text-to-speech services.
(b) No Caching. Customer will not cache Google Maps Content except as expressly permitted under the Maps Service Specific Terms.
"Customer can temporarily cache latitude (lat) and longitude (lng) values from the Directions API for up to 30 consecutive calendar days, after which Customer must delete the cached latitude and longitude values."
https://www.eventsofa.de/campus/migrating-away-from-google-m...
Shoutout to Stadia Maps CEO who was always very very helpful.
I'm cofounder of Stadia Maps, and I'm happy to field any questions folks may have.
However - it is also important to _contribute_ funds to OpenStreetMaps - for cityflo and for us.
Donations: https://wiki.osmfoundation.org/wiki/Donate
Individual membership: https://wiki.osmfoundation.org/wiki/Membership
For organizations: https://welcome.openstreetmap.org/how-to-give-back/
The OSM Foundation in general: https://wiki.osmfoundation.org/wiki/Main_Page
OpenStreetMap itself is just the base map, not a provider of traffic data or routing.
I still think location data collection for the public good should be run by someone who is not Google / Facebook / some other company that doesn't already track everything about you. I'd be happy to send my real-time location to a company like Mozilla or some mapping company (just in the tiny netherlands I know of AND, Geodan, OsmAnd... I'd be fine if any of them organize this).
We also have some open data via the government, based on detection loops in the road, but if I remember correctly it's only (or mainly) on highways.
I still can't believe Google does not have a GUI to both style a map AND add markers to it, but they have GUI for those things separately.
But now that their Google bill is only ~$50/day, it might not be worth building their own prediction system.
I used an app which predicts public transportation arrival time based on historic data and it mistakes most often than not, sometimes by a lot.
I guess you are reacting to private smug conversations with people who eulogize ML but are clueless about what they are talking about. Usually the amount of conviction they add to their comments and advice is inversely proportional to their knowledge. I am tired of those too, to the point of being allergic in spite of being an ML practitioner myself.
The parent comment could have been worded better but I think you will agree that there is scope for using ML or statistical modeling of some kind to optimize the use case both in API cost and accuracy. Making some minimal number of API calls would very likely be part of such a system.
IME congestion is also often modal, there's probably reductions in sampling effort you could make by noticing the patterns in how commuters route, whether school is on break, and what the latest roadworks are.
Disclaimer: I work in Google. But I also worked a bit on effects of cache placement in my PhD.
Then, I might be perceived as biased due to that employment. I might even be actually biased. Better disclose that up front.
Example here on how to decode: https://github.com/geodav-tech/decode-google-maps-polyline
Besides, I suspect that Google knows exactly what their data is worth and keep a very close eye on the price/demand model.
Just because something is subjectively expensive doesn't make it objectively overpriced.
Perhaps they have some margin formula for “buying” data from their internal sources.
It seems to me that they are value pricing and since they are close to a monopoly on this service they are testing the price elasticity.
I just wish they would either augment their data with OSM or augment OSM with their data: now, with proprietary data, the data was already old when we got the car, and the fixes I contributed to OSM are obviously not in there.
Instead of contracting some middle man for inferior map data, they could have used completely free (as in beer) data for the regions where OSM is more up to date and complete than any commercial provider (empirically, this is most of the land mass on earth plus countries with a lot of mapping enthusiasts and/or free government data, like Germany and the Netherlands).
Also given they seem to own the buses, they could give the drivers a rooted Android phone and get the data by reading the memory (or network traffic after reading the TLS key from memory) of their running instance of Google Maps that they are using to drive, which should be undetectable by Google.
Having worked for a company that provided transit software I am well aware that GTFS-realtime data isn’t always accurate (especially the ETAs). But it seems like combining their ETA estimate with yours and the scheduled arrival time would be a reasonable approach. Or if you just started building a model of the average speeds of the bus on each link using the GTFS-r speed and location data you could probably build up a pretty good historical model of the traffic flow between 2 stops at any given point during the day.
Yes, I am aware that scheduled times in transit can be meaningless. There are a lot of factors that contribute to that from traffic to poor schedule planning. They can however be a useful baseline estimate.
One thing to look at given you would have everyday data is to see if there is a pattern (if not already did it). Then you can suggest commuters to take the bus now and save commute time v taking after 15 mins.
For example, if you have 2 buses on the route `A-B-C-D-E-F-G-H`, using the same value of T(F-H) for the bus starting off at A and another one already at D, might not be quite right?
When google maps gives me an ETA, I assume they account for it (or their ML model does from the vast troves of past data).
I still haven't found a viable free replacement, unfortunately. I'm just doing this for personal projects, so no big budget.
First, is the fuild ETA absolutely necessary? Surely the bus has a schedule. Is it on time, or not. If not how late might it be? That's not the same as ETA. Certainly, they've collected and keep collecting enough data to infer such things. That is, 2 mins late to Stop B translates to what at Stop F and stop N.
The consumer doesn't need ETA per se. They need to know if they're going to be on time to their destination or not.
Like putting mirror next to an elevator to shorten the perceived wait; there might be other opportunities to solve this problem.
If you're interested in such things, this book from a year or so ago was intriguing.
Rory Sutherland
Alchemy: The Dark Art and Curious Science of Creating Magic in Brands, Business, and Life
4.6 out of 5 stars
https://www.amazon.com/Alchemy-Curious-Science-Creating-Busi...