Warning: $14k BigQuery charge in 2 hours
discuss.httparchive.org
discuss.httparchive.org
If you're running a business and you have lawyers, then fair enough — just play the game. But for individuals, it seems crazy that so many of us accept this sort of thing. Good luck contesting the charge with your credit card company when you already agreed to a contract that said Google could bill you thousands of dollars and then you used thousands of dollars worth of their service.
Big cloud providers are not your friend. They do not care if they destroy the lives of you and your family, unless it's happening so often that it's making mainstream news.
My advice is to go and delete your cloud accounts, and only use services that offer hard spending caps, and ideally prepaid accounts.
Maybe this doesn't leave many options. Oh well. Maybe if you can't afford big lawyers then you also can't afford the risks of using big cloud.
In some jurisdictions I think that reduces the legitimacy of their claim that you actually owe them money.
EDIT: Even better, focus on the examples where Google "forgave" the debt; you could argue that those examples prove that Google knows it's at least partly their fault.
I know that if it did work it would change the opportunity cost of forgiving debt in these cases dramatically
How often is one of these accidental debts created? How often do customers just pay up because it's small enough that it's not worth fighting? How often does AWS (or Google or whoever) decide whether to forgive the debt based on PR damage control rather than the legitimacy of the debt? Jeez I hope someone leaks those numbers one day.
It reminds me of all those horror stories of hospital visits in the USA, where the first bill you receive is just a test to see if they can squeeze that much out of you, but if you know what you're doing or just can't pay then the actual bill is way lower. It's all just yucky.
If big cloud providers couldn't selectively choose which of these debts to enforce, I bet there would be a media shitstorm and then they would suddenly discover that it's not all thaaaaat hard to implement real time billing and hard caps after all.
I think we (the developer community) need to start pushing back against this abuse, it's getting out of control.
The thing that bothers me the most is I caught this $14k charge b/c I'm a small fry and that money matters to me. How many big accounts just wouldn't notice that? I can't help but think a very non-trivial % of all cloud revenue is just obscure fees that nobody notices - engineers doing the engineering, accounts receivable pays the bills, and the cloud providers get fat.
Can not recommend privacy.com enough.
Will they pursue? Do they have enough info to purse? Who knows, but they can if they want to.
I emailed support, and here's what I got back:
> Hi, $firstname. I've been reviewing your dispute and wanted to touch base with you to explain what happened.
> It appears that the disputed charge is a "force post" by the merchant. This happens when a merchant cannot collect funds for a transaction after repeated attempts and completes the transaction without an authorization — it's literally an unauthorized transaction that's against payment card network rules. It's a pretty sneaky move used by some merchants, and unfortunately, it's not something Privacy can block.
Have you found a site that does "block" this? Did you communicate with your credit card company about this? I am wondering
Edit: and the dark Lord surely reserves a particularly unpleasant circle of hell for loan officers who encourage borrowers to consider a 5-1 variant rate because "we know rates will fall next year."
I think that might finally allow you to pay for the New York Times on your own terms and not worry about their hounds sniffing you down.
What's interesting is that they seem to be glossing over the truth. It's not unauthorized, per se, it's using a prior authorization code. And it's intended for processing offline transactions. It seems like 'force' is an industry term and a bit hyperbolic when used in lay discussion.
More discussion here: https://www.tidalcommerce.com/learn/force-sale-credit-card
For small charges they might just give up because it's not worth it, but when dealing with a $14k bill you should assume that they will at the very least hand the debt off to a collections agency if you try to just ignore it.
I used Amazon EC2 instances for years and I always felt in control. There were never any surprises. I knew even in the worst case situation I would be okay because I had faith in the Amazon support. With Google I felt insecure. I never played with any of Google cloud services since then.
Amazon's customer first policy is really true. They try their absolute best to make sure there are no surprises to a great extent. Even the UI is very intuitive.
Which part of customer first drove their egress fee policies?
Egress is basically all outbound traffic. The fee was always this. Dont act shocked when it doesn't go down when you have buyers remorse.
A personal example would be that we reserved an instance based on information given by our AWS account manager. Said instance turned out to have issues linked to my original question to the account manager who answered incorrectly.
The reserved instance team then refused to refund us but also refused to tell how much they would prorate if we were to upgrade instead.
Basically a protection racket.
I think the HTTP Archive team could set something in that regard.
PS: When I was an instructor for some cloud training in AWS, the first 2 hours were only to set up billing and budgets to avoid any kind of situation like this. No one would start training without all those locks in place in the first place.
[1] - https://github.com/HTTPArchive/httparchive.org/blob/main/doc... [2] - https://cloud.google.com/bigquery/docs/best-practices-costs
> Note: The size of the tables you query are important because BigQuery is billed based on the number of processed data. There is 1TB of processed data included in the free tier, so running a full scan query on one of the larger tables can easily eat up your quota. This is where it becomes important to design queries that process only the data you wish to explore
Could this be a bigger warning? Sure.
Is something a scam just because they don't explain the general implications of entering your payment information to a usage-billed product? Not really.
Estimating the total cost of a query is obviously fraught, but from the UI and other comments it sounds like BigQuery knows up front how much data a query will require, and there's at least a minimum cost per TB, so the UI could just say "this will cost at least $X" but the UI has a very basic "this will process X PB of data" text. So they're charging by TB but showing the usage in PB which is a) a 1000x smaller number, and b) visually similar to "TB".
It's very hard to see that as anything other than "design to obscure cost" given that there's no reason to not say "this will cost $X" when the cost is per TB, even if they don't the pricing is per TB but they're showing PB, the checkbox and the textual description are smaller that other text on the page, and there's no ability to specify a cost cap.
Last week I ran a script on BigQuery for historical HTTP Archive data and was billed $14,000 by Google Cloud with zero warning whatsoever, and they won’t remove the fee.
This official website should be updated to warn people Google is apparently now hosting this dataset to make money. I don’t think that was the original mission, but that’s what it is today, there’s basically zero customer support, and you can lose $14k in the blink of an eye.
Academics, especially grad students, need to be aware of this before they give a credit card number to Google. In fact, I’d caution against using this dataset whatsoever with this new business model attached.
https://cloud.google.com/bigquery/docs/best-practices-costs
Estimate query costs
BigQuery provides various methods to estimate cost:
Use the query dry run option to estimate costs before running a query using the on-demand pricing model. Calculate the number of bytes processed by various types of query. Get the monthly cost based on projected usage by using the Google Cloud Pricing Calculator.
If the latter... I'm not sure that it's explicitly against the rules, but coopting a name of something as your handle just to complain about it is in poor taste and probably should be.
This was the last comment in TFA, so it seems like they just used it because it was the topic...
Public datasets are hosted for free by Google (Amazon has a similar program) to take the burden off public projects.
You didn't pay for the data, you paid for the query you ran against it.
Also note that anyone can make a dataset available for public use, where they pay the storage and the consumer pays the compute. The official Google datasets are just curated and maintained by Google itself.
What it is, roughly, is a publicly-accessible data supercomputer. If you lost $14k in a blink of the eye, then I would think you consumed at least $4k of Google's actual resources -- maybe $7k. Maybe more. That thing can move some serious data, and you apparently moved around over 2PB.
Google bears some significant responsibility for not making the cost transparent to you, it's true. But on the the other hand, don't they bear some significant credit for making such an awesome power available to a lowly peon with a credit card?
If GCP would return the query cost in the API and show it directly in the console when you run a query, it would be much easier for their users but unfortunately, it's not Google's interest for obvious reasons.
The console shows you this number (in very small letters) after you have entered the query but before you press go. In the on-demand billing model, which is what you were using, you can multiply this number by $6.25 to understand your query cost, exactly.
It's a design that's hostile to new customers, I agree. But it is comprehensible.
https://medium.com/@steffenjanbrouwer/how-to-set-a-hard-paym...
This was one of the main selling points for all portfolio companies of the group to adopt AWS in their digital transformation projects.
I would love to kick the tires on some AWS stuff, but the threat of unlimited ruin is not worth it. Sure, maybe the gods would take pity on me and wipe the debt, but far easier to just run with someone who caps costs. My toy project can gladly go down if the alternative is a huge unexpected bill.
The OP is probably a good person with strong interest in data science and building projects.
If it'd be "oh here's your $500 charge, upgrade your quota for more, 'ok fair enough, I did a mistake'", but $14k is not ok without explicit quota upgrade.
Unfortunately, if the customer has written their applications in such a way that they're effectively locked to the platform... they won't have much choice until they can dis-entangle themselves.
After that though, yeah. ;)
To be clear: I'd be very in favor of the major cloud providers having a "DO NOT! DO NOT! use this for production mode and your content could be deleted at any time if you screw up. But I suspect most people wouldn't use that."
I can't say that's for certain what it is. I just know a hallmark of any business with recurring charges that are otherwise incomprehensible is so they can hit you with the charge after the fact, and you have little recourse to avoid paying it without a ton of work for yourself or your team.
https://cloud.google.com/billing/docs/how-to/notify#cap_disa...
Notice that their "solution" is to tell you how if you want you can spin up effectively your own custom service to watch spend and if it goes over some threshold delete the entire project[0] after some delay. This is the malicious compliance version of letting you add a limit.
[0] At least, that's how I interpret "This example removes Cloud Billing from your project, shutting down all resources. Resources might not shut down gracefully, and might be irretrievably deleted. There is no graceful recovery if you disable Cloud Billing. You can re-enable Cloud Billing, but there is no guarantee of service recovery and manual configuration is required."
https://cloud.google.com/bigquery/docs/custom-quotas
> Custom quota is proactive, so you can't run an 11 TB query if you have a 10 TB quota. Creating a custom quota on query data lets you control costs at the project level or at the user level.
All of those technical reasons aside though, the commercial reason is obvious - people's mistakes and overages are a great source of revenue and profit. Companies refund the times where it'd be enough to lose the customer, or when it hits HN, but they make more money every time someone pays up. They have no incentive to fix it. It's part of the business model.
Ever notice how your "1GB" data plan sometimes lets you use 5GB if you happen to be roaming in another country and downloading something fast over 5G...? Same reason.
They do not want it, that is the only reason.
Google doesn't want the bad press. Most real companies would prefer to have a big bill when their product surges in popularity than have unexpected downtime at the worst time.
That doesn't preclude it being an option.
Sure, this guy fat-fingering $10k sounds amazing.
But GCP deals with businesses paying for years of service. Multi million dollar deals are common.
Google and AWS and the like could give a flying fuck about anything under $100k.
Would you like Amazon to delete all your files, disks and backups once you hit your limit?
Also Static IPs, Load balancers, DNS zones?
To learn how to use it you’ll have to try it. Learning by doing. Trial and error.
And no you may not use blanks while learning use of footgun. No training wheels. No precautions. Has to be live - full send unlimited risk.
Actually - here are some safety squints (billing alerts) - just to give you some illusion of control.
Good luck! Yours truly big cloud
First response on the OG link covers this with a screenshot: the size of the query is previewed beforehand, and you have to check a checkbox to acknowledge it. (I dare say listing it in PB instead of $$$ is still a scummy move, etc. - but they do resolve about half of your concerns right there)
> Version 4.3 patch 30 of rn’s Pnews.SH (September 5, 1986, published to support the new top-level groups) introduced the “thousands of machines” message:
> > This program posts news to thousands of machines throughout the entire civilized world. You message will cost the net hundreds if not thousands of dollars to send everywhere. Please be sure you know what you are doing.
-- https://retrocomputing.stackexchange.com/questions/14763/wha...
AWS charges $27/hour for a server with 3TB of memory. Enough to run the queries in memory.
A regex query on response_bodies would churn through 2.5TB of data every time it's run.
Again, you could fit the whole dataset in memory in an EC2 instance and do your thing.
A simple "running queries over the whole dataset can cause significant costs due to the size of the dataset" should be enough. And I think that's a valid and fair point.
The whole part of accusing Google should just be ignored.
https://github.com/HTTPArchive/httparchive.org/blob/main/doc...
I don't think that's a proper warning on costs.
I don't know. Google could trivially solve this problem by imposing an opt-out warning on potentially expensive queries.
"It looks like your query might cost $14k. Are you sure?"
But money.
This comment kind of suggests that you do not understand how BigQuery bills. The archive pays for the storage, but you have to pay for the queries. You would also have had to attach a billing account to run those queries. Running BigQuery searches is not free.
Expensive lesson, but on the surface this one appears to be your error.
https://arstechnica.com/gadgets/2009/04/users-62000-data-bil...
One of the bigger issues is they charged my card before I literally had any notice what the bill was - it wasn't even in the dashboard yet. I would have terminated the script ASAP had I gotten *any* warning.
> Note: BigQuery has a free tier that you can use to get started without enabling billing. At the time of this writing, the free tier allows 10GB of storage and 1TB of data processing per month. Google also provides a $300 credit for new accounts.
> Note: The size of the tables you query are important because BigQuery is billed based on the number of processed data. There is 1TB of processed data included in the free tier, so running a full scan query on one of the larger tables can easily eat up your quota. This is where it becomes important to design queries that process only the data you wish to explore
> When we look at the results of this, you can see how much data was processed during this query. Writing efficient queries limits the number of bytes processed - which is helpful since that's how BigQuery is billed. Note: There is 1TB free per month
https://github.com/HTTPArchive/httparchive.org/blob/main/doc...
It should make this available.
To be fair: I'm sure they don't provide this limit to make money, because this is a rare case, but to avoid the far more common case of established business going offline because someone forgot to update a limit.
Sure, a crosswalk may have an extensive system to warn drivers of pedestrians, but that doesn't change the fact a driver hits a pedestrian there at least once a month. It only has to happen once to ruin someone's life.
For cloud providers, the obvious solution is hard budget limits. Ask people to set a hard budget limit before they get the opportunity to drown themselves in debt. Free up some workload off of the support team in the process.
Hard budget limits change the process to avoid these charges almost entirely. Warnings only inform a few people that they're aware the process lands people in debt, and to please use the broken process correctly to avoid the severe financial consequences.
Lesson #2 - if you select * on a gigundous dataset, make sure it's on your employers bill.
If your employer is a small business and you nuke them from orbit by doing this a few times, it's unlikely to go well.
Similarly that checkbox being a tiny part of the UI, and not allowing people to set up cost limits on a query (or not having them at the account level), does seem very much like an "encourage people to overspend" UX. I'm sure "overspend to the level of a $14k bill to an individual" is not intended, but that's a reasonably predictable occasional outcome for this design.
So on the one hand, yes they did click a checkbox saying they were aware of the amount of data being processed, but OTOH the UI seems to be specifically designed to encourage this kind of mistake.
Edit: since I didn't hear back from you, I've consed a 'not' onto the username 'httparchive'. If you prefer a different name, feel free to contact us at hn@ycombinator.com.
Do you know if the dataset is public? We should just offer a cheap alternative and ditch BigQuery.
Honestly I was already concerned when it was taking more than 5 minutes to return a result.
Once I saw how slow it was I found my error, not querying the sample dataset that was a fraction of the size, to make sure my filtering worked.
> This website makes it seem like this “public” dataset is for the community to use, but it is instead a for-profit money maker for Google Cloud and you can lose tens of thousands of dollars.
you didn't understand what you were doing. HA's datasets are public and free. It is not a "for-profit money maker for Google Cloud". Sorry, sucks for you but blaming the restaurant when you bit off more of the steak than you can chew is not how this works.
Their UI clearly has all the info needed in order to put guard rails in place (aka big scary warning dialog in red), as it's already giving a non-obvious warning about the expected data usage.
Blaming users for this seems like a bastard act. Talk about causing further reputational damage... :( :( :(
FWIW I use BigQuery a lot and as a rough guide I assume about 1c per GB scanned. So if I query a dataset that's 1TB, that's about $10. If the same data were stored on a relational db, the same query would take about a day (or at least a good part of a day). Because BigQuery returns a result so quickly (e.g. <1 minute) it can be easy to miss the insane amount of work it did to get there. So I could see someone accidentally putting that ~1min (but 1TB!) query into a loop or something, and boom, there's your $15k bill. Accidents happen.
Also FWIW, I've found although the big 3 cloud's pricing is tricky (since there are so many services), I find them much better than the PaaS built on top of the big 3 clouds. My suspicion is that the PaaS's have a strong incentive to obscure their pricing because customers can typically see what their costs are (e.g. if they buy some compute from AWS at $0.16/hr and sell it for $1.40/hr, that can be seen as a bit of a rip, hence they try to obscure it). But I think the big 3 are not too bad at this practice. It really bugs me when anyone deliberately obscures their prices, and it's often an indicator of more shady practices to come.
If you instead filter out the rows you are interested in (e.g. the particular "few sites" by their URL) and put that in a new table, querying the resulting, tiny table will be very cheap.
[1]: https://cloud.google.com/bigquery/docs/partitioned-tables
I’m not familiar with your use case or BigQuery but in Redshift I’d just do a COPY to a local table from S3 then do a CREATE TABLE AS SELECT with some logic to split those URLs for your purpose.
You might even be able to do it all in one step with Spectrum.
always use plain ec2 spot and s3.
lots of smaller instances. fewer larger instances. single massive instance. whatever. fancy sql thingy, awk and grep, or whatever else.
do your data processing with ephemeral spot priced compute and persist as little data as possible to s3.
$2-5/hour gets an insane amount of ec2 spot. egress aside, no surprise bill is possible.
empathy for op though. not a fun day. just a bump in the road though, keep on trucking!
One place I worked at had a table with 100 billion rows. And some other tables as well. If a manager asked for an ad-hoc query, it was 5 minutes of writing a SQL query including JOINs (which didn't need to worry about which fields were indexed etc. e.g. you could write WHERE then a regex), and $15 and 5 minutes later I'd have the answer. Apparently 100s of VMs were started and stopped to answer that query, but it all happened automatically, at very low cost.
"Running this query will touch 2PB of data so you'll be charged $20000" for it..
There is no hard billing cut, people normally tell me that no one wants that because customers and service disruption but I want that
I want a hard budget cutoff
A VM with max settings doing crypto mining can be very expensive and adding more of them is easy.
And you know what else is super hidden costly? Log digestion for metrics. Egress. Auto scaling .
You can create costs for someone by just downloading their assets _a lot_
Partitioning matters y'all!
So $14k would be about 2,240 TiB.
I wonder what sort of partitioning and clustering is used for the tables.
I know the explanations and justifications for it, but for personal use a service where I can't put a hard limit on usage is simply not acceptable for me. It's just not worth the risk.
well, other than the three market leaders (GCP, AWS and Azure)
This is victim blaming.
And my understanding is that almost none of the way of setting limits are actual hard limits, but only alerts and some hacked-together emergency abort scripts. Correct me if I'm wrong, but can you actually limit the cost robustly for services that spend that much money in an hour or so? Doesn't help much if I get an email about it and read it two hours afterwards.
Until that point, they are just an individual who got screwed by disguised billing practices.
I have very limited cloud experience, but I did make a mistake that lead to a rather slow but constant cost. The amount was small enough to not be relevant in a professional context, but the memorable part was that I could not pinpoint the source easily with the AWS tools and my limited understanding of them. The categories and labels were too broad, and it took a bit until I figured out what went wrong. There are certainly better tools to investigate this, but I didn't know them. In the end it was simply luck that the mistake still fell into an area of insignificant amounts of money, but it could have easily been significantly more if a few parameters had been different for the same mistake
Excellent idea. Please describe how to create an account on AWS or GCP that is not allowed to spend more than $100/mo. Since it is "a really easy fix" and "takes almost no time" it should be easy to explain, right?
That's probably enough for 99% of people, and if you're highly motivated, you could make that trigger an SNS notification that trips a circuit breaker.
While you can footgun yourself with hard limits I tend to think that learners/hobbyists should, in general, be able to access at least many services with an ironclad guarantee that they can't be billed for over a certain monthly amount or a total number.
I'm much more inclined to shrug if a startup screws themselves over with a hard spending limit than if a student screws themselves over because of a lack of one.
a catastrophic mistake might result in the company going bust and all the pain associated with that
but shouldn't lose you your home (assuming you acted properly, the project using the cloud provider has to be in the aims of the company, etc etc etc)
speak to a lawyer
I’ve long wondered what I can do with an LLC to protect me from debts like this but I don’t know how to get more information about it. Particularly as I’d be the sole owner I don’t really understand what the llc does/doesn’t do.
If you had just 1000$ (and made a few hundred a year) is it worth doing?
obviously the advice will cost you but may be the best couple of hundred bucks you ever spent
it's not magic though, you do have to conduct yourself properly as a company (be that non-profit or otherwise)
if you're experimenting with side project that you think has potential commercial value then this is why limited liability as a concept was invented
(as always, speak to lawyer)
The short, short version is: You have to have a reason for the LLC that isn't just "contain some risks". Something like "this is a legal entity for my side project bilombinaboloa.com, that I'm hoping will one day become a company and make me Rich" will work, "I pay my expenses via this and take my income directly" will not.
The idea that Google would give away BigQuery for free if only people would write better SQL queries.
Edit: rereading, I think this is actually for non-interactive scripts, in which case yes it should just cancel the query
Edit 2: https://news.ycombinator.com/item?id=39447499 was kind enough to point out that the resource-based version of this might actually exist, which is nice
You can set the size limit for individual queries. Plus the custom quotas and everything.
Part of the problem is that the OP wrote a script with a loop. So say you set the limit to 50 GiB per query, but then write a script that runs a 49 GiB query 1000 times...
That type of batch process should be designed much more carefully to consider costs.
Are you sure?
The article doesn't say anything about a loop, and the estimated usage by the Google responder makes it seem like the cost is from a single "SELECT *".
> I was doing historical evaluation for a few sites, so I was running a query for each month going back to 2016 for each site. I've done this before with no real issues, and if I knew the charges were rapidly exploding I'd have halted the script immediately - but instead it ran for 2 hours and the first notice I got was the CC charge.
So looks like a loop of ((6 * 12) + 2) * #sites iterations with a full table scan every time.
If nothing else, it can be an example in my SQL 101 course.
`SELECT * FROM super_wide_table_with_lots_of_text WHERE NOT filter_on_partitions_or_clusters`
Select * is dangerous because it's a column store. You really need to look at the schema and select only the things you want. And when exploring the data it's important to use sane limits and pull from a single partition.
SELECT page, url, payload FROM `{table}` WHERE page like '%{site_domain}/%' AND url like '%[EXAMPLE.COM]%'
---
There's no LIMIT on it b/c I actually need all the results.
I'd say!