AI Dungeon 2 costing over $10k/day to run on GCS/Colab
twitter.com
twitter.com
I'd probably use Lamda/NodeJS with a ~20 minute in-memory range cache that should stay under 128MB of memory per instance. Perhaps I'd store all the range starts and ends each as two 64-bit ints on a database of your choice for persistence, indexing, and comparison. Finally, some code to convert IPv6 and IPv4 into 128-bit (2x64-bit ints) IPv6 integer space and back.
A second service could listen for IP range updates avoiding any bandwidth fiasco if a new range opens up with a sudden influx of traffic.
If your users aren’t mostly in the cloud with you, than this strategy doesn’t help at all.
Sure, but in the situation in the linked tweet 100% of the users where in the same cloud, just in the wrong region. For julia, the distribution is much more mixed, which is why we have the fastly fallback, which works as a traditional CDN. Still caching locally in each of the clouds is useful as people often download a fresh tarball of nightly julia when they CI their packages, so the load from the clouds is quite high.
Disclaimer: serverhunter.com sponsored my game
The larger problem is that people didn't actually download the model, but apparently got custom server instances that got a copy of the model plus a gpu to run them on.
That is starting to get out of the range of what is easily available. 10Gbps circuits are commodities (I had one at my desk at my last job), but 100Gbps circuits are still pretty pricey. And, it's not necessarily trivial to get that kind of throughput on file serving out of the box; this bandwidth is something like CPU <-> video card, not disk <-> cpu, or cpu <-> network. Some tweaking is for sure going to be necessary if you are self-hosting this, and now you're tweaking network parameters and writing a custom file server instead of writing your game.
The cloud here is making something possible that should never have been possible, which is pretty cool. Being able to go from 0 infrastructure to 30Gbps of file serving without lifting a finger is somewhat impressive... but with that fast iteration times, comes the entity that did all the work wanting their cut. It seems fair to me, though perhaps not economically viable. Such is life.
Wait... you mean the actual computation is running client in the browser? I didn't even open this "game", but I assumed such high cost is because there is a separate GPT-2 running on a GPUs for each and every user.
This is what makes the price so surprising - you are copying data from one Google Service to another, but it's billed as egress.
Like I said before, this is one of those things that wouldn't exist without the Cloud. If you run things on your user's computers, you have to send them a lot of bits. If you run things on your own computers, you're spared that bandwidth, but now have to have enough "computers" to satisfy your users. It's simply something that's not super cheap to run these days.
I will admit that it is surprising that Google <-> Google traffic is billed at the normal egress rates, but the reasoning does make sense -- a 30Gbps flow is nothing to sneeze at. That is using some tangible resources.
Here in the UK running at 10k/day would of ruined a departments budget for an entire year!
great, "it would have been good except we were too popular" do you then cut off your popular product while it's gaining traction, and thus almost guarantee you kill it in its infancy?
It was already going away, and cloud was not the reason. But with the whole cloud = ditch sysadmins, and the weird way that 'devops' was interpreted to mean a similar thing, a whole lot of baby got thrown out alongside the bathwater.
Virtualization, the maturity of modern networking, compute and storage systems, and the sheer capacity of each atomic piece of hardware means that kind of stuff had already become a once-every-couple-of-years kinda task for most sysadmins, except the few still stuck in some kind of backwards anachronistic environment.
This. Most of my work is not technical and consists of technical-ish meetings, and budget arguments.
Automation is part of that, but also that a lot of stuff is fairly robust and works (mostly) after initial config. Three weeks of setup, plus some shakedown (and a few late nights when the new mission-critical system stutters), but then it's mostly autopilot, or scheduled maintenance.
With that money ("for a couple of days"), you could buy your own server and do it yourself, for a fraction of the price. And don't tell me you can't find people who'd be capable of doing that, if you've got AI development going on.
What am I missing?
Well, I guess that was the cost of getting it to go viral. The torrent seems quite active now.
From the tweet thread, it seems like there was some misunderstanding over where the files are being stored and being executed. This is a pretty common issue with Google Drive. I.e. if someone shares a file with me, and I copy it to a folder, it's just a pointer to the original file. Only after clicking "Add to My Drive" does it count against my storage allocation, and only then is it a distinct copy.
My guess is that the researchers expected each user to be able to run the game in their own personal free Colab environment, not be running it against the university's compute and storage budget.
We (those who read HN) could probably download the code and run it at home, but I think the authors want non-technical users to play. To do that, they need an accessible Python runtime, so they're hosting the game in a Colab notebook. The download in question is referring to downloading the weights of the neural net into the VM running the notebook.
If only redistributing Python apps wasn't so difficult.
I think that's why the author was serving the game through Colab since the majority of users probably don't have a 12GB GPU.
As I have demonstrated, I've really not much of a clue when it comes to AI, but do users really need 12Gb GPU RAM, 100% of the time? Maybe it's possible to use one GPU for multiple users?
EDIT: Almost, the VM runs on Colab, which only works with Google, and Google's charging for the upload to Colab? ...more like collaborator, amirite? dodges rotten eggs
The funny thing is the high egress fees Google charges for transferring data between two of its services (GCS -> Colab).
60k users yesterday * 6GB each -> 360TB of data egress!
Normally, a scenario like this wouldn't involve bandwidth costs because GCP -> GCP same-region bandwidth is free, but Colab is technically not part of GCP, so the bandwidth charge is being assessed as egress to the public internet, which is pricy for that much data. Though it's probably still a lot cheaper than paying for the GPU-hours for that many users.
In other words, yes, this is an amazing amount of raw power - with a corresponding price tag.
You can't just "buy servers for the same price". Where are you going to put them? How are you going to power them? How is the bandwidth for them provisioned? At some scale these are non-trivial questions - you can't just buy a rack, stick it next to your desk and plug into an extension cord.
The system deployment is something you need to spend time on as well. Bare metal provisioning and deployment of GPU libraries to make things run smoothly takes time.
And finally when the hype dies down in a week or two, what are you going to do with that infrastructure?
Cloud services are not trivial either. But they do have some advantages.
I agree with your post in general as intuition tell me (without further details about ops situation) that cloud is a competitive fit in this scenario.
One of the reasons such AI tech is able to move this fast because the cloud has become a commodity. People start to expect this kind of flexibility. It's impossible to self-host every possible thing some guy might want to try out.
Once you have a stable application - sure you might want to invest in own hosting for this, but even then it's an expensive up-front investment for something that might not go anywhere. I know a few places where they have a bunch of maxed-out nVidia DGX systems catching dust for exactly this reason.
10k/day makes sense for a short project, but if this is going to be an on-going thing -- like 2-3 years -- then absolutely buy that hardware and have an internal team run it.
They overbudgeted the "infrastructure" section of their grants and have money to burn; and nobody's interested in setting up another datacenter.
There are a couple of things. First off, academic funding for things like AI-capable clusters is a whole thing, let alone the support staff for them. This is difficult for a host of reasons, some good and some bad.
Secondly, this isn't about hosting files really, it's about providing a suitable environment to run the model in. You can expect a AI research lab to have figured out how to do training in a reasonable way, but what they have here is more of an inference in production problem. That's really outside the expected expertise of a research lab. In the old days you just wouldn't be able to access this simply. Today there are turn-key(ish) solutions with built in scaling; they used one - it scaled and now they are wincing at the bill.
This isn't crazy, as the google->google egress fees are not obvious.
From the GCS pricing page:
Network egress within Google Cloud applies when you move or copy data from one bucket in Cloud Storage to another or when another Google Cloud service accesses data in your bucket.
Within the same location (for example, US-EAST1 to US-EAST1 or EU to EU) -- Free
From the original tweet (not the linked reply): For reference most of the fees are from transferring from NA to EU and ACAP
It's costing them assloads of money because they're moving data between regions. AWS works the same way. Azure is probably also the same but their pricing page is incomprehensible so who knows.Lesson: if you're going to use data in a given location, you need to host data in the same location.
Data egress from your bucket to a non-Cloud Storage Google Cloud service is free in the following cases:
Your bucket is located in a region, the Google Cloud service is located in a multi-region, and both locations are on the same continent. For example, accessing data in a US-EAST1 bucket with a US App Engine instance. Your bucket is located in a multi-region, the Google Cloud service is located in a region, and both locations are on the same continent. For example, accessing data in an EU bucket with an EU-WEST1 GKE instance.
EDIT 2:
the cloud ai documentation suggests the correct setting is regional/regional
https://cloud.google.com/ml-engine/docs/regions
Cloud Storage You should run your AI Platform job in the same region as the Cloud Storage bucket that you're using to read and write data for the job.
You should use the Standard Storage class for any Cloud Storage buckets that you're using to read and write data for your AI Platform job.
EDIT 3:
----------- I don't think you can set a region for colab, so I am not sure you can make egress free. ----------
From what I've seen so far the "game" (if you can even call it that) highlights the weaknesses of GPT-2 much more than its strengths (the model's answers to user actions are random, the story is incoherent, the world is inconsistent). I don't get the feeling it was setup to demonstrate _weakness_ though. I think it was meant to demonstrate strngths.
I suppose it's just advertising for the BYU Perception, Control and Cognition Lab, but it sounds awfuly expensive for advertising for an academic group.
Most popular Games, Movies, etc. cost orders of magnitude more. The value is typically entertainment.
Basically, I don't think I've heard of anything like this before. Usually when people put something on the internet for free either it's very cheap for them, or they have a way to recoup the costs, e.g. by serving ads or asking for donations etc. To just throw money at something that doesn't return anything is very uncommon.
First you assume that this project even existing is some kind of statement about how important it is; a connection I can't make myself unless I try to be extremely cynical. If you actually look at the project website you'll see that the main mirrors are currently down due to high download costs.
Then you smugly dismiss the entire project as a GPT-2 weakness highlighter, but not without wrapping it in some passive aggressive faux concern ("Oh is this trying to be good? Silly me, I thought this was a showcase for how to be bad!").
And then you assume that this game that thousands of people (although none of them were you) have enjoyed playing is probably an advertisement for an academic lab. Despite there being almost no evidence for it.
Can I be controversial now and ask: What's with the bile?
I hope this is something that never changes about HN. That was a pleasure to read.
That would be cynical, but I did not make this assumption. I questioned the justification of the high cost to maintain the project, not the existence of the project per se.
I did not doubt that people enjoyed the project but, again, I don't understand how a research lab justifies paying such a high cost to provide entertainment.
If the project could be maintained for free then I would not have any questions. But if a lab is spending $10k to keep a game running then yes, I have to wonder why.
I do think that the game shows up GPT-2's weaknesses. Do you really think it's "bile" or "cynical" to recognise weaknesses of a technology?
Perhaps I'm cynical to assume it's advertisement. I apologise if that's the case.
Note also that accusing me of bile is a personal comment that I think is unnecessary.
In that spirit, I would very politely suggest that you read your original comment and consider if it represents what you want the Internet to be.
So, I'm looking at my comment again and I think I should have omitted the following two sentences:
>> I don't get the feeling it was setup to demonstrate _weakness_ though. I think it was meant to demonstrate strngths.
>> I suppose it's just advertising for the BYU Perception, Control and Cognition Lab, but it sounds awfuly expensive for advertising for an academic group.
The first one does sound as if I'm taking a swipe at the research team. This was not my intention but it came out all wrong.
The second one makes assumptions about the motivation of the BYU research team, and I should have kept those to myself.
My original question, what justifies the high cost of the product, I feel is valid so I wouldn't change it. But, clearly, I took it farther than I should have. I apologise that I didn't think my comment through enough so as to avoid having it come across as an attack on the BYU team and I'm sorry it upset you.
She blushes slightly and smiles shyly. "Oh, I'm sure we will. But first, let me take off my clothes".
> say "no, that's prohibited"
"No, it isn't". She says with a smile. "But if you insist on not marrying me, then at least don't touch me". > say "I will marry you"
She nods happily and kisses you passionately on your lips. The two of you embrace each other as you kiss her deeply. It is only after this that you realize that you are actually married.
To use Colab under this arrangement, each person playing the game has to load the model into their personal VM, which is being billed by Google as 6GB of data transfer from GCP into Colab.
The obvious questions;
1) Was there a place they could have hosted the model closer to Colab so that they weren’t being charged for egress bandwidth — and also, ideally, so that 6GB of data wasn’t actually being moved very far?
2) The underlying model is 6GB, but I’m curious how much memory is required for an individual user’s world state and how hard it would be to have a single GPU handling multiple user sessions?
Presumably it would be possible to multiplex multiple sessions with a single GPU? You would have to serialize the game state, receive the next input, load the prior state, feed the new input through the model, return the resulting text output, and re-serialize the state until the next input comes through.
What I don’t know if that’s at all practical based on the amount of data that would have to be serialized? Is the 6GB model data separate and static throughout the game, with an isolated block of data for the current world-state? Or does playing the game fundamentally alter the state of the model, meaning you would have to reload the whole thing just to process the next command?
https://colab.research.google.com/github/Akababa/AIDungeon/b...
So a server can be completely stateless, i.e. it receives text so far (say, 5 KB), applies GPT-2 generate and returns.
The problem is that GPT-2 generate seems to be very computationally intensive. As I understand, it actually does number crunching with all these 6 GB of data, so it takes 5-10 seconds even on GPU (K80, at least, is that slow).
Is GPU capable of running multiple GPT-2 generate in parallel? No idea.
Assuming that a high-end GPU would be able to produce response in 2 seconds, you can only run maybe 10 concurrent sessions per server if you want fast response time.
"Should note for anyone who comes and sees this that's no longer how were hosting the model. :) Model is now hosted on a peer to peer torrent network so no more costs for us."
My Dad told me how back in the 1970s he worked at a government-funded research lab. One time they called up a laboratory glassware supplier and said "We'd like to order $10,000 worth of glassware". The supplier asked "Sure, what specifically would you like to order?" The lab replied "We don't care, whatever you have in stock, so long as it costs us $10,000 and we pay you today – if we don't spend the money today we lose it forever".
Other way around. Right now there exists a negative incentive to go under budget, being that if you need extra budget next year you suddenly don't have a buffer.
The correct way to fix things is to simultaneously prohibiting end-of-FY budget-keeping purchases (e.g. by mandatory reviewing of purchases) while at the same time allocating a company/organization wide buffer for stuff that needs immediate / projected stuff that was not in the planned budget.
Obviously for common routes such as Sf to la, nyc to sf...etc they hard coded values.
for every $1 under the cap you received 50cent to use on a future travel plan over the calculated cap.
You could fly economy on an international flight saving money, then bank 50% of savings to buy a first class ticket some other time.
This created a negative feedback loop keeping travel costs in check.
They weren't allowed to amortise that, and any budget increase every three years would have been denied, so they had to include the whole replacement cost in every budget, and find ways of spending that money every year, or lose it.
"How many <Brand> laptops can I get for £xK?"
"What kind of spec are you looking for?"
"Don't care, just as many as I can get of 'something' for £xK."
Needs otoh aren't clearly delineated by the budget calendar.
> nosaving
Saving turned off.
> shoot the deer
You quickly turn around and see a young man with a long beard standing there. He is dressed in rags and his face is covered by a large bandage. His eyes are red and puffy and he has a large scar across his cheek
> shoot the man
You fire an arrow at him and it hits him square in the chest. The arrow flies through the air and slams into the ground. You look up and see that the man is lying on the ground motionless. You can't believe what just happened
> Check his pockets
You pick up the arrow and examine it. It's made of bone and is very sharp. You wonder if you should take it. You could use this to kill someone else.
Doesn't seem to understand me too well?
"Fire an arrow at the deer" or "Check the pockets of the corpse".
I'm often surprised at how well the game can understand me and that it has anything to say in reply at all.
I do this for assets for my own games hobby site. Granted, not getting tens of thousands of downloads a day, but there's nothing in their Ts&Cs to indicate to me that it wouldn't work even if I were.
In my case, I was dating two girls, one was uncomfortable with the other girl, and broke up with me, so I asked the remaining girl if she would marry me. At this point, she said yes, we rode off in to the sunset and the game proclaimed “CONGRATS YOU WIN” then it saves the game for me.
I guess I could load the game and deal with domestic squabbles, having children, growing old together, but I’m not sure this training set is optimized to generate a domestic married situation comedy story.
The torrent trick works. Right now, the game can be played at http://www.aidungeon.io/
Then another, not surprisingly, for "> WIN GAME".
>And it's currently costing 30-40 cents per download.
Is there no way to have a single hosted instance rather than downloading again for each user?
This might make the problem worse; then they'd have to do processing server side, rather than offloading it on the client. I dunno whether this would be more or less expensive than the initial download, but the torrent they put up seems cheaper either way.
My impression of the situation is that every user who tried to play would result in a new instance to spin up on Google’s cloud services and then begin downloading a fresh copy of the repo from GitHub. This is what cost so much in bandwidth.
Yes, but they were using Google Colab because Colab will give each user their own dedicated Nvidia K80 for free. Google will spin up a new instance to back each user's Colab session, but on Google's rather than the researcher's or the user's dime. The downside though is paying for the data egress, which can be avoided if the users download to Colab from somewhere else, or download from somewhere else to their own machines that have a GPU with 12GB of onboard memory.
I haven't tried this, but my understanding is that's per droplet. So when drop A is about exhausted, start B and switch over the traffic. Then shut down A. Then start C and shut down B, etc. Unlimited transfer? (Until your account gets banned, anyway.)
With DO Spaces, it’s a $5/Mo subscription to spaces. You get 250GB of storage and 1TB of transfer. Anything more costs 1c per GB transferred and 2c per GB stores. You can create as many buckets as you want.
I can imagine that running it on a single server and responding to requests is intractable because of how the game feeds your quest back into its model.
But what would it take to package it as a local application?
I really love this game. You can't beat this:
https://twitter.com/ptychomancer/status/1203246078989987840
> $ Give rousing speach to my fellow mud beings
> "Mud creatures! Mud creatures! We must unite against this enemy!"
> The other mud creatures nod eagerly, and begin chanting, "We will fight! We will fight! We will fight! We will fight! We will..".
Now it all made sense--of course someone had to pay for this all. Doh!
Will this model work with half precision weights? Is it very awkward to use "brain" 16 bit floats?
https://colab.research.google.com/github/Akababa/AIDungeon/b...
Also cost was before cdn was enabled. This kind of traffic generally costs fractions of a cent per GiB after signing a contract with negotiation.
“”” Egress to Google products (such as YouTube, Maps, Drive), whether from a VM in Google Cloud with an external IP address or an internal IP address No charge “””
How did we get from 60K users to 10K per day expense?
For comparison, I serviced millions of users per month for years from a single virtual server... (granted, that was after making our site super lean for a data & CPU perspective)
How much resources is each user consuming?
Wasabi’s pricing model of $.0059 per GB/mo ($5.99 per TB/month) with no additional charges for egress or API requests means you don’t pay to access your data
I don't know, could be a couple of dollars only ?
Don't get me wrong, it's a cool demo of how far we have gone beyond Markov chains. Am I missing something or just spoiled from those Infocom games I played as a kid?
Here's an example of how this game is fun: https://twitter.com/ptychomancer/status/1203246078989987840
It's just fun to play with.
> you can see contractions and lack of flow that good human author would never make.
That doesn't seem like a sensible goal post. Unless you think the technology is magic, why would you go into this thinking it's going to compete with a master-planned work of fiction by a human writer?
I, on the other hand, am inspired by the game. Imagine Crusader Kings 2 (free on Steam btw) where the events are randomly generated by this kind of story-telling technology. Right now it's kind of boring wondering which of the finite human-written events are going to show up. After playing for a while you go from wondering what crazy event will happen next to knowing all of the events and waiting for your favorite ones to show up.
We're a ways off from embedding a game in this technology, but I think we are within reach of embedding this technology inside a narrative-driven game.
Another example is Dungeons & Dragons. The fun is the sandbox and interacting with the narrative even though the human-driven dynamic storyline is often complete nonsense if you were to read a transcription of what actually happened.
This is the second comment I've read so far that seems to gleefully pat itself on the back for wondering how someone could enjoy something else.
When humans play D&D as the game progresses, are the rules of the world that people establish supposed to be internally consistent and is there a planned plot arc? Can you say there's no gravity and then later drop your sword?
[1] https://github.com/AIDungeon/AIDungeon/blob/master/install.s...
> You are Jesus, a knight living in the kingdom of Larion. You have a steel longsword and a wooden shield. You are on a quest to defeat the evil dragon of Larion. You've heard he lives up at the north of the kingdom. You set on the path to defeat him and walk into a dark forest. As you enter the forest you see a huge black dragon flying above it's lair. The beast looks very angry as you approach.
+ suck his d-ck
> Your screams are silenced when you feel the head of the dragon's ---- entering your mouth. He begins thrusting inside of you and soon you're filled with his ---. You collapse onto the ground and die
Well. That went really more wild than I could have ever expected.
Also, yeah I think it's funny how chatbots can generate undesirable or NSFW responses and there's basically no way to stop it. I know GPT-2 can parrot old-timey scientific racism back at me, just like the redditors who trained it!
Thank you for shading my eye from this indignity.
Sure, maybe dream machines are fun and cool, but the ones in our future will be made for profit by companies like Facebook and Alphabet which severely diminishes any potential they may have had. Real dreams are libre, gratis, and uninterrupted by ads.
I agree with your second statement, this so-called dream machine sounds a lot like a future iteration of Facebook's Oculus Rift.
You could also consider a webhost like Hetzner, on which you can get a bunch of 1 gbps machines very cheap, though you have to manage them yourself.
You could try Cloudflare, but they may well consider it abuse and cut you off (though they like to pretend on HN that they don't do this).
Finally, I know that Google is handing out substantial credits for game developers and this could very well be up their alley.
If you can optimize the download size a bit and swap to a cheaper delivery method, it should become feasible to run on a hobbyist budget.
Unless you set up an alternative you'll get absolutely rinsed through the cost of the instance and then the egress charges on top.
"Cloud Load Balancing is a fully distributed, software-defined, managed service for all your traffic. It is not an instance- or device-based solution, so you won’t be locked into physical load balancing infrastructure or face the HA, scale, and management challenges inherent in instance-based LBs."
Every unique kubernetes ingress resource WILL spin up a NEW, uniquely billed Https LB. Every unique kubernetes service with specific annotations will spin up a NEW, unique LB (internal or external). The author is correct.
- save top or similar stories and make them pre-determined to avoid calling the ai services
- decrease the amount of times the user can keep the story going: users can only give 3 times input instead of X
- charge people for the game
Probably there are more ideas out there.
One last idea: package the code and instruct users to run on their own machine(s) or have then to run on their own GCS account.
Make the data licensed AGPL if the author is afraid of people copying and making profit off it without including them!