The data trove is fairly unique, and valuable in being the only of its kind, but we don't need anywhere near instant access to most of it.
Hypothetical: you create a logging service for users to send all their log data to you. You promise 365 days of archives, but 30 days of data accessible at any time. You create a lifecycle rule on your S3 bucket to automatically archive data to Glacier 30 days after creation. On the 31st day, your user decides they want to look at an old log. They click the big Download button. You display a message saying they'll get an email from you when that data is ready to download.
http://docs.aws.amazon.com/cli/latest/reference/glacier/init...
From there you can submit a "DescribeJob" request, with that Job ID as the parameter, and the Glacier service responds with the state of the job.
Once the job is marked as complete, you submit a "GetJobOutput" request with that Job ID. That response is the archive body. (similar to how you'd do a GET request from S3).
You've got 24 hours to start the download of the archive before you'll have to repeat the entire InitiateJob->GetJobOutput cycle again.
I think through the API you do not leave the connection open, you check with whatever frequency you want and when it's ready, the response will include the temporary location on S3 for the file.