HNHacker News
TopNewBestAskShowJobs

obulpathi

647 karma · joined July 31, 2011

Born and brought up in India, did my Ph.D. in Big Data from the University of Florida.
submissionscomments
obulpathi··on Serverless Takes DevOps to the Next Level
> You don't emit logs?

I emit logs. I mentioned that in my above comment as well.

> So your code assumes that the network/storage/DB are fault-free?

No, it does not assume that the network/storage/DB are fault-free, completely. It assumes they are fault free with the SLA limits provided by your Cloud provider. The software is written in a way that enables the platform to know about errors when then happen and remediate them. Like, when you build a website, you make the app serving layer stateless and choose a backend datastore that is replicated across regions and is highly available (like Cloud Datastore or Spanner). If your platform detects that a disk failed and your app is returning errors, then that instance is killed and an another instance is brought up. Very similar mechanisms exist for the above-mentioned data services as well for auto scaling, sharding, ...

If your Cloud provider can not guarantee the SLA's or shows a green sign even when the service is down, IMO they are not competent enough.

obulpathi··on Serverless Takes DevOps to the Next Level
To me, the definition of Serverless is this: Developer writes code and the management of resources for running that code is taken care of by a Serverless platform. As a developer, I don't have to deal with servers, disks, networks, logs, metrics, ...
obulpathi··on Serverless Takes DevOps to the Next Level
> This kind of sounds like you've never managed a production system, sorry.

Don't be sorry. Instead, challange me with specific problems you have and I will show you how to build solutions.

> how do you process and manage lifecycle for an exabyte of classified/sensitive data

The infrastructure for storing/processing data is in place. If you doubt me, please go ahead and give the above-mentioned tech stack a try and let me know if you still have any problems. Or give me a detailed description of what you have to do and I will tell you how to do it.

> People who talk about "data platforms" in terms of size are like people who judge coder productivity in lines of code

Let me be more clear. I know how to build the infrastructure for that scales of data. I am not saying that I will write the code for processing data. That depends on what kind of data you got and what you want to do with it.

> I guarantee your 10pb "data platform" will choke and die the minute someone shows up with 10gb of historical data and you have to get the timezones right.

See my above comment for the answer.

obulpathi··on Serverless Takes DevOps to the Next Level
With Google Cloud, Serverless has been a reality from 2015! And welcome to Serverless to those coming from AWS.

Here is how Google Cloud achieves serverless. Ingesting Data: PubSub. Scales to millions of messages instantaneously, no need to spin up and spin down the capacity/shards (like as in Kinesis), exposes RESTful interface for ingesting messages from anywhere (web/mobile ... ). If you want a high-performance interface, you get gRPC as well. With RESTful interface to consuming messages, you can connect PubSub to anything you want or trigger Cloud Functions / write to storage / ingest to Stackdriver (Monitoring system) using managed services.

Querying Data: Big Query. No need to spin up a single server. Just write your query in SQL and watch the magic of a thousand servers being spun up to serve the query in a fraction of second and process a petabyte of data in a couple of minutes. btw ... you can ingest data into Big Query in real-time and analyze results within in seconds.

Process Data using Dataflow: For workflows that are more complex than SQL queries to ones that need to be running continuously on streaming data. Dataflow is the serverless version of Spark. No more running our of memory errors, much fewer hassles with hotkeys, no more manual performance tuning of buffers, no more cleaning up log folders. If you are using Spark, but have not tried Dataflow, you are missing some serious magic.

Google Cloud ML for Machine Learning: Cloud ML is a hosted solution for running TensorFlow jobs. No more manual hyperparameter optimization, no more spinning up GPUs, no more scaling up the cluster size. It's all taken care for you.

Container Engine: Hosted version of Kubernetes. I am sure, everyone is aware of what K8 is and its capabilities.

I am from a Data / Analytics / ML background. The above was my reality of Serverless since 2015. Google Cloud has serverless options for Web (App Engine) / Mobile (Firebase), and other purposes as well.

Having done PhD in Cloud Computing and Big Data and I don't find Infrastructure/Big Data a sexy problem anymore. To a large extent, it's a solved problem. Building a data platform with a team of 3 people that can handle 10+ petabytes of data is easy. What is on the horizon and unsolved yet, is AI!

Would love to see if there are any other better/compelling Serverless options.

obulpathi··on RethinkDB joins the Linux Foundation: What Happens Next
This is awesome news! Good to see Cloud Native Foundation growing to address the needs in Cloud Computing space.
obulpathi··on Making Google Data Studio Free for Everyone
Thanks! And yes, as you mentioned, the problem is that it works only for inserts. I use Datastore as my primary database. So requesting for supporting Datastore from Datastudio.
obulpathi··on Making Google Data Studio Free for Everyone
That's exactly my issue. I have a crawler that collects data and stores into Datastore. I need to periodically scoop up the snapshots and export them into BigQuery, then run the visualizations. I want my visualization to be live and also don't want the overhead of scooping the data.
obulpathi··on Making Google Data Studio Free for Everyone
Awesome! Datastore connector, please!! Cloud Datastore is the ideal NoOps and Serverless storage solution for projects from small to large. It would be great if you can build a connector for it.
obulpathi··on Snap commits $2B over 5 years for Google Cloud infrastructure
Lets not assume that Google buys storage from Seagate. Google makes its own hardware for many things (networking, TensorFlow custom ASIC). Also I don't think that have that many number of customers.
obulpathi··on Snap Inc. S-1
By the time the other company builds infrastructure, some other company would have built the product that Snap wants to build.
obulpathi··on Snap Inc. S-1
If you are doing just one or two things at a massive scale (Dropbox storing files), it makes sense. But, we are talking about Snap Inc, the infrastructure, and services it needs. A single high-speed intercontinental fiber cable costs 400 Million $ (https://techcrunch.com/2016/10/12/google-and-facebook-are-bu...). To build a service like CLoudML (Hosted TensorFlow), you need to write software (TensorFlow), build the ASIC Chips by manufacturing / by partnering with a chip manufacturing company like Intel / Samsung / Qualcomm, you need custom networking software and hardware to scale massively for data workloads (Andromeda), years of research to build these tools, you need to create your own programming language (Go and Dart) to be able to scale your development to these levels, you need to build data centers in multiple regions, you need to build DeeLearning models for cooling your data centers (https://deepmind.com/blog/deepmind-ai-reduces-google-data-ce...), ...
obulpathi··on Snap commits $2B over 5 years for Google Cloud infrastructure
I don't know.
obulpathi··on Snap commits $2B over 5 years for Google Cloud infrastructure
Nop. By user I mean per person, not company. Now take a deep breath and think about the scale of Google Cloud.
obulpathi··on Snap commits $2B over 5 years for Google Cloud infrastructure
I guess that speaks about the scale at which they are operating in Google Cloud. Diane Green mentioned in a recent conference that one of their healthcare customers collect about 2 PB / user. Lot of companies struggle with managing / extracting value from data. Thats where the bottle neck is usually. If they have capability to handle more data, overtime their services evolve to collect, store and process more data. Once Big Data became reality, many companies started collecting orders of magnitude more data. With Google Cloud its easy to handle petabytes of data. That enables large scale computing companies on Google Cloud. (Think of driverless cars / genomics / large scale machine learning / social networks / ... )
obulpathi··on Snap Inc. S-1
Expert on Cloud Computing here. Short answer is "No". Here is why. Google Cloud has infrastructure that scales to petabytes of data and millions of users. Google primary uses this infrastructure for storing, processing and communicating the Internet. Add the services like Pub/Sub, Dataflow, BigQuery, TensorFlow & CloudML and things like security, communication backbone, ... its near impossible to build Google Cloud or even few critical components like the ones mentioned above with 2 billion dollars. Also, if they focus on building infrastructure, it might slow them down significantly. They are better off building their chat platform rather than building the infrastructure.
obulpathi··on Microsoft earnings blow past estimates in every category, beats Street
> I'm not at all, that example was pulled at random as an example.

Then give me an example of what you are looking for and I will give you a solution

> My company certainly likes to do things in a unique way.

And every company has a unique way of doing things. That is why they provided all the tools you want and let you combine them to fit your custom needs. Instead, if they were to provide turnkey solutions, you would not use them because you have unique needs.

> It isn't like Google is the only company that can build a message queue or a load balancer. Nor are they the only game in town selling them.

If there are any better solutions, please let me know. Support is first class on Google Cloud. That is one of the reasons why Spotify choose Google Cloud over AWS.

> "Flat rate" does not describe the monstrosity that is that spreadsheet.

Google Cloud pricing structure is far easier to work with, as you don't need reservations/upfront payments and at the same time leaves room for scaling up and down elastically. The amount you pay is simply based on usage. For Pub/Sub its the number of messages, for BigQuery its the amount of data that is queried. I would love to see a simpler pricing structure. Would you care to share the details if you know of any?

obulpathi··on Microsoft earnings blow past estimates in every category, beats Street
If you are looking for a solution for mobile analytics: https://firebase.google.com/docs/analytics/

If you want a customizable solution, you connect the tools. Pub/Sub for streaming, Dataflow for processing data, Big Query for Data Analytics (Warehousing), Datalab is your development environment. Its hard for a single solution to serve 100% of customer needs. That why they suggest tools. Care to share what a better technology stack looks like? I would love to know if there are better alternatives for services like Google App Engine, Google Cloud Load Balancer, BigQuery, Pub/Sub, Dataflow, TensorFlow Cloud ML, Container Engine (Kubernetes), Stack Driver, ...

Speaking about cost, its all flat rate, with usage discounts (the more you use, the more discount you get) and there are no upfront payments / reservations / locking / hourly billings (like AWS)

obulpathi··on Microsoft earnings blow past estimates in every category, beats Street
> "We need to do X" and they say "Here is product Y that connects to product Z, but you need to write your own connector." And each product has it's own pricing structure (one by CPU usage, one with flat fees, one based on throughput) so working out total cost is a nightmare. Then you have to hire around those techs, project manage the build of the infrastructure, etc. At that point you start to ask "Why are we using them again?"

This is a feature(?) of AWS. Google Cloud provides fully built solutions for majority of use cases. AWS provides nuts and bolts and tell you to integrate them.

obulpathi··on Microsoft earnings blow past estimates in every category, beats Street
Google Cloud and AWS are pretty much equivalent in terms of feature parity: https://cloud.google.com/free-trial/docs/map-aws-google-clou.... There are few (10% os services maybe?) that are not available on other platform. Also, the way AWS counts service is by nuts and bolts (AWS believes in providing fundamental components, not higher level frameworks) vs Google Cloud counts things by Cars / Trains / Submarines levels. That why people tend to think that Google Cloud provides fewer services. (As an example, imagine how many number of AWS components are needed to perform the function of App Engine. May be 10?)
obulpathi··on Microsoft earnings blow past estimates in every category, beats Street
I worked with both Google Cloud and AWS. Having read Gartners Magic Quadrants on Cloud, I really question Gartners competency on understanding what Cloud really means. There is no question that Google Cloud is years ahead of AWS (ease of use, scalability, performance, vertical integration, networks, security ... ), but Gartner reports that AWS is leader. Bottomline: Don't trust what others say. Try out various cloud providers and decide for yourself, before diving head down and suffering with pain everyday.
obulpathi··on Amazon Web Services’ secret weapon: Its custom-made hardware and network
I am very happy to let you know that perspective is informed and correct too! Google Cloud is far better than AWS. Especially in things like Ease of use, scalability, big data, machine learning, NoOps, pricing Google Cloud beats AWS by miles margin.
obulpathi··on Microsoft Azure in Plain English
Agree 100%. For instances, you have standard, high memory and high CPU and custom. Super easy to make sense of rather than crawling through webpages to understand what each type of instance mean. For example: AWS has r, t, m, g, p instance types. When you look at the whole platform, naming, gotchas etc AWS creates lot of cognitive overhead for developers. I find Google cloud far easier to use, partially due to things like these.
obulpathi··on Google Infrastructure Security Design Overview
> Many of these solutions are unavailable below a certain scale

This was true before Google Cloud. With Google Cloud, you can enjoy these benefits whether you are an individual developer with a sub 100$ monthly budget / a mom n pop shop with 1000$ budget or a SMB / Startup with a 10K$ - 100K $ to spend on your infrastructure.

obulpathi··on Using TensorFlow in Windows with a GPU
Give Google Cloud ML a try. With Google Cloud ML, one does not need GPU instances or any other cluster. Just submit your ML job from your computer. Cloud ML will train your model and bill you for the by training resources. No need to spin up / spin down anything.
obulpathi··on How Google Is Challenging AWS
I completely agree with this. But, what I found more impressive about Google Cloud is that it requires a far fewer number of people to get things done compared to AWS. AWS believes in providing fundamental building blocks rather than frameworks. It takes a lot of time, money, skills, expertise and people to make it work. Google Cloud provides frameworks which are easy to use, secure by default, does not require tuning or turning knobs, high performance, scale automatically (or automagically) and are ready to use. It is an impressive feat that Snapchat, a 25 Billion dollar company, runs on Google Cloud with 2 part-time DevOps engineers and recently Pokemon Go was able to scale to facebook level user engagement in a mere month time period with 4 backend engineers (and of course with lot of help from Google). Things like these are impossible to achieve with AWS. Bottom line, if you want to get something done, Google Cloud will get you there in a fraction of time compared to AWS and can scale way better than what AWS can, at half of the AWS's price.
obulpathi··on Ask HN: Alternatives to AWS?
Google Cloud is not an alternative, but a much better option. If you have pains with AWS or want a better version of Cloud (ease of use, performance, scalability), give Google Cloud a try.
obulpathi··on Users spend more minutes per day in Pokémon Go than Facebook
They seem to be using Google Cloud and Java (https://www.nianticlabs.com/jobs/). My best guess is that they are using Google App Engine.
obulpathi··on Our nightmare on Amazon ECS
Pub/Sub at glance may look like that. But having used all the three queuing systems from AWS (sqs, kinesis streams, kinesis firehouse), Pub/Sub is far more advanced than all the three put together. It can scale instantly to millions of messages while you can only dream of that on AWS. Multi region, noops model, push delivery, no need for capacity planning are some of the great features of Pub/Sub.
obulpathi··on How Diane Greene Transformed Google's Cloud
HN comment space was not sufficient enough for me to write about the ways AWS drives me insane. So, here is my blog on 1000 cuts by AWS: https://medium.com/google-cloud/the-future-of-cloud-computin.... I feel Google cloud is much better engineered, focusing on developer happiness and productivity.

Talking specifically about my field, Cloud, Big Data and DataScience, its so painful to build a decent data stack that can handle few terabytes of data, let alone petabytes of data. Google Cloud (Pub/Sub, Dataflow & Big Query) make it a breeze to handle petabytes of data. You can literally debug a petabyte scale pipeline, while its running. Unified logs, metrics, monitoring, alerting is another feature that shows how well the Google Cloud platform is built with developer in mind.

obulpathi··on Google supercharges machine learning tasks with TPU custom chip
- Google Container Engine

- Cloud Shell

to name a couple more.

← PreviousPage 2 of 5Next →