HNHacker News
TopNewBestAskShowJobs

vvipgupta

84 karma · joined January 25, 2021

Building open-source ML monitoring tool. PhD, UC Berkeley
submissionscomments
vvipgupta··on Launch HN: UpTrain (YC W23) – Open-source performance monitoring for ML models
Thanks for letting us know. We did some docs restructuring before the launch, and missed fixing this link. It is now available here: https://docs.uptrain.ai/docs/uptrain-examples/quickstart-tut...
vvipgupta··on Launch HN: UpTrain (YC W23) – Open-source performance monitoring for ML models
Generally, MLOps helps in reducing engineering headaches. During our user interviews and customer calls, we realized very early that customization is key for ML model monitoring since all models are different. Thus, we have built the framework to lessen the engineering headache while allowing customizability (think PyTorch). Would love to know your thoughts on this.
vvipgupta··on Launch HN: UpTrain (YC W23) – Open-source performance monitoring for ML models
Thanks! Also, wondering how did you hear about Arize? Have you dealt with the pain of ML model monitoring in the past?
vvipgupta··on Launch HN: UpTrain (YC W23) – Open-source performance monitoring for ML models
Just curious, what kind of use cases do you have in mind?
vvipgupta··on Launch HN: UpTrain (YC W23) – Open-source performance monitoring for ML models
Thanks for the very relevant comment :) We provide users the option to attach their training data from csv/json (working to support loading from cloud storage provider or data lakes). We have illustrated this in some of our examples, such as the human orientation classification: https://github.com/uptrain-ai/uptrain/blob/main/examples/hum...
vvipgupta··on Launch HN: UpTrain (YC W23) – Open-source performance monitoring for ML models
Additionally, refinement is a key focus of ours. Figuring out the best data points to retrain the model upon has twin benefits:

1) It provides automated issue resolution and saves data scientists' effort to debug and fix their models. 2) It allows us to reduce false positives in alerting: we send alerts only when we see a dip in model performance, or retraining can lead to improved model accuracy.

vvipgupta··on Launch HN: UpTrain (YC W23) – Open-source performance monitoring for ML models
Thanks for the suggestion and links. Completely agree, ML production data management can be painful and to support model refinement for users that operate at scale, an abstraction at the data layer would be a useful feature.
vvipgupta··on Data consistency is overrated
The same is true for data (aka gradients) consistency while training large ML models. Asynchronous SGD is as good (and maybe even faster) than synchronous SGD: https://papers.nips.cc/paper/2011/file/218a0aefd1d1a4be65601...
vvipgupta··on Ask HN: Open-source ChatGPT alternatives?
You might also want to check out https://github.com/lucidrains/PaLM-rlhf-pytorch
vvipgupta··on How to train large models on many GPUs? (2021)
When training over multiple GPUs, it's hard not to think about Ray (https://docs.ray.io/en/latest/train/train.html). Ray, as an open-source project, has exploded over the last few years and helps with the memory bottleneck by segregating memory and computing.

FYI, I am not affiliated with Ray. However, I did write the following paper on scaling data-parallel training for large ML models ;) https://openreview.net/pdf?id=rygFWAEFwS

Also, another one of my papers talks about distributed training while reducing the communication bottleneck for distributed training: https://dl.acm.org/doi/pdf/10.1145/3447548.3467080

vvipgupta··on Google testing ChatGPT-like chatbot 'Apprentice Bard' with employees
This would be a good application of chatGPT, testing a test for whether it tests for rote learning or fundamentals.