HNHacker News
TopNewBestAskShowJobs

jiayq84

119 karma · joined April 18, 2017

submissionscomments
jiayq84··on What Happens When Hyperscalers and Clouds Buy Most Servers and Storage?
I do a startup called Lepton AI. We provide AI PaaS and fast AI runtimes as a service, so we keep a close eye on the IaaS supply chain. For the last few months we see supply chain getting better and better, so the business model that worked 6 months ago - "we have gpus, come buy barebone servers" no longer work. However, a bigger problem emerges. Probably a problem that could shake the industry: people don't know how to efficiently use these machines.

There are clusters of GPUs sitting idle because companies don't know how to use them. It's embarrassing to resell them too because that makes the images look bad to VCs, but secondary market is slowly happening.

Essentially, people want a PaaS or SaaS on top of the barebone machines.

For example, for the last couple months we were helping a customer to fully utilize their hundreds-of-card cluster. Their IaaS provider was new to the field. So we literally helped both sides to (1) understand infiniband and nccl and training code and stuff; (2) figure out control plane traffic; (3) built accelerated storage layer for training; (4) all kinds of subtle signals that needs attention. Do you know that a GPU can appear OK in nvidia-smi, but still encounter issues when you actually run a cuda or nccl kernel? That needs care. (5) fast software runtimes, like LLM runtime, finetuning script, and many others.

So I think AI PaaS and SaaS is going to be a very valuable (and big) market, after people come out of the frenzy of "grabbing gpus" - and now we need to use them efficiently.

jiayq84··on Show HN: Conversational search in less than 500 lines of Python
Full open-source code with Apache license here: https://github.com/leptonai/search_with_lepton
jiayq84··on Show HN: Conversational search in less than 500 lines of Python
Hi folks - Yangqing from Lepton here. The idea came from a coffee chat with a colleague on the question: how much of the RAG quality comes from the old good search engine, vs LLMs? And we figured out that the best way is to build a quick experiment and try it out. What we learned is that search engine results matter a lot, and probably more important than LLMs. We decided to put it up as a site and also open source the full code.

You can try plug in different search engines or even your own elastic interface, write different LLM prompts, pick different LLM models - a lot of ablation studies that could be tried out.

We appreciate your interest and happy Friday!

jiayq84··on Structural Decoding (Function Calling) for All Open LLMs
General availability of the structured decoding capability for ALL open-source models hosted on Lepton AI. Simply provide the schema you want the LLM to produce, and all our model APIs will automatically produce outputs following the schema. In addition, you can host your own LLMs with structured decoding capability without having to finetune
jiayq84··on Super AI Creativity App Run with Local GPU on Windows/Linux/MacOS
Super cool exhibition of what a local machine can already do in the AI frenzy!
jiayq84··on Show HN: Running LLMs in one line of Python without Docker
Oh wow yeah, that is a beast. Let me give it a shot.
jiayq84··on Show HN: Running LLMs in one line of Python without Docker
Thanks so much for the warm words!
jiayq84··on Show HN: Running LLMs in one line of Python without Docker
Thanks - we definitely agree that llama.cpp is great. Big fan of their optimizations. We are more or less orthogonal to the engines though - in the sense that we serve as the infra/platform to run and manage those implementations easily. For example, we support running a wider range of models - for example sdxl is one single line too:

lep photon run -n sdxl -m hf:stabilityai/stable-diffusion-xl-base-1.0 --local

It's really about how to productize a wide range of models as easy as possible.

jiayq84··on Show HN: Running LLMs in one line of Python without Docker
Thanks - the policies are listed here: https://www.lepton.ai/policies

we'll put a link on our homepage.

In short - we do not collect, record, or log any of your prompts and responses. They are computed in memory, returned and discarded on the fly.

jiayq84··on Show HN: Running LLMs in one line of Python without Docker
In theory one can have 640G = 8 * 80G A100s memory and launch it. 180B Falcon with fp16 will be 360G, so there would be enough memory. It's definitely going to be very expensive indeed.
jiayq84··on Show HN: Running LLMs in one line of Python without Docker
Great catch! Our cloud machine encountered a cuda error (the GPU fell off PCIe) - had to restart it. It's back to normal now.

All the more reason to have a managed version of services :)

jiayq84··on Show HN: Running LLMs in one line of Python without Docker
It's not only about "building a docker" but also maintaining multiple models, multiple environments and a lot of users. Imagine there is a group of engineers each needing to deploy their own models: one needs tensorflow 1.x, one needs tensorflow 2.x, one needs pytorch and one needs a very strange combination of dependencies. Trust me, things get complex very easily:

https://github.com/leptonai/examples/blob/main/advanced/whis...

I definitely agree that for a fixed use case, building a docker once and for all is probably the simplest and best approach. However, it quickly gets more complex and out of hand.

Also the basic plan is free for independent developers. You don't need to pay more than as if you were using EC2 instances, but with the platform convenience - we definitely hope it's worth it!

jiayq84··on Show HN: Running LLMs in one line of Python without Docker
To show some actual coding examples, We have made the python library open-source at https://github.com/leptonai/leptonai/. With it, launching a common HuggingFace model is as simple as a one liner. For example, if you have a GPU, Stable Diffusion XL is as simple as:

pip install -U leptonai

lep photon run -n sdxl -m hf:stabilityai/stable-diffusion-xl-base-1.0 --local

And you have a local OpenAPI server that runs it! Go to http://0.0.0.0:8080/docs, or use your favorite OpenAPI client.

We've been building AI API services using such tools ourselves. The easiest way to try out Lepton is to head to https://lepton.ai/playground and use our API service for popular models: Stable Diffusion, LLaMA, WhisperX, and other interesting showcases

We are proud of our performance. For example, we have probably the fastest LLaMA 7B and 70B model APIs, and it costs $0.8 to run 1 million tokens inference - we believe it's the most affordable one in the market. In addition, during the open beta phase, calling these services is free when you sign up for the Lepton AI platform.

Under the hood, we wrote a platform to allow you to run things easily on the cloud with ease. For example, if you find Pygmalion to be a great conversation model but you don't have a GPU, use lepton's Remote() capability to launch a service:

from leptonai import Remote

pygmalion = Remote("hf:PygmalionAI/pygmalion-2-7b", resource_shape="gpu.a10")

Wait a few minutes for the model to be downloaded and run, and you can now use it as if it were a standard python function:

print(pygmalion.run(inputs="Once upon a time", max_new_tokens=128))

If you are interested in the operational details, you can find fine-grained controls at https://dashboard.lepton.ai/ as a fully managed platform - we also support BYOC (bring your own compute) if you are an enterprise needing more autonomy over infrastructure.

jiayq84··on The End of Starsky Robotics
I don’t want to be mean, but since you mentioned RCNN - no, you are dead wrong. RCNN was open sourced in 2014, check the repo: https://github.com/rbgirshick/rcnn

Not to mention that nvidia has thrown numerous open source efforts over the years. If SR was under the impression that 2017 was a dry year for open source deep learning vision systems - I can understand why it didn’t do very well technology wise.

Disclaimer: have been doing deep learning open source and research over the years. Have touched all major frameworks in the market.

jiayq84··on The End of Starsky Robotics
Just to clarify a little bit... "At the time, very few object detection models had public implementations" - this is wrong. Almost all object detection models had public implementations starting from 2014, most notably Detectron (Caffe), GoogleNet/SSD (Tensorflow and matlab). Post 2015 when TensorFlow was released, one can find even more implementations.

Data is the problem. Everyone has the algorithm but not enough people have data (especially labeled ones)

jiayq84··on Relicensing React, Jest, Flow, and Immutable.js
We are moving to Apache 2.0 in a few days. Early draft at https://github.com/Yangqing/caffe2/tree/apache pending double check to make sure we are honoring all existing contributors.
jiayq84··on Relicensing React, Jest, Flow, and Immutable.js
Yangqing (creator and main author of Caffe/Caffe2) here. We are moving to Apache 2.0 in a few days.
jiayq84··on Relicensing React, Jest, Flow, and Immutable.js
Yangqing (creator and main author of Caffe/Caffe2) here. We are moving to Apache 2.0 in a few days.
jiayq84··on Facebook and Microsoft introduce ecosystem for interchangeable AI frameworks
So what we do is to keep syntax=proto2, but allow users to compile with both protobuf 2.x and protobuf 3.x libraries. Minumum need is 2.6.1. We kind of feel that this gives maximum flexibility for people who have already chosen a protobuf library version.
jiayq84··on Facebook and Microsoft introduce ecosystem for interchangeable AI frameworks
Learned one more thing today!
jiayq84··on Facebook and Microsoft introduce ecosystem for interchangeable AI frameworks
Interesting data, thanks!

We chose protobuf mainly due to a good caffe adoption story and also the track record of it being compatible with many platforms (mobile, server, embedded, etc). We actually looked at thrift - which is Facebook owned - and it is equally nice, but our final decision was mainly to minimize the switching overhead for existing users such as Caffe and TensorFlow.

To be honest, protobuf is indeed a little bit hard to install (especially if you have python and c++ version differences). Would definitely be interested in taking a look at possible solutions - serialization format and the model standard is mor e or less orthogonal, so one may see a world where we can convert different serialization formats (JSON <-> protobuf as an overly simplified example)

jiayq84··on Facebook and Microsoft introduce ecosystem for interchangeable AI frameworks
Thanks so much @haberman! Yep, the whole thing is a little bit confusing... We basically focused on two things:

- not using extensions, as 3 does not support it - not using map, as 2 does not support it

and we basically landed on restricting ourselves to use the common denominator among all post-2.5 versions. Wondering if this sounds reasonable to you - always great to hear the original author's advice.

Plus, I am wondering if there are recommended ways of reducing protobuf runtime size, we use protobuf-lite but if there are any further wins it would be definitely nice for memory and space constrained problems.

jiayq84··on Facebook and Microsoft introduce ecosystem for interchangeable AI frameworks
Thanks! Haven't yet, but our TPMs are going to reach out for collaborations. I wish we were grad school mode where latency is <1 hour, but it pays to get things proper across multiple companies. Kindly stay tuned.
jiayq84··on Facebook and Microsoft introduce ecosystem for interchangeable AI frameworks
Yangqing here (caffe2 and ONNX). We did use protobuf and we have an extensive discussion about its versions even, from our experience with the Caffe and Caffe2 deployment modes. Here is a snippet from the codebase:

// Note [Protobuf compatibility] // ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ // Based on experience working with downstream vendors, we generally can't // assume recent versions of protobufs. This means that we do not use any // protobuf features that are only available in proto3. // // Here are the most notable contortions we have to carry out to work around // these limitations: // // - No 'map' (added protobuf 3.0). We instead represent mappings as lists // of key-value pairs, where order does not matter and duplicates // are not allowed.

jiayq84··on Facebook and Microsoft introduce ecosystem for interchangeable AI frameworks
Yangqing here from facebook. We consciously made it a MIT license as onnx is intended to be widely shared by a lot of participants, and MIT seems to be more widely agreeable among different parties co-owning it. It's also simpler.
jiayq84··on Facebook and Microsoft introduce ecosystem for interchangeable AI frameworks
Yangqing here (created Caffe and Caffe2) - we are much interested in enabling this path. Historically CoreML has provided Caffe and Keras interfaces, and having ONNX / CoreML interop would help a lot for everyone to ship models more easily.

Earlier in the year we provided compatibility between Caffe2 and Qualcomm's SNPE library, which follows this similar philosophy.

jiayq84··on Caffe2: Open Source Cross-Platform Machine Learning Tools
There is no push to migrate from Caffe to Caffe2 for sure. After Facebook "dogfooding" our own implementation, I think it is safe to say that C2 is now much stable. I would encourage you to try migrating it, letting us know if you run into problems, and stay tuned for the nice additional features that you may be able to enjoy from C2 - like optimized computation with MKLDNN, etc.
jiayq84··on Caffe2: Open Source Cross-Platform Machine Learning Tools
Haha yeah, I definitely feel your pain - as a caffe developer it really makes me cry when things get so incompatible. I've made some improvements in caffe2 to make it more modular - checkout http://GitHub.com/caffe2/caffe2_bhtsne/, things like such will potentially make things more maintainable than the old Caffe solution.
jiayq84··on Caffe2: Open Source Cross-Platform Machine Learning Tools
PyTorch definitely makes experimentation much better. For example, if you want to train some system that is highly dynamic (reinforcement learning, for example), you might want to use a real scripting language which is Python, and PyTorch makes that really sweet.

Sometimes the line gets a bit blurred - for research that are focusing on relatively fixed patterns, such as Mask RCNN, both PyTorch and caffe2 are working great. In fact, Mask RCNN is trained in Caffe2, and that also makes things much easy when we put it on mobile - what our CTO Mike Schroepfer showed in his keynote is a Mask RCNN model trained and then deployed onto mobile with Caffe2.

jiayq84··on Caffe2: Open Source Cross-Platform Machine Learning Tools
TensorRT currently is supporting Caffe; it would definitely be interesting to make a direct Caffe2 compatibility support.
Page 1 of 2Next →