HNHacker News
TopNewBestAskShowJobs

paulcjh

138 karma · joined February 16, 2019

CEO @ Mystic - YC W21
submissionscomments
paulcjh··on Midjourney web experience is now open to everyone
We're one of the Black Forest Labs partners, you can try the models here: https://www.mystic.ai/black-forest-labs
paulcjh··on Show HN: Turbo Registry – Rust Docker Registry for AI Cuts Cold Starts by 90%
Thanks!
paulcjh··on Show HN: Turbo Registry – Rust Docker Registry for AI Cuts Cold Starts by 90%
Hey folks.

I’m Paul, the CEO of Mystic (YC W21). Today, we're launching our Turbo Registry, a new type of Docker registry built in Rust. It's designed specifically for AI/ML inference deployments, which typically have larger container sizes. Traditional Docker registries suffer from slow download times, which can be a significant bottleneck in AI inference workflows when going from 0->1. Our new registry has high bandwidth parallel downloading when connected with our new containerd adapter. This adapter changes the way images are mounted and downloaded and interfaces with a new V3 API specification we have introduced in our registry. We will be open-sourcing this later this year but will keep it closed-source as we develop it further for the moment.

With Turbo Registry, we achieved:

• 5GB Docker images loading in 10.23 seconds (down from 82.21 seconds)

• 10GB Docker images loading in 14.75 seconds (down from 147 seconds)

• 20GB Docker images loading in 23.72 seconds (down from 270.47 seconds)

You can use the registry today behind our serverless endpoints.

Information about the roadmap and current limitations can be found in the documentation. Check out our documentation here to get started with Turbo Registry and experience the speed boost in your AI projects. We’re looking forward to your feedback and contributions!

paulcjh··on OpenAI – transformer debugger release
A very meagre attempt to look like they provide open source tools help the world safely make AGI
paulcjh··on Show HN: Fractional GPUs for AI
There's not really a trade off, just less VRAM.
paulcjh··on SambaNova Says Its First with Trillion-Parameter GenAI Model
The cost of their chips need to some down first, way too big and expensive
paulcjh··on Deploy Gemma 7B with TensorRT-LLM and achieve > 500 tok/s
Really cool!
paulcjh··on VLLM with Mistral 7B guide and benchmarks (1.8k+ tokens/s)
Managed to get 1.8k tokens per second with a batch of 60 when running vLLM with Mistral 7B on an A100 40GB in bfloat16 mode. Pretty damn fast!

vllm==0.2.0 got released an hour or so ago, so it's pretty fresh. Let me know fi you'd like anything else in there.

paulcjh··on Mistral AI 7B – beats Llama 2 7B and 13B
Amen to that!
paulcjh··on Mistral AI 7B – beats Llama 2 7B and 13B
Their new 7B model beats the Llama 2 7B in all benchmarks that they provided, and in many cases the 13B variation. This is a basic demo for you to try it it, seeing the best results with the instruct variation, and JSON extraction seems good!
paulcjh··on Llama 2 chat with vLLM and tensor parallel guide
Hope that you enjoy the guide, below is also some cost/speed comparisons for running the models with vLLM:

- 7B, 1x A100, 25GB VRAM, 49 tok/s, $0.0113 /1k tok - 13B, 1x A100, 37GB VRAM, 32 tok/s, $0.0174 /1k tok - 70B, 2x A100, 150GB VRAM, 13 tok/s, $0.128 /1k tok

paulcjh··on Stopping at 90%
In the context of building a product it’s normally best to do some of that final 10% as early as possible. Fail fast, see what works fast. I’ve had so many projects not get to 100% due to things I could have validated in the first 5%
paulcjh··on Alfred-40B, an OSS RLHF version of Falcon40B
Any performance benchmarks compared to other LLMs? Also, any performance increases on the orig Falcon model in inference speed?

We ditched most of our focus on Falcon 40B after Llama 2 70B came out, both the tokens per sec and quality of results are not even close.

paulcjh··on Nvidia H100 GPUs: Supply and Demand
I know people are working on it - but IMO it's not spoken about anywhere near as much as slapping more GPUs on a problem. I'm on about enterprise workload orchestration not just making a model faster
paulcjh··on Predictive Debugging: A Game-Changing Look into the Future
Activating type checking mode in VSCode was a game changer for me when running python, help catch loads of the edge cases for correctly for some production code
paulcjh··on Nvidia H100 GPUs: Supply and Demand
By far the biggest issue it utilisation of the GPUs, if people worked on that instead of throwing more power at problems this would be way less of a problem