HNHacker News
TopNewBestAskShowJobs

sherlockxu

419 karma · joined September 20, 2022

Hello world
submissionscomments
sherlockxu··on DeepSeek OCR
Seems there’s still some confusion around what DeepSeek-OCR really does. Learn about the model, Contexts Optical Compression, and its impact on LLMs here: https://www.bentoml.com/blog/deepseek-ocr-contexts-optical-c...
sherlockxu··on LLM Inference Handbook
Thanks. We just added the example.
sherlockxu··on LLM Inference Handbook
Thanks. We have updated the image to make it more accurate.
sherlockxu··on LLM Inference Handbook
Hi everyone. I'm one of the maintainers of this project. We're both excited and humbled to see it on Hacker News!

We created this handbook to make LLM inference concepts more accessible, especially for developers building real-world LLM applications. The goal is to pull together scattered knowledge into something clear, practical, and easy to build on.

We’re continuing to improve it, so feedback is very welcome!

GitHub repo: https://github.com/bentoml/llm-inference-in-production

sherlockxu··on Navigating the World of Large Language Models
Hi HN readers,

One thing I didn't mention in this blog post is that developing vertical models tailored to specific industries may be more important than creating general-purpose models.

Actually I have been wondering why we need so many general-purpose models? People in this world come from different industries and what they need is targeted solutions. Vertical models can address nuanced problems that general-purpose models might overlook due to their broad training.

Feel free to leave your comments here :-)

sherlockxu··on LLM in a Flash: Efficient Large Language Model Inference with Limited Memory
Apple recently revealed a new method in a research paper, enabling the operation of AI on iPhones. This approach streamlines LLMs by optimizing flash storage.
sherlockxu··on Ask like a human: Implementing semantic search on Stack Overflow
"Our hypothesis is that if our semantic search produces high-quality results, technologists looking for answers will use our search instead of a search engine or conversational AI."

I am not sure about others, but as long as my problem is solved, I do not care whether the answer is provided by AI or human.

sherlockxu··on Handling 100k consumers with one pulsar topic
I asked the author the same thing. But they have not contributed it back to the community, but perhaps that's sth they will consider in the future.
sherlockxu··on Handling 100k consumers with one pulsar topic
It looks like the website has some problems. See this link: https://www.streamnative.io/blog/handling-100k-consumers-wit...