HNHacker News
TopNewBestAskShowJobs

unrahul

22 karma · joined April 3, 2017

build things fast and move on..
submissionscomments
unrahul··on Microgpt in pure C hits 10M tps on Apple m5
Yes, a quick back of the envelope math is 0.65 * (memory bandwidth of the card / (model weights in bytes + kv cache in bytes) ~ practical decode tps. Below context around 32k (depends upon the model but again can be used as a placeholder number) you can ignore the kv cache in bytes and the math becomes just about memory bandwidth and model weights in bytes.
unrahul··on Microgpt in pure C hits 10M tps on Apple m5
You could think of it as a standard decoder only LLM (almost all modern ones we use everyday), with some layers (experts) having parallel networks and conditionally based on the input token (per token) - the token is routed through some of these layers. In the case of a non MoE (dense) - each token goes through all layers, so the inference engine has to read all the layers and do a matrix (layer) times vector (token) computation, while in the case of MoE the number of layers per token that has to do the compute is substantially lesser, so one can expect much higher tps than a dense model at the same number of parameters (size - 7B, 27B etc)
unrahul··on Show HN: Ocrbase – pdf → .md/.json document OCR and structured extraction API
I have seen this flow in what people in some startups call "Agentic OCR", its essentially a control flow that is coded that tries pdf-parse first or a similar non expensive approach, and if it fails a threshold then use screenshot to text extraction.
unrahul··on LLMs can teach themselves to better predict the future
Hey Danny, Really nice read.

Do you plan to share the source code to see if we could replicate this?

unrahul··on Bypass DeepSeek censorship by speaking in hex
We don’t want hex , can ask in a language that is not popular or the first 5 in the dataset , and it would answer , but not always will work with deep think . Using a tiny translator model in front of the api can make it more ‘open’.
unrahul··on Fixing Gemma Bugs
Daniel is of the best engineers I have ever worked with. Engineer in the true sense of wanting to know how something works and figuring out ways to improve it !
unrahul··on Advent of GenAI Hackathon on Intel GPUs
AI Developers, Startup Founders, Students: Ready for a Challenge? Join the "Advent of GenAI Hackathon" by Prediction Guard, with Intel Liftoff's support. A week-long journey into Generative AI awaits, packed with daily challenges to test your skills. Dive Deep into LLMs and experiment hands-on with Intel Corporation's AI Developer Cloud. Experience the power of Intel Xeon CPUs and Intel Data Center GPU Max! Build a Jupyter Notebook-based application that could win you cloud credits and recognition. Enroll by Dec 2: https://adventofgenai.com
unrahul··on [dead]
I couldn't find online how to finetune LLMs on an Intel dGPU, so i made a simple version. This particular one can be used to generate text based on your favorite book (for eg). I hope you find it useful if you are having an Intel discrete GPU.
unrahul··on Unofficial guide to setup Intel dGPUs in Linux
If you are one of those folks who use Linux and intel dGPUs (a tiny minority, I am sure :)) like me and is finding it difficult to set up a functional dev environment for the GPUs (Arc Alchemist, Datacenter Flex, GPU Max cards). This repo will help you set it up. I would love to get your feedback on this and on improvements that can be made. I made this for myself after I was tired of doing this many times. If there are any changes to the intel gpu docs, the repo gets auto-updated, so you can be sure this setup will work (to an extent).

I have also written a verification tool(https://github.com/rahulunair/xpu_verify) that can test if the setup is correct and help you fix it if it is not. The verification scripts will automatically run some C++ sycl parallel programming examples, AI examples using TensorFlow and PyTorch, and a few others. I would once again greatly appreciate any feedback on this.

unrahul··on Show HN: I wrote a book about using data science to solve “everyday” problems
Congrats!

Really nice work, I just bought it and put in on my reading list for today

unrahul··on [dead]
I was playing around with PyO3 and thought of building a UUID wrapper for Python using Rust's UUID library.
unrahul··on Show HN: Peek into a remote repo from GitHub or Gitlab quickly
Also, i am looking at ways to make pulling the content the fastest way possible using cdn cached files where possible and things like that..
unrahul··on Show HN: Peek into a remote repo from GitHub or Gitlab quickly
for the core functionality, the tool pretty much does that, although I am thinking of adding trending repos and things as such.. and any other features that other folks have in mind, keeping the core utility of the tool as opening a repo in your editor
unrahul··on Show HN: Peek into a remote repo from GitHub or Gitlab quickly
Yup that is the intent, i designed it keeping in mind, I download very many repos and open it on vim, search is then offloaded to your tool of choice.

Future improvements I am looking at is to add a search for trending remote repos and some caching mechanism

unrahul··on Show HN: GitHub-peek – quickly peek a remote repo locally in ur favorite editor
With github1s getting released last week, I have seen a few implementations that help in quickly viewing a remote codebase. This is my attempt to build one, the tool pulls a remote repo locally and opens it in your favorite editor. It takes care of deleting the repo after you close your editor as well.