HNHacker News
TopNewBestAskShowJobs

shihab

1,715 karma · joined September 19, 2018

PhD student, Computer Science & Engineering
submissionscomments
shihab··on Ask HN: What is interviewing like now with everyone using AI?
I understand this sentiment for experienced developers. It is an imperfect signal. But what is in your opinion a better signal for junior or new grads?

Every alternative I can think of is either worse, or sounds nice but impractical to implement in practice at scale.

I don’t know about you, but most interviewers out there don’t have the ability to judge the technical merit of a bullshitters’s contribution to a class or internship project in half an hour, specially if it’s in a domain interviewer has no familiarity with. And by the way, not all of them are completely dumb, they do know computer science, just perhaps not as well as an honest competitor.

shihab··on Ask HN: What is interviewing like now with everyone using AI?
To get an idea of just how advanced cheating tools has become, take a look here:

https://leetcodewizard.io/

I think every interviewer, hiring manager ought to know or be trained on these tools, your intuition about candidate's behaviour isn't enough. Otherwise, we will soon reach a tipping point where honest candidates will be at a severe disadvantage.

shihab··on When Greedy Algorithms Can Be Faster [C++]
yeah, that's very likely the explanation. All these functions are pretty high latency instructions, vs rejection sampling which only involves a multiplication. On Nvidia GPUs, mul has latency of 1-4 cycles while others are 16-32.
shihab··on When Greedy Algorithms Can Be Faster [C++]
On GPU, the rejection sampling approach wouldn't come close to analytical one.
shihab··on DeepSeek's AI breakthrough bypasses industry-standard CUDA, uses PTX
I had the impression that gpu isn’t a good fit for ultra low latency usecases. Can you please elaborate on what sort of work hft firms do with gpu?
shihab··on DeepSeek R1 671B running on 2 M2 Ultras faster than reading speed
That doesn't necessarily mean final weights are 8-bit though. Tensor core ops are usually mixed precision- matmul happens in low precision but accumulation (i.e. final result) is done in much higher precision to reduce error.

from deepseek v3:

"For this reason, after careful investigations, we maintain the original precision (e.g., BF16 or FP32) for the following components: the embedding module, the output head, MoE gating modules, normalization operators, and attention operators...To further guarantee numerical stability, we store the master weights, weight gradients, and optimizer states in higher precision. "

shihab··on DeepSeek R1 671B running on 2 M2 Ultras faster than reading speed
Please note that it’s using pretty aggressive quantization (around 4 bits per weight)
shihab··on DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
Here [1] is the leaderboard from chabot arena, where users vote on the output of two anonymous models. Deepseek R1 needs more data points- but it already climbed to No 1 with Style control ranking, which is pretty impressive.

Link [2] to the result on more standard LLM benchmarks. They conveniently placed the results on the first page of the paper.

[1] https://lmarena.ai/?leaderboard

[2] https://arxiv.org/pdf/2501.12948 (PDF)

shihab··on DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
I have been using it to implement some papers from a scientific domain I'm not expert in- I'd say there were around same in output quality, with R1 having a slight advantage for exposing it's thought process, which has been really helpful for my learning.
shihab··on TikTok says it is restoring service for U.S. users
Yes, polls are an imperfect tool. But I think they remain the only tool we have to gauge what decisions coming out of Washington are product of broad popular support vs ones product of intense lobbying from shadowy powers.
shihab··on Boosting Computational Fluid Dynamics Performance with AMD MI300X
I helped develop a hydro solver coupled with radiation at Los Alamos this summer. We observed from 7x upto 15x performance improvement on a single GPU compared to 64-core intel CPU. [1]

Now granted, the flops to byte ratio for this program might be better than an avg fluid simulator. Also, our performance tanked when we moved to multi-node system. But I am aware of underlying reasons behind the scalibility issues and they don't feel like problems that can't be overcome.

[1] https://github.com/lanl/HARD/tree/main

shihab··on TikTok says it is restoring service for U.S. users
Is anyone aware of any opinion poll among US population about banning tiktok? This to me feels like one of the issues with potentially largest disconnect between voters and politicians

Edit: found one from Pew. "The share of Americans who support the U.S. government banning TikTok now stands at 32%." Sept 05, 2024. In contrast, 87% US lawmakers voted for the law that caused this.

shihab··on Boosting Computational Fluid Dynamics Performance with AMD MI300X
That's the smallest of 4 experiments. It goes upto 140 million cells, where MI300X retains similar performance advantage of around 10% over Nvidia's H100.
shihab··on The Missing Nvidia GPU Glossary
I was looking for a simple table recently- outlining say how the shared memory or total register size/SM varies between generations (Something like that Wiki table). It was surprisingly hard to find those info.
shihab··on Tabby: Self-hosted AI coding assistant
I think that example says more about the company that chose to put that code as a demo in their homepage.
shihab··on Yemeni Coffee Shops in Texas
First part is true here in Michigan too. Wish they stayed open late.
shihab··on Ask HN: Who is hiring? (January 2025)
CS Phd student here: took an online assesment recently (for internship role). 3 questions in total. None of them were puzzle-type question, or required advanced data structure or algorithm knowledge, but they were more of implementation focused. One of the questions took me about 80 lines, in Python, which is quite a lot for these sort of tests.

I think they were the kind of problems where if you didn't break them down into smaller chunks, individually test each of them- it'd be quite hard to get them working in such a short time.

shihab··on Things we learned about LLMs in 2024
The last OpenAI valuation I read about was 157 billion. I am struggling to understand what justifies this. To me, it feels like OpenAI is at best few months ahead of competitors in some areas. But even if I am underestimating the advantage, it's few years instead of few months, why does it matter? It's not like AI companies are going to enjoy the first-mover advantage internet giants had over the competition.
shihab··on Making AMD GPUs competitive for LLM inference (2023)
I have come across quite few startups who are trying a similar idea: break the nvidia monopoly by utilizing AMD GPUs (for inference at least): Felafax, Lamini, tensorwave (partially), SlashML. Even saw optimistic claims like CUDA moat is only 18 months deep from some of them [1]. Let's see.

[1] https://www.linkedin.com/feed/update/urn:li:activity:7275885...

shihab··on Fast LLM Inference From Scratch (using CUDA)
Excellent, amazing article.

To the author, if you're lurking here, I have a tangential question- how long did it take you to write this article? From first line of code to the last line of this post?

As someone who works in GPGPU space, I can imagine myself writing an article of this sort. But the huge uncertainty around time needed has deterred me so far.

shihab··on The GPU is not always faster
An otherwise valid point made using a terrible example.
shihab··on ICC issues warrants for Netanyahu, Gallant, and Hamas officials
The biggest condition behind US aid to Jordan and Egypt is them continuing friendly relations with Israel. In 1970s when this aid was started- this condition was made very explicit by USA.

So in other words, these two at least are nothing but indirect aid to Israel.

shihab··on ICC issues warrants for Netanyahu, Gallant, and Hamas officials
EU foreign policy chief said the court's decision should be implemented. Ireland also indicated they would comply with the warrant.
shihab··on ICC issues warrants for Netanyahu, Gallant, and Hamas officials
> Hamas is a terrorist group that was elected by Gaza’s residents.

"Prime Minister Benjamin Netanyahu gambled that a strong Hamas (but not too strong) would keep the peace and reduce pressure for a Palestinian state." - From "Buying Quiet: Inside the Israeli Plan That Propped Up Hamas", NYTimes [1]

[1] https://www.nytimes.com/2023/12/10/world/middleeast/israel-q...

shihab··on ICC issues warrants for Netanyahu, Gallant, and Hamas officials
"Defense minister [Gallant] announces ‘complete siege’ of Gaza: No power, food or fuel". [1]

[1] https://www.timesofisrael.com/liveblog_entry/defense-ministe...

shihab··on ICC issues warrants for Netanyahu, Gallant, and Hamas officials
This is common and expected. Even when a serial killer suspected of 20 murder is apprehended, arrest is often made based on one or two confirmed cases, more charges are later added as investigation deepens.

Also, keep in mind foreign journalists are completely banned by Israel from entering Gaza- complicating evidence gathering.

shihab··on Optimizing a WebGPU Matmul Kernel for 1 TFLOP
Yes, that's why I was focusing on percentage of peak hardware performance, not actual flops.
shihab··on Optimizing a WebGPU Matmul Kernel for 1 TFLOP
Great article!

For context: this WebGPU version achieves ~17% of peak theoretical performance of M2. With CUDA (i.e. CuBLAS), you can reach ~75% of peak performance for same matrix config (without tensor core).

shihab··on Hezbollah pager explosions kill several people in Lebanon
do you realize that nurses in hospital, civil servants workers are among people carrying this device? That not all, not even majority of Hizbollah personnel have no military responsibility whatsoever?
shihab··on Hezbollah pager explosions kill several people in Lebanon
Hezbollah is a political org, part of the government. Many hospitals, nurses were carrying those pagers.

And, the only well-known fake child death scandal was the one fabricated by Israel aka 40 beheaded babies, babies baked in oven etc- scandals used to justify the real mass killing of over 10,000 palestinian children by now. Can you point to any well known, well-distributed "Pallywood" incident involving children?

← PreviousPage 3 of 5Next →