HNHacker News
TopNewBestAskShowJobs

ismailmaj

243 karma · joined August 9, 2022

ML Infra

[[username]][[hardy number]] at g mail

submissionscomments
ismailmaj··on Claude Code’s suggested message feature: I think the real customer is the model
The idea is that a thread can have many reasonable follow-ups that the user would've accepted, so it is wrong to punish the model for predicting a follow up that is different from the user message, as that prediction could've been accepted by the user if it was given.
ismailmaj··on Mistral Large 4
it's extremely unlikely that they re-use anything from a chinese model, that would be obvious quickly, what's more likely is using documents produced by a better model to create synthetic data.
ismailmaj··on Mistral Large 4
Those are comments from Europe. The US is waking up now and I expect them to be much harsher.

I really want them to win as that's our last horse in the AI race, but ~200 research-oriented devs out of 1800 employees? I believe they agree it's pretty doomed and have pivoted.

ismailmaj··on Gemini 4 Argon
They got heavily hit in the January 2023 layoff cycle.
ismailmaj··on Claude is a Contrarian
5 for sure is unbearable, 4.8 is a reasonable sweet spot, 4.6 tends to just agree with whatever I say.

But lately I haven't downgraded because 5 is so much better at tool use, so I just accept the cost of Fable for chatting and hope Opus 5.1 fixes this mess.

ismailmaj··on OpenArch – PyTorch implementations of modern LLM architectures
FYI Mistral at launch just dropped the weights without any model architecture mentioned.

Most of the OSS models follow the same architecture which is Llama +- a few things, so it wasn't too hard for people to make it work.

ismailmaj··on I resigned from Anthropic today
Assuming that current LLMs don't have a "self" whatever that means, what makes you think it won't emerge after enough intelligence?

Also it's not even necessary, it's sufficient for it to be aligned with human values of survival (likely in its training) and act accordingly.

ismailmaj··on Mistral raises €3B
Pay isn't the main problem with Paris, it's quality of life.

Housing near work is mostly old Haussmannian buildings that are rent capped and where demand far exceeds supply, so most of the stock is unmaintained. You either accept bad housing in the city or live in the suburbs and commute — not ideal.

To make matters worse, French companies (especially old-school ones) have a culture of presenteeism for white collar jobs, it's uncommon for people to leave before 6PM and staying late is rewarded as high engagement.

Lastly, you might consider buying and renovating a house so you can escape this dilemma, but it isn't cheap: 2-bedroom (T3) around 700k euros, that represents 20 years of frugal savings on 90k gross.

Vacation, cheap and good healthcare, and unemployment benefits are great in France, but the current government has been eroding these social benefits, and those are not unique to France for well-paid tech workers anyway.

ismailmaj··on Mistral raises €3B
All original scientist co-founders of H have left.

From what I've seen in the past few months, Mistral is the unique European lab still trying to compete (I wouldn't count Poolside as EU).

ismailmaj··on Apple caught off guard by AI demand for Mac Mini and Mac Studio
Obviously they do not, the same way Amazon isn't treating their drivers poorly, it is just third party contractors "surprisingly" not doing their job well.
ismailmaj··on Nvidia agrees to acquire Hugging Face for $13B
It feels really weird to see this when I keep seeing the CEO and CTO of HuggingFace push for local inference to win.
ismailmaj··on Python Polars Cheatsheet (based on our O'Reilly book)
even for just in-memory quick data analysis?
ismailmaj··on Rust SIMD on the GPU
There is something very SIMD-coded in GPU programming which is coalesced stores/loads, if a warp (32 threads) handles contiguous memory, it will create ~4 transactions instead of 32.
ismailmaj··on U.S. economy lost 23,000 jobs in July, a sudden reversal
Given the trend of revisions, I wouldn't be surprised if the real loss is around 6 figures, May 2026 revision went from 172k to 63k.
ismailmaj··on Why Large Language Models Fail at Tabular Prediction
Unsure if it's LLMs that fail at tabular data or its just that tree boosting are spectacular at that task.
ismailmaj··on AI doesn't generate working products, that's still your job
For my projects, I can produce with LLMs a codebase that I'd be proud of, but it requires extensive use of the plan mode and nagging anytime the implementation looks much more complex than the feature at hand requires.

Anytime I tried spec driven development, it produced total garbage, the problem is that the model usually swings too far, if they write unnecessary complexity, a nag about it and you risk it code golfing, this problem across 100 lines of specs and you're guaranteed it will swing too far on some segments you just wanted X slightly more than Y (usually for me it was reliability dropped for better readability since # users = # developers = 1)

Another thing I dropped is steering implementation design too early, start with the goal and nag in the direction you want, it works well for me but might not work well for big complex codebases.

For now I love LLM coding but I'm sure my opinion will change if I have to review code from a developer that does the 1 prompt = 1 PR without looking at the code jutsu.

ismailmaj··on Advancing the price-performance frontier with GPT‑5.6
it's 80% less cost, not 80% in efficiency gains, could be that Luna was overpriced to begin with, we don't have much info on the models themselves.

Assuming the efficiency gains are real, I feel like something has to give, maybe worse quality due to aggressive quantization/kv cache compression?

ismailmaj··on LearnVector – Andrew Ng's AI company building one‑to‑one learning experiences
While the model itself is great at hiding the answer, the AI summary keeps giving it away, it's so annoying and AFAIK it cannot be disabled.
ismailmaj··on Claude Opus 5
I'd pay good money to see OpenAI "oh fuck" war rooms.
ismailmaj··on Postgres rewritten in Rust, now passing 100% of the Postgres regression tests
I think I would be horrified looking at your Claude API bill.
ismailmaj··on GPT-5.6
Apparently plus users do not have access to Sol, so I'm really worried about the ugly Terra Pelican.
ismailmaj··on A software engineering interview question I like: computing the median
I don't know what the day to day would be like but it was a ML research engineer role.
ismailmaj··on A software engineering interview question I like: computing the median
I'm pretty confident I would've been able to come up to the solution in better circumstances, maybe even without hints.

But it was clear in this case that the interviewer just took a question from the company's bank of questions and was alt-tabbed for half the interview, I have felt the energy early and I was also half-checked out.

I'm aware I'm saying this post factum, but I had a very fun first interview with that company and matched well with the first interviewer so my expectations were high, and then I got hit by the big tech style interview when it was an early stage startup.

ismailmaj··on A software engineering interview question I like: computing the median
I got that recently, I really didn't like that question.

For the n-th percentile version, the obvious solution is sorting and it takes 10 seconds to get to that point, 5 minutes of implementation with tests. Good. It's all downhill from here.

Then you get hit with the "it's a data stream" and you realize you have to implement a balanced tree on the spot which I wouldn't describe as fun.

You may or may not be able to implement that. I did not. Blabbered something about Rust having sorted B-Trees and I don't think Python has them -- they do not on the standard library.

Then the interviewer leaned heavily on the "reduce memory usage" and I couldn't come up with a solution (no shit it's Ω(n) and he didn't even tell me to go fetch for a randomized algorithm). I later understood he expected the reservoir sampling solution which is basically keeping a representative group of size K that is a good proxy of the whole stream, it goes like this: keep the K first elements, any elements after that replaces any element of our sample at random.

What I did after 10 minutes of weird silence is to assume the data stream follows a normal distribution and computing the P-percentiles by computing the running mean and standard deviation.

I felt frustrated at the end of the interview because it really felt like a big gotcha of either you know the reservoir sampling "leetcode trivia" or you don't.

ismailmaj··on Tell HN: Who wants to be hired" posts outpace "Who's hiring" 2 to 1
High interest rates and instability makes it a bad environment to invest, it raises the bar on speculative/growth investment, and tech is mainly that.

Mix that with heavy AI bills, there isn't a lot of budget left for hiring.

ismailmaj··on Leaking YouTube creators' private videos
From talking to someone that worked at YouTube for 15 years, they still had a lot of core Python code in 2016 that was legacy from the OG company/team and that code needed to be transitioned to follow the Google way of doing things in C++/Go.

I don't think it was distinct enough from the Google culture like Android was at the start of the acquisition but it seems they had leeway to do their own thing.

ismailmaj··on Leaking YouTube creators' private videos
I'll pick small company, thank you.
ismailmaj··on Fable 5 Is Back
I'm not getting the Opus 4.8 switch for coding, supposedly given how fast I reached the usage limit, which is kind of nice.
ismailmaj··on Zig's new bitCast semantics and LLVM back end improvements
It's great for defining fancy floats used in machine learning

e.g. https://github.com/zml/zml/blob/33ced8fa078b3c7c8c709bd526ae...

ismailmaj··on OpenAI unveils its first custom chip, built by Broadcom
they're still kings for training, though I've heard Anthropic is training now on JAX+TPU setup, so might not be a monopoly in that segment.
Page 1 of 5Next →