HNHacker News
TopNewBestAskShowJobs

cdavid

2,233 karma · joined November 10, 2009

Hi, my name is David Cournapeau.

I used to be a data/stats geek and numpy/scipy/scikit learn contributor Those days, I dabble in engineering management in the areas of search, recommendation and ML. http://github.com/cournape twitter @cournape

submissionscomments
cdavid··on Float Exposed
Maybe I am too mathematically enclined, but this was not easy to understand.

The ELI5 explanation of floating point: they approximately give you the same accuracy (in terms of bits) independently of the scale. Whether your number if much below 1, around 1, or much above 1, you can expect to have as much precision in the leading bits.

This is the key property, but internalizing it is difficult.

cdavid··on De-Clouding: Music
I am surprised to see those discussions w/o a single mention of roon. As a music lover, roon is a software I've happily paid 100 of $ for.

While not OSS, roon 1) can run on linux 2) supports large local libraries (I have > 2k albums in FLAC, and it supports much more) 3) have roon arc that allows you to listen from phone anywhere 4) has a very good system to link metadata and recommendation within your library.

The metadata support is truly wonderful, you can easily browse your music like wikipedia, can find music per composer, performer, discover related musicians, etc. I strongly recommend people serious about music to try it out.

I've happily replaced spotify with it a few years ago, and will never go back.

cdavid··on Hashed sorting is typically faster than hash tables
the big O copmlexity makes assumptions that break down in this case. E.g. it "ignores" memory access cost, which seems to be a key factor here.

[edit] I should have said "basic big O complexity" makes assumptions that break down. You can ofc decide to model memory access as part of "big O" which is a mathematical model

cdavid··on ML needs a new programming language – Interview with Chris Lattner
I agree ability to use python to "script HPC" was key factor, but by itself would not have been enough. What really made it dominate is numpy/scipy/matplotlib becoming good enough to replace matlab 20 years ago, and enabled an explosion of tools on top of it: pandas, scikit learn, and the DL stuff ofc.

This is what differentiates python from other "morally equivalent" scripting languages.

cdavid··on Derivatives, Gradients, Jacobians and Hessians
I agree it is confusing, because starting with notation will confuse you. I personally don't like the partial derivative-first definition of those concepts, as it all sounds a bit arbitrary.

What made sense to me is to start from the definition of derivative (the best linear approximation in some sense), and then everything else is about how to represent this. vectors, matrices, etc. are all vectors in the appropriate vector space, the derivative is always the same form in a functional form, etc.

E.g. you want the derivative of f(M) ? Just write f(M+h) - f(M), and then look for the terms in h / h^2 / etc. Apply chain rules / etc. for more complicated cases. This is IMO a much better way to learn about this.

As for notation, you use vec/kronecker product for complicated cases: https://janmagnus.nl/papers/JRM093.pdf

cdavid··on The Framework Desktop is a beast
Yeah, memory bandwidth is often the limitation for floating point operations.
cdavid··on The Framework Desktop is a beast
I was surprised at previous comparison on omarchy website, because apple m* work really well for data science work that don't require GPU.

It may be explained by integer vs float performance, though I am too lazy to investigate. A weak data point, using a matrix product of N=6000 matrix by itself on numpy:

  - SER 8 8745, linux: 280 ms -> 1.53 Tflops (single prec)
  - my m2 macbook air: it is ~180ms ms -> ~2.4 Tflops (single prec)
This is 2 mins of benchmarking on the computers I have. It is not apple to orange comparison (e.g. I use the numpy default blas on each platform), but not completely irrelevant to what people will do w/o much effort. And floating point is what matters for LLM, not integer computation (which is what the ruby test suite is most likely bottlenecked by)
cdavid··on Hiroshima (1946)
It is very likely: the atomic bomb was initially built to defeat Nazi Germany, and the Manhattan program was started before Pearl Harbour. When you read The Atomic Bomb book, many scientists who worked on the bomb justified their effort by defeating the Nazis, and many had to escape from Europe.

Once it became clear Nazi were about to be defeated, there were discussions from the scientists about sharing the knowledge w/ all countries. But at that point, the scientists had long lost control over the project.

cdavid··on My Ideal Array Language
If I understand correctly what is meant by rank polymorphism, it is not just about speed, but about ergonomics.

Taking examples I am familiar w/, it is key that you can add a scalar 1 to a rank 2 array in numpy/matllab without having to explicitly create a rank 2 array of 1s, and numpy somehow generalizes that (broadcasting). I understand other array programming languages have more advanced/generic versions of broadcasting, but I am not super familiar w/ them

cdavid··on Skip the exit interview when you leave your job
Or a good one who cares about their naive reports not to get burn.

Like any advice, it is contextual. Especially when working for large organizations, the OT is the right default. If you're leaving because things are bad, it will be a mix of 1) people know but did not care/could not do anything about it and 2) people did not know about specific issues. Younger me thought it was often 2), but actually it is almost always 1).

cdavid··on P-Hacking in Startups
There are cases where A/B testing does not make sense (not enough users to measure anything sensible, etc.). But if the A/B test results were inconclusive, assuming they were done correctly, then what was the point of launching the underlying feature ?

As for the HIPPO pushing for an A/B test because of lack of confidence, all I can say is that we had very different experiences, because I've almost always seen the opposite, be it in marketing, search/recommendation, etc.

cdavid··on TPU Deep Dive
SVD/eigendecomposition will often boil down to making many matmul (e.g. when using Krylov-based methods, e.g. Arnoldi, Krylov-schur, etc.), so I would expect TPU to work well there. GMRES, one method to solve Ax = b is also based on Arnoldi decomp.
cdavid··on P-Hacking in Startups
A/B testing does not have to involve micro optimization. If done well, it can reduce the risk / cost of trying things. For example, you can A/B test something before investing a full prod development, etc. When pushing for some ML-based improvements (e.g. new ranking algo), you also want to use it.

This is why the cover of the reference A/B test book for product dev has a hippo: A/B test is helpful against just following the HIghest Paid Person Opinion. The practice is ofc more complicated, but that's more organizational/politics.

cdavid··on The Missing Manual for Signals: State Management for Python Developers
well it is both an easy way to compute in a dataframe context and a reactive programming paradigm. When combined, it gives a powerful paradigm for throwing data-driven UI, albeit non scalable (in terms of maintenance, etc.).
cdavid··on The Missing Manual for Signals: State Management for Python Developers
One of the largest, if not the largest python codebase in the world, implements similar ideas to model financial instruments pricing: https://calpaterson.com/bank-python.html.
cdavid··on I think I'm done thinking about GenAI for now
The issue about executive mandate is likely coming from the context of large corporations. It creates fatigue, even though the underlying technology can be used very effectively to do "real work". It becomes hard for people to really see where the tech is valuable (reduce cost to test ideas, accelerate getting into a new area, etc.) vs where it is just a BS generator.

Those are typical dysfunctions in larger companies w/ weak leadership. They magnified by a few factors: AI is indistinguishable from magic for non tech leadership, demos that can be slapped quickly but that don't actually work w/o actual investments (which was what leadership wanted to avoid in the first place), and ofc the promise to reduce costs.

This happens in parallel to people using it to do their own work in a more bottom up manner. My anecdotal observation is that it is overused for low-value/high visibility work. E.g. it replaces "writing as proof of work" by reducing the cost to write bland, low information, documents used by middle management, which increases bureaucratic load.

cdavid··on Getting AI to write good SQL
My observation is the latter, but I agree the results fall short of expectations. Business will often want last minute change in reporting, don't get what they want at the right time because lack of analysts, and hope having "infinite speed" will solve the problem.

But ofc the real issue is that if your report metrics change last minute, you're unlikely to get good report. That's a symptom of not thinking much about your metrics.

Also, reports / analysis generally take time because the underlying data are messy, lots of business knowledge encoded "out of band", and poor data infrastructure. The smarter analytics leaders will use the AI push to invest in the foundations.

cdavid··on Ask HN: How do you talk about past jobs you regret in interviews
Given the context, I am assuming this is on the "behavioural" side of the IV (aka what most companies call culture fit). And I am assuming you are applying to "traditional" companies, that is companies that have a defined hiring process and are large enough. This includes all FAANG and what not.

My advice:

  - write down the stories (use cases) before the actual IV
  - for each story, focus on what you learnt / succeeded
  - for the really negative ones, focus on the learning
  - for the other ones, focus on the outcomes, mentioning  things that worked and maybe some things that did not work  and how you did it
This is the part where you have to act the game and avoid being too transparent. Mentioning too much the negative will be seen as a red flag by most hiring managers or recruiters.
cdavid··on Knowing where your engineer salary comes from
I am not saying mature companies are "YOLO-ing" this randomly, but that many assumptions are made about how the input metrics trickle back to revenue/profits, and those can change. The attribution exists, but how it is done is far from an objective thing. E.g. how do you translate CTR into revenue ? How do you value an additional user ?

This can also be seen with cost saving. There are numerous examples on HN when people wonder why reducing the cost of something by X millions was not recognized (e.g. https://x.com/danluu/status/802971209176477696). Based on my own experience, most likely explanation is that's because there was no item related to this in the financial planning to be recognized.

cdavid··on Knowing where your engineer salary comes from
Sure, but things like DEI, OSS, are tiny minority in most companies. At least they were in the companies I've worked at.

You mentioned that attribution is to be decided when house is burnt, but that certainly not my observation. Which department is responsible for what revenue is what senior leadership fights over all the time, whether times are good or not.

cdavid··on Knowing where your engineer salary comes from
Say you are meta. You know that a big stream of revenue is ads. You are an engineer working for one of the myriad ML model around feeds, ads click prediction, whatever. Those ML models are in production, and cost a lot of money to maintain / operate. How much you are a cost center or a money maker will depend on a lot of non objective choices.

The essential issue is attribution, which fundamentally requires some choices about how money is actually made. Even when everybody is in good faith, there are reasonable ways to agree. And people are rarely in good faith around those things.

cdavid··on Knowing where your engineer salary comes from
This article makes the naive assumption that "what is making money" is an objective thing.

Once a company reaches a certain size, basically once it has a financial planning department w/ different VPs owning their PnL, who is making money increasingly becomes a social construct. "how to get promoted" by spakhm is much more closer to how a large, successful org works IMO: https://spakhm.substack.com/p/how-to-get-promoted

I've seen this in my career multiple times. For example, when I was involved in search for some companies, we would demonstrate through A/B testing that we would make X more money per month. Executive team changed, they decided that "A/B test does not work and slows us down", the definition of making money changed, we overnight went from a money-aking org to a cost center.

Nowadays, most companies are pushing GenAI everywhere. Most of those things don't make money, and yet a lot of promotions will be obtained across the board until the tune ends.

cdavid··on Negotiating a Job Offer
It is at least in part because of how recruiting works, especially in large tech companies. Companies have many candidates for each role, and can't easily review all of them. Between bureaucracy, how busy the hiring manager is, etc. many candidates get lost in the process.

This is why 1) it really helps to have a referral and 2) tell the recruiter when you have competing roles. Neither will change much about getting or not a role, but they will really help the prioritization and avoid you getting lost.

Paradoxically, that effect is bigger nowadays when the market is not as hot as it used to be, because recruiting is more stretched. Even some FAANG are very understaffed on the recruiting side

cdavid··on A year of uv: pros, cons, and should you migrate
Yes, that's what I do today.

The UX improvement would be to have a centralized managemend of the venv (centralized location, ability to list/rm/etc. from name instead of from path).

cdavid··on A year of uv: pros, cons, and should you migrate
That's my main use case not-yet-supported by uv. It should not be too difficult to add a feature or wrapper to uv so that it works like pew/virtualenvwrapper.

E.g. calling that wrapper uvv, something like

  1. uvv new <venv-name> --python=... ...# venvs stored in a central location 
  2. uvv workon <venv-name> # now you are in the virtualenv
  3. deactive # now you get out of the virtualenv
You could imagine additional features such as keeping a log of the installed packages inside the venv so that you could revert to arbitrary state, etc. as goodies given how much faster uv is.
cdavid··on Startup Winter: Hacker News Lost Its Faith
negotiation is very much a thing even at FAANG. If anything, it can help you being at the top of your assigned band. Other things can be negotiated.

Source: I am no great negotiator, but I've always negotiated my salary and got 10-15 % more than what I would have had w/o asking anything. This compounds after a few stints. And I've been a manager in startup/mid size/big tech: always negotiate.

cdavid··on Conda: A Package Management Disaster?
A lot of path dependency, but essentially

  1. A good python solution needs to support native extensions. Few other languages solve this well, especially across unix + windows.
  2. Python itself does not have package manager included.
I am not sure solving 2 alone is enough, because it will be hard to fix 1 then. And ofc 2 would needs to have solution for older python versions.

My guess is that we're stuck in a local maximum for a while, with uv looking like a decent contender.

cdavid··on How to give a senior leader feedback without getting fired
One of the surest way to get your manager's back is to help them make their goals and put some stuff off their plates. Like other people said, most semi competent managers are aware of issues happening in their team. If you come up w/ some proposal to solve those issues, it will improve the team much more effectively than some feedback.

It also depends on your goals, but fixing some issues encountered by your manager is one of the most reliable way to promotion in up to mid size companies, unless your manager is a a*hole.

cdavid··on How to give a senior leader feedback without getting fired
> Imagine the difference between "I want to give you feedback that you aren't spending enough time with new hires" vs "I know you've been wanting to spend more time with the new hires, why don't you take them for lunch and send me to your status meeting over Tuesday lunch time this week."

This is the proper answer. Ultimately, feedback should be about changing something. My experience is that most people are neither good at giving or receiving feedback, and that includes myself. There are more effective ways to change things.

OP's is useful when you have to give feedback, which is expected in most large companies in some form or other (evals, etc.).

cdavid··on The Influence of Japanese Archaeology on the Legend of Zelda: Breath of the Wild
As other said, gonna depend on what you like.

Meiji jingu + walking around harajuku/omotesando gives you a great contrast and is very doable in half a day. Not original, but very "tokyo vibe".

Walking around Jimbocho, with lots of old books shops around

If you can pass the evening, I would recommend eating out in ebisu yokocho

← PreviousPage 2 of 27Next →