HNHacker News
TopNewBestAskShowJobs

d3m0t3p

151 karma · joined November 8, 2021

submissionscomments
d3m0t3p··on IBM Quantum Nighthawk R2
They are currently targeting more material science, such as quantum chemistry for drugs discovery or new material discovery. They can’t and do not aim to do prime factorization on those.
d3m0t3p··on IBM Quantum Nighthawk R2
I agree that they should be doing a press release and a separate technical deep dive.

Here they mix both and don’t get the public. It has always been IBM problem to correctly target the product to the audience.

We see it here again

d3m0t3p··on Tailscale didn't stop the Hugging Face intrusion
When everyone push on friday, and you have 400 CICD pipeline triggers spawning that many nodes. How do you know if this is unexpected ?

Their cloud compute might be on demande, someone starts training a model and 50 machines are spawned. Knowning when something is unexpected is hard

d3m0t3p··on Is One Layer Enough? A Single Transformer Layer Matches Full-Parameter RL Train
They are using Qwen, so this is decoder only.
d3m0t3p··on ZCode – Harness for GLM-5.2
This is exactly the same with providers from the USA.
d3m0t3p··on IBM debuts sub-1 nanometer chip technology
It was 15 years ago. Whole management got replaced, they are quite ambitious. Let's see if how this works out now.
d3m0t3p··on Switzerland wil have a referendum to cap population at 10M
I don't think so, faster trains are overtaking slower trains. There is simply not enough space between the station to overtake without having an acceleration that would damage the trains or the tracks. For example in western switzerland the maximal train speed for the fastest trains are ~130 Km/h while the same train can go up to 200 in some swiss-german part, only due to more congestion on the western part. Trains cannot be bigger, some of them are already too big for the smaller train station and in case of rerouting / unexpected stop this causes issue. You cannot make them higher too.

You could get ride of the smaller train , only allowing big city to survive or decrease the commodity traffic or increase the rail network or increase the train station (more tracks allowing to overtake there, and have bigger trains) There is no easy solution otherwise it would have been done.

d3m0t3p··on PyTorch Landscape
Interesting to see clearml but not its bigger counterpart mlflow
d3m0t3p··on No More JetBrains Products for Me
I think this is due to their AI insight, they run locally a model and it start to burn the whole computer.
d3m0t3p··on Moving away from Tailwind, and learning to structure my CSS
It is really fun that the navbar has unaligned elements. (Docs is lower)
d3m0t3p··on AGENTS.md outperforms skills in our agent evals
Yea but the goal it not to bloat the context space. Here you "waste" context by providing non usefull information. What they did instead is put an index of the documentation into the context, then the LLM can fetch the documentation. This is the same idea that skills but it apparently works better without the agentic part of the skills. Furthermore instead of having a nice index pointing to the doc, They compressed it.
d3m0t3p··on I asked AI researchers and economists about SWE career strategies given AI
Same, Firefox iOS
d3m0t3p··on History LLMs: Models trained exclusively on pre-1913 texts
The model is fined tuned for chat behavior. So the style might be due to - Fine tuning - More Stylised text in the corpus, english evolved a lot in the last century.
d3m0t3p··on F-35 Fighter Jet's C++ Coding Standards [pdf]
Is that really the only thing you managed to remember ?
d3m0t3p··on Nvidia DGX Spark: When benchmark numbers meet production reality
Because the ML ecosystem is more mature on the NVidia side. Software-wise the cuda platform is more advanced. It will be hard for AMD to catch up. It is good to see competition tho.
d3m0t3p··on The Missing Semester of Your CS Education (2020)
In my own studies, software engineering was mostly about structurig code, coding pattern such as visitor, singleton etc. I.E how to create a maintainable codebase
d3m0t3p··on How does gradient descent work?
Would you have some literature about that ?
d3m0t3p··on How does gradient descent work?
This sounds a lot like what the Muon / Shampoo optimizer do.
d3m0t3p··on Apertus: An open, transparent, multilingual language model
Interesting to see that they enforce retroactive opt out for data collection. I wonder how they do that, what if the model is already trained with your data and you opt out.
d3m0t3p··on Open models by OpenAI
You can batch only if you have distinct chat in parallel,
d3m0t3p··on Adding lookbehinds to rust-lang/regex
Nice to see a master thesis highlighted on the research groupe page
d3m0t3p··on Cognition (Devin AI) to Acquire Windsurf
Your first link is (in my opinion) highly biased in the samples they choose, they hired maintainers from open-source repos (people with multi years of experience, on their specific repo).

So indeed, IF you are in that case: Many years on the same project with multiple years experience then it is not usefull, otherwise it might be. This means it might be usefull for junior and for experienced devs who are switching projects. It is a tool like any other, indeed if you have a workflow that you optimized through years of usage it won't help.

d3m0t3p··on Show HN: Refine – A Local Alternative to Grammarly
It is Gemma 3n, I can't give feedback yet on the battery hit, But I would not expect anything bad as these models have been developed for much smaller devices (Phones)
d3m0t3p··on ETH Zurich and EPFL to release a LLM developed on public infrastructure
Hey, really cool project, I’m excited to see the outcome. Is there a blog / paper summarizing how you are doing it ? Also which research group is currently working on it at eth ?
d3m0t3p··on A non-anthropomorphized view of LLMs
Do they ? LLM embedd the token sequence N^{L} to R^{LxD}, we have some attention and the output is also R^{LxD}, then we apply a projection to the vocabulary and we get R^{LxV} we get therefore for each token a likelihood over the voc. In the attention, you can have Multi Head attention (or whatever version is fancy: GQA,MLA) and therefore multiple representation, but it is always tied to a token. I would argue that there is no hidden state independant of a token.

Whereas LSTM, or structured state space for example have a state that is updated and not tied to a specific item in the sequence.

I would argue that his text is easily understandable except for the notation of the function, explaining that you can compute a probability based on previous words is understandable by everyone without having to resort to anthropomorphic terminology

d3m0t3p··on Occurences of swearing in the Linux kernel source code over time
You can check company names too ! It's interesting to see that by default, the graph shows google,apple. But adding meta, and IBM really changes the plot.

Meta went from 2K to 10K+ from 2018 to 2025. While IBM seems to have stopped contributing in 2008. Since they the merging with RedHat, I would have expected to see them increase again but none of RedHat / IBM seems to have increase. https://www.vidarholen.net/contents/wordcount/#redhat,oracle... Not sure if their name appearing means that they are contributing tho.

Really cool project,

d3m0t3p··on Show HN: WhatsApp MCP Server
Apparently this is the case: https://github.com/tulir/whatsmeow/discussions/199
d3m0t3p··on Ask HN: Who is hiring? (March 2025)
Hi, do you offer visa / allow remote from the EU (GMT+2)
d3m0t3p··on CTO / cofounder exit deal after 1.5y at 600k revenue without SHA
Would you mind sharing your company name? I'm a master's student in AI, and after finishing my master's thesis at IBM this summer, I'll be looking for jobs.
d3m0t3p··on Beyond Gradient Averaging in Parallel Optimization
Well, it's just like stochastic gradient descent, if you think about it. The normal gradient descent is computed using the whole training set. The stochastic gradient is trained on a batch (a subset of the training set), and in the distributed case, we compute two batches at once by doing the gradient on each in parallel. The intuition works IMO, but indeed, having the first batch update and then the second, is not equal to having the mean update.

This is indeed super cool !

Page 1 of 3Next →