HNHacker News
TopNewBestAskShowJobs

jfim

4,794 karma · joined January 30, 2013

I press buttons. More often than not, the right ones.

https://blog.jean-francois.im/about.html

submissionscomments
jfim··on PISA 2025 Students' reading and mathematics performance declined across the OECD
The breakdown for AI usage is four categories though, not three. While the first three are negatively correlated with increased usage, the one for using AI "to help me learn" is much more flat when it comes to usage frequency.
jfim··on An Inside Look at the Token Reseller Market
There are probably various metrics like language used to prompt the model, number of hours per day spent prompting, and many others.

There are quite a few other mitigations that could be done by providers that aren't mentioned in the article.

jfim··on Police removed prominent scientists from American Diabetes Association meeting
Why would it be bureaucracy? In places with universal healthcare, you go to the doctor, show your card, then go to your appointment. Go to the pharmacy with your script, show your card and pay whatever is left to pay, if any.

That's it. No open enrollment, no surprise bills, no trying to figure out which ppo or hmo works best, no wondering if it's in network or not or any of that BS.

jfim··on Waymos cause "way mo" injuries per mile driven than for-hire vehicles in NYC
Part of the issue is that KSI events are so rare, and it's the same issue that AV companies face. If it takes you millions of miles to encounter a single event, how can you even start to quantify the rate of events, and even prove a different rate?

To scale it in human terms, say that you're a human driver and you're 2x better than the average human driver. If you wanted to prove that you're better than the average driver, you'd need to drive multiple hundreds of millions of miles to statistically demonstrate so.

Given that the rate is significantly lower for taxi drivers in aggregate, to prove that they're safer than taxi drivers if Waymo is 2x safer, they'd still need to drive multiple billions of miles to statistically prove that they're safer.

jfim··on Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling
The base weights can't be updated but from what I recall it allows adding a low rank adapter to customize the model a little bit.
jfim··on Advances in Real Time Rendering in Games
Their YouTube channel has the excellent deep dive on the unreal engine's nanite virtual geometry system: https://youtu.be/eviSykqSUUw

The one on idTech8 global illumination from last year is pretty interesting too.

jfim··on The Kimi K3 Moment
I can't recall the last time that it was useful knowledge when writing code. Reservoir sampling, online softmax, Otsu, sure, Tianenmen square not really.
jfim··on Kimi K3: Open Frontier Intelligence
It could generate code that's plausible but has intentional flaws, kind of like the defunct underhanded C contest [0], except through a LLM.

[0] https://en.wikipedia.org/wiki/Underhanded_C_Contest

jfim··on Why American ambulance rides are so expensive
The cost isn't about the actual mileage though, it's having two paramedics each earning about 100k/yr per ambulance, while having coverage 24x7x365. So fully loaded, the labor for one ambulance might be in the high six figures to seven figures.
jfim··on Why American ambulance rides are so expensive
The article argues the exact opposite:

> The standard answer is greed: rapacious ambulance operators, owned by villainous private equity firms, exploit patients at their most helpless. But I don’t think that’s actually what’s going on. Ambulance providers are chronically unprofitable businesses; margins are thin, crews are underpaid, and operators exit the industry every year.

jfim··on Distributed system is slower than a laptop
> No CFO approves a $1.4 million annual commitment on the strength of a chart showing that the commitment scales.

Except the company probably approved a budget for AWS or another cloud provider, and basically gave a blank check to developers to deploy whatever is needed. So developers are going to just deploy MSK or whatever is trendy, instead of trying to get the most throughput from the servers they got from IT.

jfim··on Show HN: Yamanote.fun – A complete soundscape for Tokyo's Yamanote line
No real feedback other than it's pretty awesome. It'd be cool to have a version of the display above the doors that shows the upcoming stations, but I'm not sure in practice if that would be that useful since I assume most people would have that in the background as you point out.
jfim··on Ask HN: Another "Hacker News" with less AI and more human-focused hacking news?
It's also kind of hard to know if you already know someone who could give you an invite, since I don't really know the online handles of people I know in meatspace. I've resigned myself to maybe someday getting something posted that gets picked up there and asking for an invite at that time.

The RSS feeds though are pretty neat, it's what I use to fetch articles for archiving so that I can get a curated set of things to read on the go.

jfim··on Amazon without the knockoffs
We do the same for groceries though. Only a few generations ago, oranges or bananas were a luxury if you didn't live in an area where those fruit grow. Now we get avocadoes and berries way out of season, shipped from across the world.
jfim··on Resetting Xbox
Would they be counted as active players on steam if they're being played on game pass though?
jfim··on “Beyond the limit”: Satellites and mirrors in space pose threat to the night sky
Light scatters in the atmosphere, it's the same reason you can shine a laser beam and see it even though the light should be collimated. With enough sources of light, you end up with more background light pollution.
jfim··on Fable created novel 4D splat format
That already exists though. I believe Braindance VR uses a rig with a couple dozen cameras to capture the same scene from multiple viewpoints then converts it to a gaussian splat that can be walked around.
jfim··on Senior SWE-Bench: open-source benchmark that assesses agents as senior engineers
I wonder how they're planning for the benchmark to stay relevant over time.

If the benchmark is to implement features that are part of an open source project, and LLMs have those changes as part of their training dataset, it seems that they could just give a verbatim or slightly modified version of the change in their training data.

And if one updates the benchmark to only incorporate code changes that are past the models knowledge cutoff, then the benchmark is less comparable over time, since the changes in the benchmark at time T and T+1 aren't the same.

jfim··on Department of Commerce has lifted export controls on Claude Fable 5 and Mythos 5
For programmatic usage oftentimes SOTA isn't useful.

For example, I have software that summarizes articles and classifies links on webpages to build a synthetic RSS feed, both of which use LLMs, neither of which need a SOTA model.

I'll probably use LLMs to bootstrap a dataset of native ads in articles, and there again, I don't really need a SOTA model.

If it's for more open ended tasks like writing code though, I agree that at this point SOTA models make more sense to use.

jfim··on Professor denounces mass AI fraud on an exam at Brown
$30/month is likely a rounding error in the budget of students at the schools mentioned in the article.
jfim··on The gap between open weights LLMs and closed source LLMs
True, but the capabilities and knowledge of that model are also frozen in time, so the value of that model declines over time.

A model that writes code without knowledge of any language or library changes for half a decade is less useful. A 2021 era chatgpt would be quite quaint in 2026.

Right now the Chinese labs might have incentives to release their models for free, and maybe Google is happy to release open weights today, but I'm sure there are already bean counters at Google salivating at the idea of having Gemini in Chrome as part of a Google AI monthly subscription just like YouTube premium and other Google subscriptions.

jfim··on Show HN: OpenKnowledge – open source AI-first alternative to Obsidian/Notion
Obsidian is a lot more than "just markdown" though.

For example, with the appropriate plugins like dataview and charts, it's possible to create dashboards, lists, and tables that update automatically based on data elements present in documents or documents themselves. I use it to have views over my to-do lists (daily routine items, tasks that are overdue, upcoming tasks, etc), make dashboards, and show lists of documents edited on a particular date.

I'd love to migrate away from Obsidian towards something that's not proprietary, but I haven't seen anything that allows querying other documents.

That doesn't mean it's a design direction that open knowledge should go in, but just a data point that reducing Obsidian vaults to "just markdown" misses what some users use it for.

jfim··on Steam Machine launches today
Yeah, I hope there's a gradient between scalper buying a single 99 cent game and actual person who has time played on multiple games and multiple purchases. The probability of someone with say a 3000 hours played across multiple games with hundreds or thousands of dollars of purchases is far less likely to be a scalper than someone with a single purchase many years ago and maybe some completely f2p play in say dota/cs, since the latter is likely to be a bot account.
jfim··on Ubiquiti: Enterprise NAS, Built on ZFS
Some do. I got the TS-873A a few years back, it works. Their software is kind of weird, and I wouldn't connect it to their cloud offering, but it does work.
jfim··on Ask HN: What are tools you have made for yourself since the advent of AI?
Cool! Just a heads-up that some of this is in a pretty rough state, but shoot me an email if you have any questions or issues.
jfim··on Ask HN: What are tools you have made for yourself since the advent of AI?
A pile of various tools:

A self hosted web archiving tool with support for extendible processing pipelines (eg. extract article -> translate -> summarize -> generate tags, download video -> split audio track -> transcribe -> summarize), which led me to make a managed chromium browser with extensions and warc support for archiving, and a RSS feed synthesizer (take random article listing page that doesn't have RSS and generate a feed for it) so that I can plug it into my archiver. An active learning loop for a model to clean up articles by removing junk like native ads and sponsored blocks.

A tabbed terminal with project management features like launching the database, app server, and claude code in different tabs with one click, and split browser/terminal panes (eg. opening a browser automatically at the correct URL when the terminal reads http://localhost:4000/).

A modular MCP server with a MCP proxy and OAuth2 dcr so that I can easily add new random ideas for MCP servers in a few minutes with Claude and deploy them such that it's available to Claude by refreshing the tool list.

A small tool to render Claude conversations so that I can link to them from my obsidian vault with something like convo://claude-code/-home-jfim-projects-foo/<guide>

And overall just deploying docker containers for my self hosted setup

Most of it is on GitHub, in various states of readiness.

jfim··on Valve P2P networking broken for more than 2 months
The Epic store is horrendously slow though. I bought a few games there but in practice the client is just so slow that I avoid it if I can.
jfim··on The Case for Space Datacenters
The article makes a lot of points about cost viability, but says nothing about what happens at the end of life for space datacenters.

On Earth, the materials and equipment in the datacenter can be repurposed, recycled, or properly disposed of. In space, EOL'ed stuff either stays in orbit, burns in the atmosphere on reentry, or moved out of useful orbits.

I'm not sure I'm thrilled at the idea of more space junk in orbit or more aerosolized metals in the stratosphere.

jfim··on "Maybe later" was a feature
That's a pretty good point, and I assume at $work they wouldn't appreciate throwing away $n dollars worth of code.
jfim··on How LLMs work
Indeed. It's pretty interesting to realize after implementing GPT-2 that the frontier models are scaled up versions of that, with various tweaks to improve performance, model-wise.

The secret sauce though is all the datasets, RL training, knowledge of what works from doing all kinds of ablation experiments, and a massive compute moat.

Page 1 of 34Next →