4,794 karma · joined January 30, 2013
https://blog.jean-francois.im/about.html
There are quite a few other mitigations that could be done by providers that aren't mentioned in the article.
That's it. No open enrollment, no surprise bills, no trying to figure out which ppo or hmo works best, no wondering if it's in network or not or any of that BS.
To scale it in human terms, say that you're a human driver and you're 2x better than the average human driver. If you wanted to prove that you're better than the average driver, you'd need to drive multiple hundreds of millions of miles to statistically demonstrate so.
Given that the rate is significantly lower for taxi drivers in aggregate, to prove that they're safer than taxi drivers if Waymo is 2x safer, they'd still need to drive multiple billions of miles to statistically prove that they're safer.
The one on idTech8 global illumination from last year is pretty interesting too.
> The standard answer is greed: rapacious ambulance operators, owned by villainous private equity firms, exploit patients at their most helpless. But I don’t think that’s actually what’s going on. Ambulance providers are chronically unprofitable businesses; margins are thin, crews are underpaid, and operators exit the industry every year.
Except the company probably approved a budget for AWS or another cloud provider, and basically gave a blank check to developers to deploy whatever is needed. So developers are going to just deploy MSK or whatever is trendy, instead of trying to get the most throughput from the servers they got from IT.
The RSS feeds though are pretty neat, it's what I use to fetch articles for archiving so that I can get a curated set of things to read on the go.
If the benchmark is to implement features that are part of an open source project, and LLMs have those changes as part of their training dataset, it seems that they could just give a verbatim or slightly modified version of the change in their training data.
And if one updates the benchmark to only incorporate code changes that are past the models knowledge cutoff, then the benchmark is less comparable over time, since the changes in the benchmark at time T and T+1 aren't the same.
For example, I have software that summarizes articles and classifies links on webpages to build a synthetic RSS feed, both of which use LLMs, neither of which need a SOTA model.
I'll probably use LLMs to bootstrap a dataset of native ads in articles, and there again, I don't really need a SOTA model.
If it's for more open ended tasks like writing code though, I agree that at this point SOTA models make more sense to use.
A model that writes code without knowledge of any language or library changes for half a decade is less useful. A 2021 era chatgpt would be quite quaint in 2026.
Right now the Chinese labs might have incentives to release their models for free, and maybe Google is happy to release open weights today, but I'm sure there are already bean counters at Google salivating at the idea of having Gemini in Chrome as part of a Google AI monthly subscription just like YouTube premium and other Google subscriptions.
For example, with the appropriate plugins like dataview and charts, it's possible to create dashboards, lists, and tables that update automatically based on data elements present in documents or documents themselves. I use it to have views over my to-do lists (daily routine items, tasks that are overdue, upcoming tasks, etc), make dashboards, and show lists of documents edited on a particular date.
I'd love to migrate away from Obsidian towards something that's not proprietary, but I haven't seen anything that allows querying other documents.
That doesn't mean it's a design direction that open knowledge should go in, but just a data point that reducing Obsidian vaults to "just markdown" misses what some users use it for.
A self hosted web archiving tool with support for extendible processing pipelines (eg. extract article -> translate -> summarize -> generate tags, download video -> split audio track -> transcribe -> summarize), which led me to make a managed chromium browser with extensions and warc support for archiving, and a RSS feed synthesizer (take random article listing page that doesn't have RSS and generate a feed for it) so that I can plug it into my archiver. An active learning loop for a model to clean up articles by removing junk like native ads and sponsored blocks.
A tabbed terminal with project management features like launching the database, app server, and claude code in different tabs with one click, and split browser/terminal panes (eg. opening a browser automatically at the correct URL when the terminal reads http://localhost:4000/).
A modular MCP server with a MCP proxy and OAuth2 dcr so that I can easily add new random ideas for MCP servers in a few minutes with Claude and deploy them such that it's available to Claude by refreshing the tool list.
A small tool to render Claude conversations so that I can link to them from my obsidian vault with something like convo://claude-code/-home-jfim-projects-foo/<guide>
And overall just deploying docker containers for my self hosted setup
Most of it is on GitHub, in various states of readiness.
On Earth, the materials and equipment in the datacenter can be repurposed, recycled, or properly disposed of. In space, EOL'ed stuff either stays in orbit, burns in the atmosphere on reentry, or moved out of useful orbits.
I'm not sure I'm thrilled at the idea of more space junk in orbit or more aerosolized metals in the stratosphere.
The secret sauce though is all the datasets, RL training, knowledge of what works from doing all kinds of ablation experiments, and a massive compute moat.