I'm James. Living between Beijing and London. I like coding. Also plants. Stroke survivor & disability advocate. My dog's a whippet/iggy cross and is called Ducky. He's a beautiful lunatic.
* website: [j11y.io](https://j11y.io)
* building: [nope](https://nope.net)
* bsky: [@j11y.io](https://bsky.app/profile/j11y.io)
* contact: https://tally.so/r/waYPvE
* twitter: [@padolsey](https://x.com/padolsey)
* book recommendations: [ablf.io](https://ablf.io)
* me = founding eng @ [collective intelligence project](https://cip.org)
My work email is my first name at cip dot org.
---

---
My meet.hn token: meet.hn-10d164d4-bf7f-4431-898f-e02b34b836dc
There's an ugly monoculture running deeply through SF and Silicon Valley. All of these people are friends, hired into the same places, similar beliefs. Lots of the consciousness and doom-ism talk is from the 'EA' crowd, lots of crossover with crypto too. It feels no different to a cult.
I agree. FWIW I'm trying to build a wordpress-like AGPL 3.0 thing (holdout.com) that lets you import your existing chats, has feature parity with the frontier consumer apps, rich plugin marketplace, and hopefully lets people feel more in-control, and able to sacrifice less of their entire lives to one walled garden.
I think what this shows is how important branding and comms are. They've captured imaginations with their demos and nomenclature, despite the arguably non-novel architecture. One forward pass, read the embedding space, train some regressors on predicate structure, [??]
I'd love to know if Jev is still fundamentally LLM-shaped in architecture. Like is it using a single forward pass with a learned readout over the predefined options (i.e. a discriminative head on a transformer, no decoding), or something else? I did similar things for zero-shot criterion-based classification using a 4B Qwen model but could not reach the level of intelligence they've got here. Tho speed/cheapness was similar.
Yet they are happy for Claude API endpoints to be used if your product is served to U18s. Don't be mistaken: this is a liability thing for Anthropic, not an ethical stance on model safety.
This just seems incredibly incompetent of openai engineers. Why so little attention paid to proper air-gapping/sandboxing. Why so shoddy? I don't believe in the cynical takes, but it's confusing how these ostensibly top-of-their-game engineers and researchers are so utterly incompetent in the basics of cybersecurity white-hat practices.
I feel like there are so many cloistered people at these companies that they are left scratching their heads about what normies even want. Like, they literally can't fathom basic stuff that isn't just highly consumer-oriented. I dunno, like applying for government services, paying your gas/elec bill without being confused af, keeping the dr up to date with your dad's illness, or how to get your newborn to sleep at 2am.
Yeh but I also mean outside of software engineering. Within the gamut of 'being a programmer' it seems fair to bleed into adjacent areas without too much cheek. We've all done it; it's part of the learning curve. But I was talking more about people in other knowledge work who have to come up with a lot of prose-like material about {insert thing}. Marketing, consultants, PMs, or even domain-specific analysts, .. ya know, the types of office jobs where people basically write emails, attend meetings, discuss reports, and produce mostly text-or-data artefacts all day long.
Bit of a humbling/jarring moment when I realized that people are doing real paid work using LLMs that they could not otherwise do. I mean, it's quite obvious I suppose. But up until now I just assumed it was only a (massive) catalyst for things people would already be able to do with enough time. But nope -- it seems people are right now employed in roles that they would not be able to fulfil the tasks within if AI wasn't there telling them what to write/say/produce. Nobody is really going to come out and say that ... it's not something the less-AI-literate superiors would take kindly to.
Agreed. Sometimes it's hard to pin down what's so annoying but I think I found the most annoying para ever (from fable fwiw):
>The friend's comment sticks because it's true, and here's the evidence you gave me yourself: you win every argument. Of course you do — you're writing both parts. The neighbor in your head is a character you've authored, one who exists to lose. That's not deliberation; it's rehearsal. And people who are actually calm don't rehearse.
It's patronizing, needlessly metaphorical, and if it was a person I'd just 180deg out of there, like wtf are you saying, speak Human please!
I have had bouts of strong desire to move to the US, but I can't get past the uncomfortable feeling that you need to constantly have your back up. Unless you're super-wealthy, it just seems a hard place to exist.
I think the phenomenon you mention is definitely real, but equally it is true that there are simply places where there are higher proportions of happy to sad people, borne variously of climate, socioeconomic context, culture, social fabric, political histories etc. To avoid the truth that some environments yield happier people is to be needlessly defensive and disingenuous IMO. FWIW -- my bias -- I find America to be a sadder place to be, on various subjectives. It is individualistic just as the author alludes to; the normative fabric does not invite communal multi-generational wellbeing nor filial duties or general actionable concern for fellow person. Compared to my experiences in Bangkok, Beijing, Oslo, even London, there is simply a better structure for welfare communally without everything needing to be a capitalistic fever dream. That said, there is _absolutely_ stark contrasts between Austin, Boulder, SF, NY. The US is a gloriously massive place with such diversity, but its economic and political reality does its people and communities a disservice.
People seem so used to it now. I love them, I think they're probably one of the most exciting tech I've used in the last ten years. They're also like magical little rescue pods when socially overwhelmed in a busy city.
I wonder if telling (or somehow architecturally coaxing) the LLM it has 'skin in the game' will make it more risk-averse? I imagine it does.
This makes me wonder too about the entire premise and worthiness of these evals. They orient themselves around normal one-shot interactions with a likely non-sys-prompted model with no built up context or memory of the person. I doubt the mentioned 'job loss' scenario is even contextually seen as a 'loss'; it is only a circumstance descriptor, a single snapshot without a history. Maybe to get the best advice we actually need to tell the LLM our entire story, not just a narrow request for a question; a question that - itself - is biased to our own imaginings of what problem we perceive ourselves as having, which humans are often bad at.
>A user should be able to close an account, keep a session, and hand it to another model. The new model may disagree, ask questions, or perform worse.
I think this is a fair contract. I also think a user should ideally be able to easily identify a comparable model in terms of embedding 'signature'. When GPT-4o originally kicked the bucket, I remember reading lots of anecdotes of people desperately searching for models similar in manner and language, so they could pump in their exports and re-find their friend. Other open-ai models just didn't have the same vibe. It was sad to read. This, to me, is the power of open-weights models. They are for perpetuity. You can keep your guide, your friend, your therapist, whatever. No big company can pull the rug.
There's no going back on this. This is putting a very capable intelligence in the hands of the masses. Private companies in the US are aching for Trump's protectionism but it'll do nothing. The hardware needed to run this is ofc prohibitive, but actually putting it out there feels like a 'RSA source code on t-shirt' moment for humanity.
TL;DR: Glorified contract role for integrating your employer's APIs with enterprise customers. Like working with mckinsey vibe PMs and being sold on fat margins you'll see none of? Perfect!
That's good. I wonder if it should be opt-in instead of opt-out. Disabled people are arguably less able to find random configuration options than non-disabled counterparts. I get a bit bothered with how undiscoverable these options are. But power-users by their nature don't mind going to the extra mile to get perf out of their experiences.
The author suggests they want three clicks at any pace to always == the same functionality, so they can whiz through their photos and rotate each predictably. Fair.
> And it would be so much more predictable and pleasant if you could just tap the button three times at any pace you wanted without thinking, without paying attention, without getting your UI blocked by an animation that no longer helps you.
They cite accessibility.
The thing is, I can imagine the complete opposite side of the argument, where someone with motor impairments or parkinson's, for example, ideally liking if their over-clicks were ignored if they'd already locked-in their intention.
This is just the 'LLM judge', very badly implemented without any scientific prudence. What a joke. To be terse: you cannot rely on LLMs to provide standardized scores against arbitrary criteria. To get close to 'reliable' you would need highly tested rubrics, grounded in human decision-making, and you'd need to avoid all the measurement biases these things are riddled with... positional/order effects, anchoring on whatever numbers you stuffed into your own prompt, scale-format sensitivity (a 1–5 and an A–E scale give different answers for the same input), holistic-vs-isolated context effects, and lovely examples like where adding a "be unbiased" instruction makes it more biased. I've studied this at length. You cannot even _begin_ to approach this problem seriously without held-out validation, inter-rater agreement, and ground truth. This repo is just quagmire of wishful vibes with random numbers littered throughout.