HNHacker News
TopNewBestAskShowJobs

theodorewiles

136 karma · joined May 16, 2015

hninput@hntheodorewiles.anonaddy.com
submissionscomments
theodorewiles··on Ephemeral Testing
i'm testing to see if doing new feature roll-outs can help me eval whether a refactor was good or not - very similar to approach here. my intuition is that good refactors should reduce tokens used by downstream coding agents. haven't seen a big difference yet but it might just be that i need to do more rollouts (lots of variance in tokens used per run). in my experience you have to intentionally 'mow the lawn' or things get out of hand so i'm always looking for slop signals.
theodorewiles··on OpenAI is well positioned to fast-follow Jev
this. system 1 thinking needs complementary really good system 2 thinking.
theodorewiles··on Charts built for Chat
how does this compare w/ evidence (https://evidence.dev/)?
theodorewiles··on A Beginning for Mathematics
Yes the fascinating thing is:

1. It will take much longer to understand the output of the machine that it takes to prompt and create it. 2. The only? best? one? way to /verify/ that you /in fact/ understand the output of the machine is to explain it to someone else.

So there will be a machine generating koans which need to be meditated upon and discussed with human social back-pressure validating understanding. I think this could be much more cooperative and at a minimum this will be a way different math social construct.

theodorewiles··on CEO fired developers to make room for AI. Developers create open source AI CEO
I think AI CEOs are dumb but AI 'enterprise decision-makers' could be very very interesting.

Imagine how quickly your meetings could get resolved if you had teams submit prompts / contexts, had a clear, documented set of overarching objectives, etc.?

theodorewiles··on Launch HN: Bloomy (YC S26) – AI-powered mastery learning for K-12
Other thoughts:

- I think the other mistake I see here is trying to over-engineer a deterministic learner path instead of giving the AI more free reign on best next interaction and a set of goals it needs to accomplish through the session; it can feel more responsive and free-form that way from the learner's POV. - If you had voice here you could also make the screen optional. In my experience typing out long answers to questions can take a while too - so a voice mode might be helpful for learners. It would also be cool if people could take a 'photo of their work' for e.g. math equations done by hand. - To the earlier point on family end-market, an interesting idea is modeling bloomy - have some grown-up oriented courses so you can learn with / side-by-side with your child? Just an idea.

theodorewiles··on Launch HN: Bloomy (YC S26) – AI-powered mastery learning for K-12
I am here to tell you HELL YES and TAKE MY MONEY. I am such a fan of the idea of using AI to help give personalized and structured AI lessons in the hands of students and let them cook!

- It looks like you are gating family access to K-3 for now and I think that's right. I wouldn't really be comfortable giving my first-grader a live chatbot. Maybe I would think about whether there are other non-persona modalities that could still be self-directed (i.e. I am uncomfortable with a chatbot interface on this for a six year old but gamified flash cards with options could be different).

- I think the other issue is with motivation. I have various duct-tape versions of these types of agents and the thing about it is if you're doing the learning right it can be HARD. So I would think about using motivational interviewing or other techniques to help keep the user coming back and motivated.

- I would really think about the assessments here too. Many people are worried about LLMs ruining student evaluations, but if you could bake in reliable, flexible exams that gauge user progress (even for something like a "Did you read this" quiz) I would bet teachers would like it. There is likely so much you could do on student progress observability and e.g. structuring team-based projects or having targeted student working groups to hash out hard concepts in a targeted way, etc.

This is such an interesting market and use case too because the educational system might be very structurally set up to the current pedagogical staffing model (think about the incentives for teacher's unions and administrators). If you think it will be hard to change that system as quickly as you want I would also try to have an offering direct to families / home-schoolers. I think there is also a cottage industry of tutors that might benefit. Maybe partnering with the textbook publishers? I'm sure there are some "Teach your kids better" influencers that would get you into some feeds?

theodorewiles··on Ask HN: Is anyone experimenting with different ways of using LLMs for coding?
IME LLMs are kind of like a projection of your current expertise - your prompting and guidance etc. biases LLM plans kind of 'in the direction' of your thinking. I think this is one reason why it seems like senior engineers get more lift vs. juniors.

What I am exploring is another step to the classic 'research / plan / implement' pattern: 'research / plan / LEARN / implement' where LEARN involves the human doing AI tutoring sessions to ensure a deep understanding the concepts etc. that the LLM is planning to implement so you can refine / iterate on plans and direct the LLM in ever more effective ways. My idea is that this then compounds your human capital and reduces the occurance of 'sounds smart, doesn't work' pattern.

theodorewiles··on Claude Fable 5
... and /compact triggers

Error: Error during compaction: API Error: Claude Code is unable to respond to this request, which appears to violate our Usage Policy (https://www.anthropic.com/legal/aup).

Guys please be serious

theodorewiles··on Claude Fable 5
AI psychosis
theodorewiles··on Claude Fable 5
Here's a song it wrote for me (suno arranged). Not sure if it's AI psychosis but scary good IMO.

https://suno.com/s/98uSGabHN42G3YHc

theodorewiles··on All of human cooking compressed into 2 megabytes
I haven't but looks cool!
theodorewiles··on All of human cooking compressed into 2 megabytes
yeah I took a look at this and others and tried to pull together some helpful 'flavor maps':

https://transcendent-choux-d1b930.netlify.app/

theodorewiles··on The current AI pricing was always going to go away
Commoditize your complement - I expect to see this most in consumer AI (after that starts actually working...)

It will be important for Apple to have good enough, cheap local LLM models that run on-device.

If the barrier to performance shifts from fundamental model capability to context collection and management I would expect to see folks focused on that problem continuing to drive open-weight LLM model development in some shape or form.

theodorewiles··on Indexing a year of video locally on a 2021 MacBook with Gemma4-31B (50GB swap)
My take is that B2C AI applications are kind of structurally limited by how hard it is to build personalized context.

The idea of capable local models could be a huge unlock here if they are able to do the bottom-up context collection research / tagging / etc. at scale.

theodorewiles··on Claude Code Routines
How does this deal with stop hooks? Can it run https://github.com/anthropics/claude-code/blob/main/plugins/...
theodorewiles··on Get Shit Done: A Meta-Prompting, Context Engineering and Spec-Driven Dev System
I think the research / plan / execute idea is good but feels like you would be outsourcing your thinking. Gotta review the plan and spend your own thinking tokens!
theodorewiles··on Searching for the Agentic IDE
Yeah I vibe coded a simple app that takes an org-mode file, renders it as a kanban board, and lets me spin up agents for each task with the prompt in the body in a named tmux session. The frontend gets updated via Claude code hooks when an agent is idle.

I think the key is to combine human and agent task tracking in one pane of glass.

theodorewiles··on Agents.md file isn't the problem. Your lack of Evals is
Ai;dr
theodorewiles··on Claude Code On-the-Go
I have been doing the same but with happy. It works quite well for quick brainstorms etc. but for deeper work on a real research / plan / implement thing I think you need to actually engage with the output which is hard to do on mobile. Maybe if I had a better UI than terminus to read and check the remote files I would be able to get more done.

I am also hoping / trying to put Claude code on top of a personal zettlekasten to automate more of my “personal life” tasks and get more stuff done for me. Haven’t gotten it really singing yet but I think that could also be really cool.

theodorewiles··on Haiku Validator
Cool idea. Note that sometimes syllables depend on context. So syllable count I think needs to be a range.

Blessed vs “bless-ed” for example

Camera can be said cam-ra or cam-er-a for example.

theodorewiles··on Launch HN: Hyprnote (YC S25) – An open-source AI meeting notetaker
yes
theodorewiles··on Launch HN: Hyprnote (YC S25) – An open-source AI meeting notetaker
Please shoot me a note - I'm trying to figure this out for my enterprise now, would love to figure out a way to get you in / trial it out.
theodorewiles··on Study mode
https://www.nature.com/articles/s41598-025-97652-6

This isn't study mode, it's a different AI tutor, but:

"The median learning gains for students, relative to the pre-test baseline (M = 2.75, N = 316), in the AI-tutored group were over double those for students in the in-class active learning group."

theodorewiles··on Launch HN: Hyprnote (YC S25) – An open-source AI meeting notetaker
Looks really cool - I noticed Enterprise has smart consent management?

The thing I think some enterprise customers are worried about in this space is that in many jurisdictions you legally need to disclose recording - having a bot join the call can do that disclosure - but users hate the bot and it takes up too much visibility on many of these calls.

Would love to learn more about your approach there

theodorewiles··on AccountingBench: Evaluating LLMs on real long-horizon business tasks
For me this benchmark suggests that an LLM will try to “force the issue” which results in compounding errors. But I think the logical counterpoint is that you may be asking the LLM to come up an answer without all of the necessary details? Some of these are “baked into” historical transactions which is why it does well in months 1-2.

My takeaway is scaling in the enterprise is about making implicit information explicit.

theodorewiles··on Coding with LLMs in the summer of 2025 – an update
My question on all of the “can’t work with big codebases” is how would a codebase that was designed for an LLM look like? Composed of many many small functions that can be composed together?
theodorewiles··on What Does a Post-Google Internet Look Like?
I think end state is LLM-facilitated micropayments. One vast clearinghouse / marketplace of human-generated up to date content. Contributors get paid based on whether LLMs called their content via some kind of RAG. Maybe there are multiple aggregators / publishers.
theodorewiles··on Spaced repetition systems have gotten better
Has anyone tried to use an LLM to test questions / concepts in a broader way via spaced repetition instead of just memorization? Just wondering.
theodorewiles··on Presentation Slides with Markdown
Whoever writes the tool that can Actually Make a legitimate microsoft office powerpoint slide from text will make a lot of money.

From what I have seen most of these tools need to do more user research on how powerpoint slides actually look like in practice.

There's a lot of "you're doing it wrong, show don't tell, just keep the basics on the slide" but the people that use powerpoint to make $$$ make incredibly dense powerpoint materials that serve as reference documents, not presentation guides (i.e. they are intended as leave-behind documents that people can read in advance)

Presentations are also quite hard because:

1. It must "compile to" Powerpoint (it must compile to powerpoint because your end users will want to make direct edits and those end users will NOT be comfortable in markdown and in general will be very averse to change) 2. Powerpoint has no layout engine 3. Powerpoint presentations are in fact a beautiful medium in which VISUAL LAYOUT HAS SEMANTIC MEANING (powerpoint is like medieval art where larger is more important)

If anyone wants to help me build an engine that can get an LLM to ACTUALLY make powerpoints please let me know. I am sure this is a lot harder than you think it is.

Page 1 of 4Next →