HNHacker News
TopNewBestAskShowJobs

jameswhitford

46 karma · joined December 31, 2024

submissionscomments
jameswhitford··on Agents feel free to take desktop screenshots?
I set my agent (Sonnet 5) to work on my KOReader extension on auto mode.

It started autonomously taking screenshots of my screen and reading them, without my consent or instruction.

Screenshots were being fired on whatever my screen was focused on, not just whatever the agent was working on.

A tad too relentlessly proactive for my taste. Agents deciding to screen capture whatever I am working on and sending it to a third-party without my consent?

This should probably be denied by default in auto mode.

jameswhitford··on Show HN: One More Letter
Awesome game! A potential UX update: if you fail a guess, the letters go back down below in the same configuration as your guess.

But others might disagree!

jameswhitford··on Show HN: Agent Draw: An agent draws while you talk, built on TLDraw
Eish, probably needs some bug hunting. During testing I had some issues with less capable models, if you have your own key for a more capable model it should perform better.
jameswhitford··on Ergo: Long Form Philosophy Lectures
Agreed! The world is a little unbalanced in this way right now unfortunately.
jameswhitford··on Ergo: Long Form Philosophy Lectures
This is sick where have you been all my life!
jameswhitford··on Show HN: Agent Draw: An agent draws while you talk, built on TLDraw
This is awesome! I love mermaid but have struggled to get my agents to be any good at it. Maybe I just need to context engineer a little better
jameswhitford··on Show HN: Agent Draw: An agent draws while you talk, built on TLDraw
It would be great if it could have examples of your own diagrams in its context, so it knows what points you like to map out and how you like to visualize them
jameswhitford··on Show HN: Agent Draw: An agent draws while you talk, built on TLDraw
That’s awesome! My inspiration was the IBM YouTube channel where they write on the board in front of them, maybe you could use the white canvas as like a green screen or something to get a similar effect?
jameswhitford··on Show HN: Agent Draw: An agent draws while you talk, built on TLDraw
The potential use cases was my favorite part of designing it, it’s not perfect now but it’s so fun to think what people might do with these kinds of designs in the future
jameswhitford··on Show HN: Agent Draw: An agent draws while you talk, built on TLDraw
Yes I would love to see some work on the design, because I think it could be useful, just needs some speed and accuracy improvements, and maybe some design towards specific use cases
jameswhitford··on Show HN: Agent Draw: An agent draws while you talk, built on TLDraw
Yes it could get a lot better with some design improvements, maybe live streaming the transcription, and maybe after some trial and error testing the best + fastest model for the job
jameswhitford··on Show HN: Agent Draw: An agent draws while you talk, built on TLDraw
Facts!
jameswhitford··on GLM 5.2 vs. Opus
That is a great suggestion that I am definitely going to look into, thanks!
jameswhitford··on GLM 5.2 vs. Opus
I hear you
jameswhitford··on GLM 5.2 vs. Opus
Cool to hear, what kind of tasks have you been using GLM for? And what other models have you found useful through Ollama?
jameswhitford··on GLM 5.2 vs. Opus
I see your point. Just the fact that one model does have vision and one does not might be an interesting point of comparison, however.
jameswhitford··on GLM 5.2 vs. Opus
This is excellent feedback thank you! These LLMisms in writing are a challenge I am living with currently and trying to improve on. The technical writing industry is taking a huge knock right now with companies demanding more work in less time with a big drop in quality, day to day I get less and less time to work on the quality in the prose of my work. We are working at the frontier of this right now, so we are the most heavily effected, but also get to experiment with the changes first which can be both stimulating and very frustrating.
jameswhitford··on GLM 5.2 vs. Opus
Hi, author here, can you link? I would love to read about this.
jameswhitford··on GLM 5.2 vs. Opus
Yes I agree 100%. My next guide would do better to use identical harnesses.
jameswhitford··on GLM 5.2 vs. Opus
GLM 5.2 is text only, not multi modal. And Opus is multi modal.
jameswhitford··on GLM 5.2 vs. Opus
Hi, author here, I cannot give an exact number for how many token the verification step took, but the verification GLM 5.2 ran was very stupid and definitely a waste of time. It read the pixel color data to try and verify the scene rendered properly. Which is really bad. Opus opened the game in a Playwright browser and took screenshots to verify the actual image. Which helped a lot.

Pro tip: You could use a multi-modal model to verify images as a subagent spawned by GLM 5.2, to get around this issue.

jameswhitford··on GLM 5.2 vs. Opus
Yes I 100% agree. Time-taken can be improved (with harnesses, subagent workflows etc.) and varies based on task.
jameswhitford··on GLM 5.2 vs. Opus
Yes, part of the reason I chose the one-shot test was really to test long-running tasks. A lot of people seem to be experimenting with this format, for example in the now trending loop-writing workflows. And really I am interested in diving into the murky waters of these novel workflows.
jameswhitford··on GLM 5.2 vs. Opus
I appreciate the feedback!
jameswhitford··on GLM 5.2 vs. Opus
Yes this is true. This test was run on a $20 pro Claude subscription. I would definitely love to try use both models on the highest plans for a whole month and compare the two, great format for a future head-to-head comparison.
jameswhitford··on GLM 5.2 vs. Opus
Hi, I am the author, I completely agree! I set out to run a vibe test on this one, not a benchmark, the real benchmarks are listed. My test shows what the models can do when both tasked with a long-running, technically difficult, one-shot task.

I think your test you describe (collaborative, task delegation, task completion, TTD, steerability) is a great format for a future test that I will definitely try out.

jameswhitford··on Claude is skeptical about OpenClaw
I asked Claude Code to research Openclaw. It spawned a subagent, got back detailed results, and then flagged them as unreliable and/or hallucinated before I could read them.

TL;DR:

Claude isn't trained on openclaw data due to its knowledge cutoff, but this is the first time I have been asked to look at research myself to verify it isn't hallucinated or unreliable.

I am not making any claims about Anthropic training their models to perform worse when dealing with information about competitors...

But I am worried about this behaviour of flagging certain sources as unreliable for what seem like arbitrary reasons.

It could also be a case of prompt poisoning at one of the research URLs.

jameswhitford··on The Case That A.I. Is Thinking
Who would not want to say their product is the second coming of Christ if they could.
jameswhitford··on The Case That A.I. Is Thinking
This submarine isn’t swimming, it’s us that are submarining!

I think I hear my master’s voice..

Or is that just a fly trapped in a bottle?

jameswhitford··on How to Migrate from OpenAI to Cerebrium for Cost-Predictable AI Inference
It's a demo project using the free tier hardware from Cerebrum, demonstrating how to migrate with a few lines of code from OpenAI. The cost is never going to beat OpenAI on an A10, there are more powerful options available.
Page 1 of 2Next →