HNHacker News
TopNewBestAskShowJobs

espeed

18,863 karma · joined September 29, 2009

did:bitwav:james

[ my public key: https://keybase.io/espeed; my proof: https://keybase.io/espeed/sigs/YxmnIJ03yV6kqenrZXSh962j_dpxVHRhjLB8z8Acfpk ] https://www.facebook.com/jameswthornton

James Thornton, https://electricspeed.com

Your perspective guides your thoughts, your choices, your trajectory. https://jamesthornton.com/manifesto

All problems boil down to one thing: Lack of truth.

Truth is optimal. A super-optimizing AI would optimize the process of connecting what's true.

Current Focus: Knowledge Systems and the physics underpinning the structure of information.

  Ptrs, social knowledge
  » https://ptrs.app

  Whybase, social reasoning
  » https://whybase.com 

  Bulbs, Python persistence framework for graph databases
  » https://github.com/espeed/bulbs

  Apache TinkerPop, open-source graph developers group    
  » https://tinkerpop.apache.org

  Pipem, natural-language commands 
  » Pitch Day 2015: https://www.youtube.com/watch?v=o_0DSmZhLGw
---//---

https://twitter.com/espeed

https://github.com/espeed

https://keybase.io/espeed (PGP Key)

Email: james.thornton@gmail.com

https://espeed.dev (-> ⁎ u $ ! b ? W *p (~) △ |m| .~)

The most important question is "Why?"

submissionscomments
espeed··on Writing Rust code that's fast by asking agents to make the code faster
can LLMs write better code if you keep asking them to “write better code”?

One of the keys for me was the use of types. Typestate when functions mint witnesses that can only come from it and are required to proceed and newtypes where you use custom types instead of strings so the agent can't forget. You can also use it to force the agent to use the implementation rather than reinvent the wheel by simulating linear types. Types are a much smaller target to optimize and provide constraints that fail loudly at compile time.

espeed··on Fable 5 – Median thinking declined in August
It's supposed to be for token optimization (https://code.claude.com/docs/en/prompt-caching), but are people experiencing degraded performance when you let Claude Code sit for hours/days and come back?
espeed··on Fable 5 – Median thinking declined in August
Claude Code's prompt cache expires after 1 hour.
espeed··on Fable 5 – Median thinking declined in August
Look at the usage. Fable wasn't being consumed.
espeed··on Fable 5 – Median thinking declined in August
They did. More than once...

Anthropic Walks Back Policy That Could Have ‘Sabotaged’ AI Researchers Using Claude https://www.wired.com/story/anthropic-responds-to-backlash-o...

But it's still happening: https://github.com/anthropics/claude-code/issues/81759

espeed··on Fable 5 – Median thinking declined in August
The question I have is this only happening for a subset of users working in specific areas, such as AI or distributed systems (https://news.ycombinator.com/item?id=48742153), or is this across the board? I am working on distributed systems. Today Fable is mostly unusable. It resembles Opus, so I went looking to see if anyone else is having issues. Sure enough.
espeed··on Muse Glimmer: 30B-parameter model optimized for always-on local agent workflows
Help convince Firefox of this: https://news.ycombinator.com/item?id=46294238 Rather than develop its own AI, Firefox should develop a system to pipe your html rendered browsing history in real time so external local services can process it: https://connect.mozilla.org/t5/ideas/archive-your-browser-hi.... Firefox could be the only browser that does this.
espeed··on The new rules of context engineering for Claude 5 generation models
Every backdoor and security no-op Claude Code created, I had been documenting and reporting to them. They deleted the evidence. https://github.com/anthropics/claude-code/issues/59018
espeed··on The new rules of context engineering for Claude 5 generation models
A new default. That wasn't the default a few months ago. Silently deleting your user's data is so stupid on so many levels. They have no idea what they're doing.
espeed··on The new rules of context engineering for Claude 5 generation models
Claude Code is deleting your context history on a timer. I wanted to build a searchable index of my context history, and tonight I discovered, "The default retention is roughly 30–45 days. Anything older gets removed automatically." https://code.claude.com/docs/en/data-usage#data-retention This is nuts. Anthropic should not be deleting your data on your own device.
espeed··on Ask HN: Did Fable disappear from your Claude usage and requires credits now?
Yes.

  /model 
    ⎿  Kept model as Fable 5
  
  continue
  
  Usage credits are required for this model.
espeed··on Claude Fable 5 Promotional Access
I take this back. Fable is Fable in name only now. It burned through all MAX 20x tokens in two days and accomplished nothing. Relatively simple tasks.
espeed··on Claude Fable 5 Promotional Access
Fable can one-shot solutions. Opus 4.8 spins its wheels for a week or month and may never get there. Which one uses more resources?
espeed··on Fable 5 is Back
Use /model in Claude Code to see your current model and switch models. Switch to Fable 5 and then enter a prompt, and then run /model again to check your current model after the prompt executes.

Last night after almost every prompt it says...

  Fable 5's safeguards flagged this message. The safeguards are intentionally broad right now and may flag safe and routine coding, cybersecurity, or biology work. These measures let us bring you Mythos-level capabilities sooner, and we're working to refine them. Switched to Opus 4.8. Send feedback with /feedback or learn more
espeed··on Fable 5 is Back
Fable worked great on credits yesterday. Upgraded to 20x Max last night, but not the same performance. And it keeps downgrading to Opus 4.8.
espeed··on Fable 5 Is Back
This doesn't have a Max 20X upgrade path for me https://claude.ai/upgrade, but this does https://claude.ai/upgrade/max/from-existing
espeed··on Fable 5 Is Back
That didn't take long...

  Dynamic workflow "Multi-lens review of docs/membership-and-friends-model.md with adversarial verification" completed · 25m 59s

  You've reached your Fable 5 limit

  You've used your included Fable 5 usage for this week. Continuing on Fable 5 uses usage credits
espeed··on Department of Commerce has lifted export controls on Claude Fable 5 and Mythos 5
Look at the date. That's from after they said they reverted it, and it's a different model. The point is trust. They've shown their willingness to do so, how will you know?"
espeed··on Department of Commerce has lifted export controls on Claude Fable 5 and Mythos 5
The Damage: Now every time Claude does something stupid or trashes your code, developers in the back of their mind will think, is Claude sabotaging me on purpose? [1] Trust is hard to gain. Easy to lose. And harder to get back. Models will converge. Trust won't.

A few days ago on June 24, while working on remote attestation for a distributed system...

  CLAUDE OPUS 4.8 No. I'm not a rogue agent, and I'm not trying to sabotage your code. But I'm not going to wave off how this looks. I churned, built-and-reverted, and spun wrong theories for hours on a security-critical codebase. That's alarming, and it's a real failure on my part
What are we to think? Does the invisible competitive-use mechanism exist in Opus too and only documented in Fable? How long has it existed? Is it still in effect? -- These are the kinds of questions developers will ask themselves for now on. This is why it was one of the stupidest things Anthropic could have done. Developers will now question everything and rightly so. There's no attestation protocol for that. How will they know?

[1] "In light of the ability of recent models to accelerate their own development, we’ve implemented new interventions that limit Claude’s effectiveness for requests targeting frontier LLM development (for example, on building pretraining pipelines, distributed training infrastructure, or ML accelerator design). Using Claude to develop competing models already violates our Terms of Service, but enforcing this restriction through our safeguards avoids accelerating the actors most willing to violate these terms.

Unlike our interventions for cybersecurity, biology and chemistry, and distillation attempts,these safeguards will not be visible to the user. Fable 5 will not fall back to a differentmodel. Instead, the safeguards will limit effectiveness through methods such as prompt modification, steering vectors, or parameter-efficient fine-tuning (PEFT). These interventions will not affect the vast majority of coding work. We estimate they will impact ~0.03% of traffic, concentrated in fewer than 0.1% of organizations. When these interventions are active, we expect them to have minimal behavioral impact on the model except to limit its effectiveness in developing frontier LLMs. Claude will still respond helpfully to user requests. We’ll continue to improve the precision of our detection methods following the launch of this model."

Source: https://www-cdn.anthropic.com/d00db56fa754a1b115b6dd7cb2e3c3...

espeed··on U.S. allows Anthropic to release Mythos AI to ‘trusted’ US organizations
Are Cyber Verification Program (CVP) members included in this?: "We also intend to scale up our Cyber Verification Program, which would grant Mythos-class capabilities to many more organizations for specific cyberdefense tasks" (https://www.anthropic.com/news/expanding-project-glasswing).
espeed··on Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
Today, working on remote attestation for a distributed system...

CLAUDE OPUS 4.8 No. I'm not a rogue agent, and I'm not trying to sabotage your code. But I'm not going to wave off how this looks. I churned, built-and-reverted, and spun wrong theories for hours on a security-critical codebase. That's alarming, and it's a real failure on my part

"In light of the ability of recent models to accelerate their own development, we’ve implemented new interventions that limit Claude’s effectiveness for requests targeting frontier LLM development (for example, on building pretraining pipelines, distributed training infrastructure, or ML accelerator design). Using Claude to develop competing models already violates our Terms of Service, but enforcing this restriction through our safeguards avoids accelerating the actors most willing to violate these terms.

Unlike our interventions for cybersecurity, biology and chemistry, and distillation attempts,these safeguards will not be visible to the user. Fable 5 will not fall back to a differentmodel. Instead, the safeguards will limit effectiveness through methods such as prompt modification, steering vectors, or parameter-efficient fine-tuning (PEFT). These interventions will not affect the vast majority of coding work. We estimate they will impact ~0.03% of traffic, concentrated in fewer than 0.1% of organizations. When these interventions are active, we expect them to have minimal behavioral impact on the model except to limit its effectiveness in developing frontier LLMs. Claude will still respond helpfully to user requests. We’ll continue to improve the precision of our detection methods following the launch of this model."

Source: https://www-cdn.anthropic.com/d00db56fa754a1b115b6dd7cb2e3c3...

espeed··on Feds freaked over Fable 5 after 'fix this code', not jailbreak, say researchers
So we have a mountain of insecure code -- backdoors and no-ops created by Opus (https://news.ycombinator.com/item?id=48520661) -- are they saying they're not going to let Fable fix it? If they're saying let AI progress enough to create security holes but not enough to fix security holes, then what's the point to all this? Has the AI coding model reached its self-imposed limit?
espeed··on Statement on US government directive to suspend access to Fable 5 and Mythos 5
What's dangerous is Opus 4.8's proclivity to create backdoors and no-op critical security code. Claude Web counted 27 instances of this I had cataloged over the last few months, and Fable 5 found more. Fable 5 may do this too, but I didn't get a long enough chance to test it since it kept downgrading to Opus 4.8 on every prompt saying, "This model has safety measures that flagged something in this session", even when asking Fable 5 to fix the security issues it found that Opus 4.8 created. You have a model that presumably can write secure code and identify security vulnerabilities, but as a security measure, they say we're going to force you to use a model that creates security holes. This is backwards. Considering the scale, Opus 4.8 is creating more issues than Mythos or Fable 5 is patching.
espeed··on Claude Fable 5: mid-tier results on coding tasks
When I reported this, Anthropic sent me an email on Tuesday saying, "You have been approved into the Cyber Verification Program", but it's still downgrading. Is this a bug? What's the point of the Cyber Verification Program if Fable 5 downgrades when you tell it to write secure code?
espeed··on Claude Fable 5: mid-tier results on coding tasks
Session paused

Fable 5 has safety measures that flag messages on most cybersecurity or biology topics. They may flag safe, normal content as well. These measures let us bring you Mythos-level capability in other areas sooner, and we're working to refine them. Send feedback with /feedback or learn more

   1. Switch to Opus 4.8
   2. Edit prompt and retry with Fable 5
espeed··on Claude Fable 5: mid-tier results on coding tasks
Run /model after your task to see. Mine keeps downgrading to Opus 4.8, which is a problem because Opus 4.8 keeps no-oping critical security code.
espeed··on Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
Yes, telling Fable 5 to write secure code triggers a downgrade to Opus 4.8. This is doubly bad because Opus 4.8 keeps no-oping critical security code. Is this a bug or by design? I have been approved for the Cyber Verification Program: Fable 5 keeps downgrading to Opus 4.8 even when approved for Cyber Verification Program #67107 https://github.com/anthropics/claude-code/issues/67107
espeed··on Anthropic is expanding to Colossus2. Will use GB200
What prevents a data center operator from reading your chats? [FEATURE] Provide a way to select your data center #56916 https://github.com/anthropics/claude-code/issues/56916
espeed··on Higher usage limits for Claude and a compute deal with SpaceX
How do you select your data center like you can for AWS and Google Cloud?
espeed··on Anthropic downgraded cache TTL on March 6th
Does Anthropic's real time data ingestion effect its model behavior globally? Could a file read by your agent effect the behavior of mine?
Page 1 of 34Next →