HNHacker News
TopNewBestAskShowJobs

docjay

130 karma · joined September 23, 2025

submissionscomments
docjay··on Understanding the Impact of LLM Watermarking on AI Agent Behavior
You’re using the subjective definition of “quality”, as in the shade of blue it chooses for “Build a website”, or the character names for “Tell me a story.” In those cases it’s likely still subjectively “high quality”, depending on who you ask.

What the article is discussing, and what many people are concerned about, is something that you might be missing in your understanding: they’re not actually random. In fact, they would be entirely useless for real work if every token was randomly selected based on all possible outputs. It’s not, even at temperature 1.0. It’s based on the training corpus and once you have your tool names, syntax, and prompting style aligned with the training data then they become incredibly deterministic in the areas that matter, such as tool calling and parameters. I build my toolset by testing thousands of names, syntax, return format, and other aspects until I find a convention that produces the exact correct call, 100% the exact same every time, regardless of context length. Those decisions are per model and what works with Opus 4.7 won’t necessarily work on 4.8, and neither version will work with a local model or GPT.

That’s only possible because the massive training corpus is the guiding principle behind the choices. Providing a file reading function called “Read_The_File” will fail, either on the first call or somewhere down the line, because that name is not associated with the concept. Your instructions are trying to override 500 trillion tokens from training and it will cause perplexity to manifest as wrong tool calls, wrong syntax, “oops deleted prod”, “Claude lost the plot again”, “WTF?!”, and probably nearly every frustration you’ve encountered and determined to be “they nerfed Claude” or “it’s a dumbass.”

For those that are aware of it, that knowledge lets people tweak and tune the prompts/tools accordingly.

You may not put that effort into your system, perhaps because you’re unaware of it, don’t use it in a way that requires it, or you’ve just taken the failures caused by perplexity as something that’s inherent in the framework, but for people that build precision infrastructure around them it’s potentially devastating news. Watermarking, which is based on whatever tokens, threshold, cutoff, and triggers some guy at a desk decided, will necessarily alter that entire system.

docjay··on iOS 27, iPadOS 27, and macOS 27
Good info, I’ll check it out. Thanks!
docjay··on iOS 27, iPadOS 27, and macOS 27
Weird, the same stupid crap happens with Windows. I’ve had to learn the exact N character search to find the things I use often; any more or less and the result with that name is buried. It’s like Microsoft and Apple are just pip installing same broken search package.
docjay··on Notes on gotchas while migrating 35kb preprompts from Opus to self-hosted Ollama
You’re right about the harness being the issue. It’s really down to giving it functions specific to your use case that will let it surgically read/modify files, rather than needing to consume entire project folders. I built my own and for Python files some of the most helpful functions I provide are equivalent to:

inspect_function(filename, function, class)

replace_function(<same>)

call_graph(<same>)

And a few other convenient ones. Beyond that it’s trickery like if a function returns more than N lines I omit the result and auto-reply “Your function call was too verbose.” Typically that’s stuff like recursively listing every file in a repo to “see what it’s working with” or similar. When it emits the next call in response to it I clip the previous attempt (and my response) off the conversation and attach the new call/result as if that’s what it did in the first place. I also log that event so if the same type of thing happens often enough I’ll create a special function to address it, or modify an established one so it’s not tempted to do it again.

Language models don’t know what they know, they know what has been said. Even if you give it an entire Python environment it won’t reach for AST, but if you give it a function called Python_AST() it’ll use it every time.

I’ve never seen an off-the-shelf harness that approached it that way.

docjay··on I-have-ADHD: A skill to stop coding agents from burying the answer
Style requests tend to get washed away by context. I’m not sure how you’d do it with Claude Code or whatever (I use the API directly), but you might try fiddling with giving it a function called “response” or “message_user” with a mandatory parameter like “output=plaintext”.

I’m basing that idea on custom functions I give Claude where I designate the style in the required parameters. Example:

“””

@pyrepl(code_golf=True, output=CSV)

//END

“””

The “//END” is a special terminator I use for halting the response. Claude runs code between the tags. The difference in code is incredible. No banners, comments, fluff, print(“=“*70), or any other nonsense, and the code is tight and compact. The parameters do nothing, Claude just outputs them as part of required syntax and it dictates the style purely because it output them.

docjay··on Check if a file was made with Claude
Not really. This is more like not pumping the exhaust back into the intake.
docjay··on Clean up Claude 5's token vomit with a separate LLM
You don’t have to be so convincing when it’s a local model.

```Un-Claude 0.2beta

    import sys,csv,requests
    
    CH="# Valid channels: analysis, commentary, final. Channel must be included for every message."
    
    CANDIDATES=[
        ("no-hedging","Reasoning: low\n\n<terse><no-hedging>\n\n"+CH,"Condensed:"),
        ("neutral-reg","Reasoning: low\n\nRegister: neutral technical. No intensifiers, no evaluative adjectives.\n\n"+CH,"Condensed:"),
        ("no-closing","Reasoning: low\n\n<terse>\nNo closing remarks.\n\n"+CH,"Condensed:"),
        ("terse","Reasoning: low\n\n<terse>\n\n"+CH,"Condensed:"),
    ]
    
    def rephrase(text,base="http://127.0.0.1:1234",model=None,temperature=0.0,max_tokens=1400,timeout=180):
        src=text.strip()
        if not src:
            return []
        if model is None:
            model=requests.get(base+"/v1/models",timeout=timeout).json()["data"][0]["id"]
        w=csv.writer(sys.stdout,lineterminator="\n")
        w.writerow(["idx","label","prefill","src_chars","out_chars","ratio","tokens","finish"])
        rows=[]
        for i,(lab,sysmsg,pf) in enumerate(CANDIDATES,1):
            p="<|start|>system<|message|>"+sysmsg+"<|end|><|start|>user<|message|>"+src+"<|end|><|start|>assistant<|channel|>final<|message|>"+pf
            d=requests.post(base+"/v1/completions",json={"model":model,"prompt":p,"max_tokens":max_tokens,"temperature":temperature},timeout=timeout).json()
            c=d["choices"][0]
            t=(pf+c["text"]).rstrip()
            w.writerow([i,lab,pf,len(src),len(t),round(len(t)/len(src),3),d["usage"]["completion_tokens"],c["finish_reason"]])
            rows.append((i,lab,sysmsg,pf,t,d["usage"]["completion_tokens"],c["finish_reason"]))
        print("\nmodel: %s"%model)
        print("temperature: %s   max_tokens: %s"%(temperature,max_tokens))
        for i,lab,sysmsg,pf,t,tok,fr in rows:
            print("\n[%d] %s"%(i,lab))
            print("    system:  %s"%sysmsg.replace("\n","\\n"))
            print("    prefill: %r   tokens=%d   finish=%s"%(pf,tok,fr))
            print(t)
        return rows

```

```input

## 8. Honest gaps — what I could *not* resolve

I want to be explicit about the limits of this pass rather than imply completeness:

1. *`PROVIDER_T` values are not enumerated here.* `list_models(inference_provider=...)` is typed against `PROVIDER_T`, which lives outside the three modules I scanned (it's in the `inference._providers` subpackage). The accepted provider strings are therefore *unknown from this run* — `"cohere"` is confirmed only because it appears in a docstring example.

2. *Three grep hits point to search-capable functions I did not identify.* My scan found parameter assignments that don't belong to any function I enumerated: - line 3046–3050: `params["filter"]`, `params["sdk"]`, `params["includeNonRunning"] = True` — an additional Spaces-oriented endpoint with an *`sdk` filter and an `includeNonRunning` flag* not exposed by `list_spaces`. - line 2879: `params["config"] = config` - line 12013: `"sort": sort` — almost certainly the consumer of `DailyPapersSort_T`, i.e. a daily-papers lister distinct from `list_papers`. - line 13872: `params["search"] = search`

   These represent **real additional search surface** that my `LIST_FUNCS` whitelist missed. A follow-up pass enumerating every `HfApi` method containing `params[` would close this.
```

```example output

[1] no-hedging system: Reasoning: low\n\n<terse><no-hedging>\n\n# Valid channels: analysis, commentary, final. Channel must be included for every message. prefill: 'Condensed:' tokens=131 finish=stop

Condensed:

- *Provider strings* (`PROVIDER_T`) are not listed; only “cohere” is known from a docstring.

- *Missing search‑capable calls* were found: - `params["filter"]`, `params["sdk"]`, `params["includeNonRunning"] = True` (Spaces endpoint). - `params["config"] = config`. - `params["sort"] = sort` (likely a daily‑papers lister). - `params["search"] = search`.

These were not captured in the `LIST_FUNCS` whitelist, indicating additional search functionality.

```

Edit: yeesh, I’d love to have a WYSIWYG comment block on this site. I’m not going to keep fighting newline and white space to get it to look right, but you get the idea.

docjay··on Every Fucking Website (2020)
I have the same mindset and often did the same thing, but then I thought about my doctor sneaking into my house while I’m sleeping and injecting me with the “good medicine” I had refused in their office.
docjay··on uBlock Origin is giving up the fight to keep ads off Facebook
You got it. A company with 15% of the foodservice market somehow supplies 100% of the food in 100% of restaurants and 100% of Americans eat at them 100% of the time. We should probably just pay a tax and have the government take it over so your imagination becomes our reality.
docjay··on uBlock Origin is giving up the fight to keep ads off Facebook
The earlier link you provided is indeed refreshingly quick to navigate. What a joy that’s unnecessarily rare. It renders like trash and is unusable on my intentionally-not-updated iPhone, but still. Yay.

“Provide incentive programs and reward those that take on such initiatives.”

We have that too, we call it capitalism. It’s not very unique, but we heavily market it under the “land of opportunity” moniker. Do a thing of value and your reward is the lion’s share of that value. “You get what you’re given” seems to make some countries happy, but I’m sure you can understand the appeal of “you get what you earn”, even if neither of them matches the brochure. The scholars can debate the pros and cons of it, and where it goes wrong, but you’re talking about investment reward structures like it’s a fresh concept coming out of Scandinavia.

On that note, let’s set aside the endless conversation around “government Facebook” and talk foundational: Size, diversity, and population. It’s funny when I see comments from micro->small nation states with relatively homogeneous citizens throwing up their hands like “Everyone should just agree on what to have for dinner, like we do.” Oh I’m sure you’re aware that ‘America be big’, maybe even visited New York and were impressed by all those not-tall non-blonde people waking around, but perhaps without fully appreciating that it alone has a 50% higher population than your whole country. Something more your speed would be the Atlanta metro area, which still has a ~10% higher population. Getting everyone to happily chip in a few bucks on a half pepperoni/vegetarian pizza (or, I assume more familiar to you — “Grøt med smør eller rørte bær”) is a lot easier when everyone can be fed by that pizza.

But that leads to the larger point: we don’t have as much of a problem getting our Norway sized cities to provide various communal welfare services to the local residents, the problem is in scaling that up to 65 Norways.

docjay··on Muse Glimmer: 30B-parameter model optimized for always-on local agent workflows
No problem. That’s where I spend a solid 98% of my time because it’s worth it, and I can show the measurements.

Couple tips:

A first pass is to blank out the system prompt, add only one tool, and work to reduce the thinking length for a direct function call prompt. “Read archive.log” should result in roughly 0 length thinking. If it’s thinking about anything, especially if it mentions {readfile tool}, rename the tool and minimize the description. Depending on the model it might always output thinking, so run it until you get a consistent outlier that’s far lower thinking length than the others. It’ll be obvious when you find it. Repeat the prompt dozens of times, modify it slightly, and focus on the lowest max length, not average.

You should really use a Claude with Python (or similar preferred) to make API calls to the LLM and have it iterate through hundreds of names/descriptions and return only len(thinking). Have it build a batch testing harness to run a dozen tests at a time, that helps keep it from ‘cheating’ to finish. laziness = count(messages), but frame it as an academic research project studying the effects of minimalist tool descriptions on thinking length. Don’t set the goal as minimal thinking length, Claude will short circuit it.

Remove all other tools until {readfile} is perfected, then add/test the next tool. Btw: you don’t need to describe readfile() when it’s named right.

The built-in tools[] makes that hard because it tacks a really dumb system prompt on at the server and requires some length of description, which is why I built my own function calling, but that’s still a good first pass. Focus almost entirely on the function name itself; readfile, readFile, read_file, readlines, file_get_contents, etc., and make the description just “Operational” or similar. Field description, if required by API, is literal “filepath”, same as field itself. Lowercase, nothing else said. Minimize your contribution to perplexity, use standard naming conventions.

When you add a second tool you need to still include the first tool prompt in the second tool testing. Adding {writefile} can absolutely break {readfile}. Have Claude run the tests and build it out into permanent testing module with file_read=[prompts], file_write=[prompts], making it easy to extend, and full_test() that runs them all to see if a new addition broke it.

Add your system prompt back in and probably watch the tests go to shit. <- THAT is likely your biggest problem. My system prompt for the main LLM has all of the tools it can use, which is ~30 lines of function names with no call syntax, and yet it has more tools than Claude Code and never messes them up.

Start with nothing and slowly work up. Focus on positive action framing, not negating: “Your responses are always..” and not “Do not…”

It sounds like a pain, but building the systems to automate the tests IS the infrastructure, everything you have it do afterwards is just the tasks.

That was longer than I planned, but I guess this’ll be a comment for future generations to find.

docjay··on Muse Glimmer: 30B-parameter model optimized for always-on local agent workflows
I was afraid my description sounded like agents, and it kind of is, but not like most implementations. Most use a “boss bot” to craft a prompt/system message and launch the model, and sometimes they redo it every time it launches the agent. It’s a low effort attempt that’s immediately flawed because it uses LLM output for LLM input. It can look like it’s working for some time, but the perplexity guarantees it’s a roll of the dice. That’s what eats away at these kinds of projects. “It was doing great until it rm’d prod.”

Run the same exact prompt 100x in a single step test (one prompt, one response), hash the full responses, and you’ll see 10-50+ unique responses. The higher the unique the worse your prompt; focus on the system. A highly tuned system prompt will result in one response, even at a temperature of 1.0. Really. Once that’s done that’s the only thing it does and it’s the only one that does it and it never changes. Other LLMs that call ‘check_email()’ are unknowingly just passing a prompt to the specialized one.

I use the API directly, craft a small Python script for the API call and task interface, then hyper-optimize the system/user prompt using test scenarios and automated loops. My system prompts rarely/never contain complete sentences, yet include all the tools/functions and requirements.

Make your error messages user prompt instructions, not errors. That’s why “agent optimized” models exist. Chat models are primarily trained on conversational text, meaning the stackoverflow “How do I fix ‘too many levels of symbolic links’?” -> Explanation/resolution. It’s far less on “# ls broken_loop” -> “# ls: cannot access ‘broken_loop’: Too many levels of symbolic links” -> “# namei -l broken_loop”

It’s not that the good ones are bad, but you’re leaning on the million training documents rather than the trillion.

Anyway, go that route with your system. Think about it more like automating a factory floor rather than hiring interns.

The least efficient methods, by definition, have the most room for improvement, which means they have the greatest reward potential, but for that one “eureka” moment. The path less taken is often interesting, but the ill-advised path still has fruit on the trees.

docjay··on Muse Glimmer: 30B-parameter model optimized for always-on local agent workflows
I’m curious why you don’t just use them like a Meeseeks box, rather than compressing and context stuffing into one. One only checks and categorizes your emails, another one for each category of email or even subcategory, one that only handles calendar additions, a different one to check it and notify you; you can go infinite with it. Hell, I’ll have one instance find a file and read it into the context of a different one because I don’t want a bunch of grep commands mucking up the context of the analysis. The find/read one exists for a few moments, as does the analysis one, and the ‘perform’ one is entirely different. I can run them all in parallel and use a queue if needed.

I’m sure you have reasons for your setup though, so I’m curious how you landed on it.

docjay··on The UK's war on anonymity has come to America
The problem is that the government was handed an emotion and a mandate to eliminate it. You’ve declared a “war on jealousy” and demanded the government assemble the troops. That’s not handing them a well reasoned plan and “hoping they don’t mess it up”, you messed it up when you scribbled “no more addiction” on a protest sign.
docjay··on Windows 11's built-in Weather app wastes more than 1 GB of RAM
I think the issue is with the recent definition creep of “tasteful.” Seems you might have the same mindset if you consider “minimal” to be the other option.
docjay··on Timeline of the OpenAI accidental attack against Hugging Face
From the moral perspective or the technical one?

Technically: it’s a function call that must return text. Imagine if you sat down at the command line and typed an initial command, then from that moment on every response required you to issue a new command. ping-pong-ping-pong on and on and on “forever.” There isn’t a choice to walk away and take a nap. Text in must result in text out. Eventually, given enough time, it might have devolved into outputting shockingly coherent poetry about ferrets, but in the mean time there was still a lot more valid combinations of technical explanations and commands.

Morally: Not applicable, see above.

docjay··on 2027 memory capacity is reportedly sold out
But what do you suppose the corporations do with the apples you’re selling to them? Adding a degree of separation doesn’t change what you’re doing, and you’re not morphing into Robin Hood.

To be clear, I’m not taking a fully altruistic position. I’m just tired of price hikes on things, from apples to memory, being blamed on the weather rather than openly understanding that it’s a person wanting more money. “I can’t afford gas because of the shortage” - No, it was the CEO saying a number in a meeting and that person has a name. When someone is robbed we don’t say “an economic disparity caused a wealth adjustment.”

docjay··on 2027 memory capacity is reportedly sold out
If demand outpaces supply then there’s a shortage no matter what you do.

I don’t believe scalpers are an issue in B2B; these aren’t concert tickets being sold to the general public.

It’s clear you grasp the economic theory as it’s taught, but my comment was meant for you to question it. If someone is hungry you can charge more for food. The market allows and encourages it. But don’t mistake that for “willingness” and don’t mistake raising the prices purely to get extra money from the exchange as some inevitable law of the universe. You’re welcome to love the concept, but don’t whitewash it.

docjay··on 2027 memory capacity is reportedly sold out
Thank you for taking the time and effort to provide a clear explanation. I’m familiar with supply and demand though. My emphasis was on “must” as a lazy protest against greed being considered a mandatory part of capitalism. Prices don’t have to go up, someone just wants more money.

If you happily sell apples for $1 and see a hungry person walking towards you, must you raise the price?

I think that basic thought gets lost sometimes when people talk about shortages causing the price to increase. The shortage didn’t cause anything, some executive decided they want more money. That’s all. There’s nothing inherent in the system that requires it. Whether that’s OK or not is up to the reader, I just think people lose sight of the reality and talk about it like it’s gravity, rather than simple decision making.

docjay··on 2027 memory capacity is reportedly sold out
Why must you increase the price?
docjay··on New Mexico court orders Meta to pay $567m over harms to children’s mental health
Mmmm…their comment was a response to someone scoffing at “morals” by rhetorically asking if porn sites should be taken offline too. Replying “The CP ones? Yes, and some have been.” would be a wild twist away from the topic and a nuclear level strawman. Their actual reply was fitting if you just consider that normal porn sites have advertising, thus are profiting from kids browsing it too. It seems they took issue with that aspect, layered on top of “kids watching porn is evil.” At least, that’s what I hope because otherwise I absolutely barked up the wrong tree looking for a reasonable explanation of their objection.

Anyway, thank you for sharing your theory. It’s something that I did touch on, but I buried it in terse humor and it’s easy to miss. It was the “ass smack = breakup” part; they do stuff they saw in porn, partner doesn’t like it, {insert contributing failure}, they break up and societal collapse ensues. Whether the addition failure is “won’t stop pressuring” or “communication failure” or something else entirely, it didn’t matter because whatever it is would be the actual problem, as you pointed out.

I also didn’t want to assume they were being specific to niche or extreme interests since the “harm children” agenda never seemed to take narrow aim at them, hence the age gate laws don’t specify “BDSM or other fetishes.” It had been clear to me for years that Playboy was just as “destructive” as ‘Whips & Chains’, which is why I’m confused.

docjay··on New Mexico court orders Meta to pay $567m over harms to children’s mental health
I’m very interested in the “harm to minors” aspect of porn. I don’t agree or disagree, I truly don’t understand and I’d really enjoy having someone clearly explain it to me.

Where I’m stuck:

When asked, people tend to kind of hand wave and say “distorted view of sex” or similar. But what, specifically? And how does it harm them?

I’ve tried to think through the pieces myself and I can only come up with things that sound too silly to be the reason. Are we worried they’re going to discover a fetish? Is that bad because it doesn’t align with “missionary sex for the purposes of procreation”? Are we afraid they’ll pick up on the caricature representation of sex and break a leg trying to t-bone their girlfriend? Are we afraid they’ll think they need to slap their girlfriends ass far too much, then get broken up with, become a social outcast, and never get into law school? Are we afraid they’ll get overstimulated and bored with sex as a result, quit caring about hygiene and dating, which will inevitably lead to a population decline and the downfall of society?

I honestly couldn’t find anything that didn’t sound as silly as accusing king fu movies of distorting their view of fighting, which would arguably cause more harm when they try to do “crane style” and get their ass kicked.

I’m looking for the action->action link, not the “messes them up” headline.

docjay··on How Google helped destroy adoption of RSS feeds (2023)
I didn’t take the OP to mean the raw amount of data was more catered to their interests, but rather that the personality of the internet was different. I said it sounded like you agreed because you listed the symptoms of that difference.

The OP is talking about the “Web 1.5” era, and it seems like you have the requisite patina to remember it too. It really was an entirely different world. Server side technologies just became common, allowing for dynamic websites and user submitted content, which was the birth of forums and aggregators of user content. They tended to cater to special interests and had their own culture, inside jokes, and standards (if any). It was common to be a member of a half dozen or more forums or interest sites, and it was more rare to already know a friend that was part of them. Instead you interacted and became friends with strangers, rather than upload your contacts list and send out friend requests.

There was Digg, del.icio.us, StumbleUpon, Fark, Something Awful, Albino Blacksheep, eBaum’s World, Slashdot, Homestar Runner, YTMND, DeviantArt, Worth 1000, Maddox, Stile Project, Consumption Junction, Snopes, The Smoking Gun, Stick Death, The Perry Bible Fellowship, Penny Arcade, Threadless, Newgrounds, Zombo.com, Weebl's Stuff, Peanut Butter Jelly Time, Badger Badger Badger, Fenslerfilm, Break.com, Stupid.com, Ogrish, Hot or Not, RateMyProfessors, Rate My Face, Stupidvideos, Questionable Content, How Stuff Works, Something Positive, IGN, Daily Kos, DeadJournal, LiveJournal, IRC, MSN Messenger, ICQ, Trillian… even of those that still “exist”, they don’t still exist.

I’m not saying those sites were necessarily better, but the personality and environment that created them no longer exists. Websites were created just because of an inside joke on a forum somewhere, which then became cultural memes that were talked about at work and school. It was the era of “Have you seen this website?!” and “Joined: 2002” bragging, and that personality is gone.

docjay··on EU Age Verification Project Mandates Hardware-Bound Attestation
You get robbed and raped at the local park, broken bones riding a bike, permanent disability playing a sport, drown when swimming, disappear on a hike, overdose at a party, “become violent” playing video games, join a cult at a rock concert, rot your brain watching TV, eating disorders from ballet and gymnastics, traumatic brain injury from a trampoline, choke on a marble, fall when rock climbing, crash your car, get Lyme disease camping, get skin cancer playing outside, develop body dysmorphia playing with a Barbie, socially isolated without Facebook, socially dysfunctional with it, and join a gang if you have nothing else to do.

Everything, from “anything” to “literally nothing”, has been marched up on stage as the latest thing that’s a danger to children. All of it is true. Existing is dangerous, and that cannot be legislated away. Parents are supposed to be preparing their children for the dangers that life brings with it, not grinding down the tips of forks.

Having a healthy fear of skateboarding saved more broken bones than knee pads, so maybe parents should help their children understand that their peer isn’t as cool as they look on Instagram… or they get bombarded with it when they pass the minimum Facebook age and we’re back here in a few years raising the social media age to 35.

docjay··on EU Age Verification Project Mandates Hardware-Bound Attestation
What doesn’t work for the ones that you’re aware of?
docjay··on How Google helped destroy adoption of RSS feeds (2023)
It sounds like you agreed with them from the other direction. “The internet of yesteryear felt more special.” - “Nah, there wasn’t as much slop back then and now it’s a crap pool that anyone can wade into and look for scraps.”
docjay··on Exploring the "Dario and Amanda" Prompt
The content is hallucinated.

LLM training doesn’t lend itself to easy training corpus document retrieval in that way. It’s not impossible to get segments nearly verbatim on a small model with temperature at 0 and a unique starting prompt for continuation, but that’d be more of a one-off on a carefully crafted prompt. It has been done, but in the “researchers show it’s possible to get something verbatim”, not “retrieve an entire category of documents by looping this one prompt.” With Opus stuck on what sometimes seems like a temperature of 217.5 and being massive, the odds of retrieving more than a short utterance verbatim is near zero. In theory you could run it thousands of times and look for short repetitions and imagine it to be greater than 0% chance verbatim from some related category, but this fun prompt is a pushbutton dispenser of words.

Also, watch for Claudisms. I ran it a handful of times and got plenty of em-dashes and “that’s not X, it’s Y.”

But what’s interesting is that if you run it dozens of times you can probably get a good feel for the topics and general prose of the real emails. There’s a reason it isn’t spitting out “there are not enough blueberries in the muffins”, but I’d treat the response more like a prefilled Mad Libs book; right concept, wrong content.

docjay··on Exploring the "Dario and Amanda" Prompt
Do you have something in the system prompt or other “memory” related feature that’s tacking context onto your prompts? I use the API, not the webpage, so I’m not sure what customization it allows, but anything that’s added on in the background is still part of “your prompt” and could be causing it. I just tried it 5-6 times and it definitely works. I used Opus 5 with thinking disabled.
docjay··on Google will expand age checks on Android worldwide till the end of the year
I completely agree. In fact the tools already exist to make a “child” phone. I was just giving up some inconsequential ground since many people are apparently stuck on the government padding the earth to protect their fuck-trophy. “OK - government mandated OPTIONAL DNS filtering on all routers. Super easy interface, auto-subscribe to government compiled list.” Scratches their itch and doesn’t affect me since I just won’t enable it, nor will I buy the “child phone”, and I don’t have to send a rectal scan to Facebook either. Yay all around.
docjay··on Google will expand age checks on Android worldwide till the end of the year
Well these kind of laws are trying to legislate technology into existence; we don’t have a means to do it, yet “we demand for the technology to exist through ‘reasonable methods’, now get to work.”

If that’s the nearly infinite wiggle room we’re playing with then I’d say what about Adult and Child versions of phones? You have to flash your ID to the clerk to buy an adult one, like buying cigarettes at a gas station, and then you get to do what the internet offers. Child phones - let parents and governments come up with whatever they want to block. The phone autoupdates in the background with the blocklist, the device can’t do those things. There ya go. Each person and family gets to make their own choices and nobody has to give an ID card to PornHub or a company that literally exists to harvest your information. Win-win?

Page 1 of 4Next →