1,466 karma · joined January 30, 2016
It has never made sense to me that people are willing to ignore the trends and argue in platitudes.
https://en.wikipedia.org/wiki/List_of_countries_by_carbon_di...
[1] - https://en.wikipedia.org/wiki/List_of_countries_by_carbon_di...
YC has advised startups in the past that it's easier to sell a single $100k customer than 100 $1k customers.
It would also be relatively surprising to learn that i.e. the Chinese providers are OOMs better at inference than OAI/Anthropic (like their prices would imply if they were in a perfectly competitive market).
1 - https://x.com/Kimi_Moonshot/status/2035074972943831491?lang=...
I also got this,
> A useful next habit is to end each correction with a concrete acceptance test, owner artifact, or stop condition: “write it into PROGRESS.md,” “make nix run .#bench fail until this is real,” “rerun this exact command,” or “do not proceed until these two choices are explicit.”
> You already do this well in the biggest penance sessions. Apply it to the smaller ones too.
Which I have found to be counterproductive in my personal work. Current models can generally infer acceptance tests of this level of granularity (not true for larger project-level prompts, but those don't produce good enough code for me yet -- even with specific acceptance criteria).
I also got penalized for using claude in read-only mode for the same validation reason?
> For read-only work, end with one of:
“turn the top finding into a PR-sized plan”
“mark these as accepted/rejected/deferred”
“write a cleanup checklist”
“give me the exact command I should run safely”
“stop, no action recommended”
No thanks, I'm literally just exploring the codebase. I don't want any of these.It's a little sad to be honest, I would actually enjoy a product that helped me improve prompting + ai usage.
Put another way, the thing we are all concerned with is the complete circumvention of safeguards that is normally possible with llms. If you _aren't_ arguing that this isn't possible, you're not engaging in discussing the the thing that is concerning to regulators or those discussing the regulation.
That,
A. Anthropic solved the llm jailbreak problem with mythos (despite no claim to have done so on their part)
B. That a full jailbreak of mythos is possible.
My hope is that this gets us some real concern for things that have been defended with de facto arguments (i.e. privacy) going forward.
edit: Anthropic argues that your Crayola analogy is fundamentally incorrect.
> Legally, a supply chain risk designation under 10 USC 3252 can only extend to the use of Claude as part of Department of War contracts—it cannot affect how contractors use Claude to serve other customers.
https://www.anthropic.com/news/statement-comments-secretary-...
Disregarding who is right or wrong for a moment, if the DoW are right (which I'm not personally inclined to believe, but we're ignoring that for the moment) -- how else can they avoid secondhand Claude poisoning?
Supposing they really want to use their software for things disallowed by Claude's (now or future) ToS, it seems like designating it a supply chain risk is the only way they can ensure that their contractors don't include Claude (either indirectly as a wrapper or tertially through use of generated code etc)