HNHacker News
TopNewBestAskShowJobs

anuramat

156 karma · joined July 17, 2023

submissionscomments
anuramat··on There Is No AI (It's Just People) with Jaron Lanier
> You are conflating

"It doesn’t even exist until you ask it to do something", not my words

anuramat··on Due to concerns about malicious applications, GPT2 will not be released (2019)
why?
anuramat··on There Is No AI (It's Just People) with Jaron Lanier
how are humans autonomous by this definition then? did you ever do anything before you were born?
anuramat··on Using an open model feels surprisingly good
do you think typing `make` into a terminal is some sort of an insane intellectual achievement?
anuramat··on Kimi K3-256k
why can't they just increase the price for cached input instead?
anuramat··on Claude Opus 5
could it be that the leading AI lab is decent at making LLMs?
anuramat··on Using an open model feels surprisingly good
what's there to hack that you couldn't do with $3 in openrouter credits?
anuramat··on Banning AI will not make it go away
> productivity findings

are you referring to the early 2025 METR study?

anuramat··on Using an open model feels surprisingly good
> you can run a small model

but why?

anuramat··on Claude Opus 5
I think you mean mostly-stateless-but-with-prompt-caching-and-batching, ie not really
anuramat··on Claude Opus 5
you think they're doing inference at a loss even with the API prices?
anuramat··on Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard
yes; fyi usage limits on the $200 claude sub correspond to at least $1.2k/week in api tokens
anuramat··on Are AI Labs Pelicanmaxxing?
"benchmaxxing by generalizing" is not really benchmaxxing
anuramat··on Who's afraid of Chinese models?
> Russia

lmao

anuramat··on Coding agents think ahead of time
> possibility of "LLMs can reason not like a human"

what would be the difference between not reasoning and reasoning not like a human?

anuramat··on Coding agents think ahead of time
I keep asking the same question, and I think the steelman version would be "has metacognitive patterns similar to humans"
anuramat··on I love LLMs, I hate hype
I'm somewhat serious -- if you think AI will scale that well, you can't really make predictions like that

I personally don't think the weight efficiency will improve that much; if anything big does happen, I expect it to be about scalable architectures and continual learning

anuramat··on GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps
just found a decent looking benchmark for iterative development: https://swe-milestone.com/

surprised it isn't a bigger thing, eg artificial analysis doesn't report anything like that

still doesn't measure the human-agent interaction part, but that's pure vibes atp

anuramat··on I love LLMs, I hate hype
> nowhere near readteaming

I think any security-related task triggers it to think about the threat model and thus hit the guardrails

anuramat··on I love LLMs, I hate hype
what languages are you working with? I imagine if it's something like C, you'd hit guardrails every time you manage memory
anuramat··on The git history command
is there a way to use it with lazygit? unfortunately I'm addicted atp, and "autorebase branches" is exactly what I was missing the entire time
anuramat··on I love LLMs, I hate hype
why would you need a local fable at that point? AGI will surely solve all the problems in the world at that point
anuramat··on I love LLMs, I hate hype
what are you working on? I only hit the guardrails twice after burning through two weeks of 20x max plan, both times on ML stuff; still more than I'd want to, but not unusable
anuramat··on Fable extended until 19 July
ml research, both brainstorming and running experiments

eg I'd throw a hypothesis at it in the evening, and overnight it would write the code, do a sanity check, start a run, monitor the metrics, identify and fix a bug, propose a new hypothesis based on the results, write the code and start the second experiment, etc

still not that good at generating ideas or drawing conclusions, but much better at criticism than opus

even the api price is not that expensive for this

anuramat··on OpenBSD has a use-after-free allowing local privilege escalation to root
security researchers are using the free tokens that they get to do useful security work, AI labs are giving away free tokens to maximize their profits; is it really that hard to imagine that different parties might have different goals?

> We shouldn't look at this and think

you're gonna tell me what to think now?

anuramat··on Entire: Distributed Git for Agents
so, git?
anuramat··on OpenBSD has a use-after-free allowing local privilege escalation to root
> it's not actually a financially efficient way ... unless you profit from creating demand for LLMs

well, they do? it's a win-win, you can't really criticise an AI lab for doing AI instead of straight up giving money to security researchers

> If you factor in all the money spent on training

why would I? it's not a cybersec-specific model

anuramat··on GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps
I imagine one could one-shot a basic app and then feed feature requests one by one, sounds like an obvious way to benchmark architecture/maintainability
anuramat··on AI 2040: Plan A
> data centers being built in their communities

> golf courses are a traditional green space where people in a community

I have a feeling those two sets of communities are disjoint

anuramat··on AI builders outnumber AI governance hires 7:1 in Europe
if you don't make your own tech, standards will be defined by people that do
Page 1 of 8Next →