HNHacker News
TopNewBestAskShowJobs

svara

3,637 karma · joined March 25, 2016

Building microscopy analysis software at ariadne.ai (co-founder and CEO).

My username at ariadne.ai if you'd like to chat.

submissionscomments
svara··on Show HN: I made Google Trends for Hacker News by indexing 18 years of comments
When searching Fastly it seems to match "fast".
svara··on Only 16 Percent of Americans Think AI Will Have a Positive Impact on Society
I think one positive thing that might come of this is for AI to act as a sort of counterweight to the fragmentation of reality into different filter bubbles.

It might be difficult to make models that have useful, high intelligence, but also are very biased. It could create a sort of grounding in logic and reality.

Grok might actually be early evidence of this. Despite the bad press it gets, it's really not so bad.

One can always hope ...

svara··on Only 16 Percent of Americans Think AI Will Have a Positive Impact on Society
I'm legit pretty excited about applying AI to accelerate biological and medical discovery.

It's already happening right now, still in relatively mundane ways, but there's so much to do.

svara··on Amazon CEO's Talks with U.S. Officials Triggered Crackdown on Anthropic Models
Okay but if I understand correctly what you did, you measured the performance with automatically rewritten prompts on Fable vs. original on Opus? This might be where the difference in performance that you saw came from.
svara··on The computer science degree isn’t dead
While I can only peripherally relate to the specifics of your story, I think it beautifully illustrates how interesting and mind expanding it is to spend time in different cultural contexts, and that different cultures can very much co-exist in the same countries or even in the same people.

Everyone should do it more, it really helps put the uncompromising convictions of people around you into perspective and see them as what they often are: a lack of understanding for the breadth of human experience.

svara··on "Don't You Just Upload It to ChatGPT?"
The things is, this is almost certainly what's happening.

You can (could, maybe they 'fixed' it by now) get sota LLMs to reproduce entire novels near verbatim.

The idea of giving it parallel texts of those novels in different languages, to train it on translation, is so obvious it'd just be strange if the AI labs didn't do it.

In fact DeepL was doing basically that more than 10 y ago.

svara··on Claude Fable 5
Unfortunately useless if you do anything related to biology. It doesn't try to flag dangerous queries, it just flags queries as biology-related wholesale.

It's absurd. To see how far the filter goes I asked it "Are trees a monophyletic group?" and that does trigger the filter.

svara··on I built a vulnerable app and spent $1,500 seeing if LLMs could hack it
This is strange to me, did you really ask like this and which model did you use?

I just tried your no. 1 and 3 verbatim and Opus gave fine answers; no. 6 I've done in the past with no issues. The other ones we can't really replicate without more details, but based on my experience with Opus I don't see what the issue would be.

The reason I'm really surprised by this is I do a lot of biology prompts and the guardrails used to be quite problematic up until some time late last year. Many legitimate prompts would trigger its biosafety filters.

But I haven't seen such filters trigger at all anymore in more than half a year.

svara··on Domain expertise has always been the real moat
Chess and proofs only work as comparisons to the extent that you can find parts of your job that share their key property: A solution is sought to a problem that can be stated with relatively little information.

What prompt would someone have used to get a superhuman coding agent to output the Linux kernel or GTA5?

Before you accuse me of moving the goalposts, that's not my point: The examples are there to help think about what humans would still need to do to build complex projects even if the coding itself was perfectly reliable.

Both the Linux kernel and GTA5 contain a large amount of incompressible information; humans thought long and hard about how to design them, i.e. about what that thing they were building was even supposed to be.

svara··on Domain expertise has always been the real moat
I had the same thought recently, I've had it happen to myself.

I've been working on something relatively large and greenfield recently.

A big chunk of my time is spent thinking about the hard parts. The raw information processing rate needed to keep up with the state of the project is high.

It feels almost like mental athleticism, whereas coding used to be a rather chill activity.

svara··on SQLite is all you need for durable workflows
Word on HN is that you're either paying more money than you expected for temporal's managed solution or taking on substantial ops burden ultimately running their very heavy system yourself.

I wouldn't know, I've not done either, but I'd like to learn more from your or other's experience.

svara··on I think Anthropic and OpenAI have found product-market fit
I agree with most of what you're saying, but I think the point I was trying to make wasn't as high-flying as you and others understood it.

I'd pay a premium for even just a model that's 20% better, no ASI required, and I think a lot of people would. I wouldn't call that marginal, if it means I'm getting frustrated on 20% fewer tasks.

A recurring pattern that I've seen in myself and others is to at first be very impressed by a new model's coding capabilities, and then desensitize quickly and start being frustrated by the shortcomings.

svara··on I think Anthropic and OpenAI have found product-market fit
There's still a lot of room for the best models to get better at coding .

Your argument rests on the "for marginal gains" part but it's really not clear that the gains are marginal in the foreseeable future.

svara··on Dropbox CEO Drew Houston to step down
The desktop client used to be just terrible. Has that changed? The Dropbox client does have its issues but it's really amazing at... Syncing files. I use it pretty creatively with large numbers of files and large volumes and it just works reliably.
svara··on Google changes its search box
So they finally have become AskJeeves?

On a more serious note, the on demand UI chrome could actually be cool UX, curious to try that out.

I see no change to look and feel so far, has this rolled out to anyone yet?

svara··on OpenAI Adopts Google's SynthID Watermark for AI Images with Verification Tool
I think something like this will need to become ubiquitous, widely supported in software and understood by lay people: https://en.wikipedia.org/wiki/Content_Credentials
svara··on It is time to give up the dualism introduced by the debate on consciousness
I don't think the question of whether subjective phenomena are casual has any bearing on whether there's a 'hard problem' and I also don't think the term 'hard problem' is used like that by others.

Whether subjective experience is casual or not is a different, additional question.

svara··on It is time to give up the dualism introduced by the debate on consciousness
There's nothing wrong with that chain. This is what some philosophers would call the 'easy ' problem of consciousness, to distinguish it from the 'hard' problem, which is the next step:

How do you get from a physical model of brain physiology and behavior to subjective experience of mental states?

svara··on It is time to give up the dualism introduced by the debate on consciousness
> My experience is that I'm conscious, and math cannot result in consciousness, therefore consciousness is a separate thing." Question: who says math cannot result in consciousness? Do you have empirical proof of that?

A lot of people, myself included, have the intuition that thinking that this might be possible is a sort of type error, to put it in CS terms.

A bit like asking "Have you proven that ice cream? Are you sure maths can not prove that ice cream? Do you have empirical evidence?"

Asking for empirical evidence seems beside the point, since the issue is a logical one.

svara··on How Claude Code works in large codebases
I use Claude Code quite a bit and quite enjoy it, so I'm a bit confused by how often it's mentioned that you should have CLAUDE.md.

I mean: If there was something you could add to the prompt to consistently increase performance why isn't it in the system prompt already?

If it's all about clarifying a couple of local idiosyncrasies, shouldn't it be able to quickly get them by looking through the repo?

Does anyone have an example of a CLAUDE.md that really makes a difference for them?

In general, this article would really have profited massively from examples of good applications of those patterns.

svara··on Princeton mandates proctoring for in-person exams, upending 133 year precedent
> but when 'anyone can be anything' it creates hyper competition, anxiety

Not sure if you intended this but this is basically exactly Byung-Chul Han's point in The Burnout Society.

svara··on A HN post with negative points – how?
That's probably the submitter's friends voting.
svara··on Cloudflare to cut about 20% of its workforce
Isn't the most likely explanation here that they needed to show in their earnings call how their bet on becoming AI infrastructure is leading to high revenue growth expectations, and that isn't happening (yet)?

The stock is currently at -17% in after hours trading.

So you need to do something that's good for your margins to show investors.

svara··on Norway Set to Become Latest Country to Ban Social Media for Under 16s
I mostly agree with you, I think what you're implying is correct on average, but I'm probably not the only one to whom HN is more addictive than Instagram, Tiktok and all the other classic social media apps.

They get boring much more quickly and also make me feel guilty about spending time on something so shallow, so it's very self limiting.

svara··on GPT-5.5
Do we know if this is another post training fine tune or based on a much larger new pretraining run (which I believe they were calling 'Spud' internally)?

The large price bump might indicate the latter.

svara··on Renewables reached nearly 50% of global electricity capacity last year
This is correct in the sense that, if you were to build a zero emissions energy system from scratch with today's technology, your conclusion would be that you'd eventually have to do this.

But in much of the world, setting up PV is economically sound simply because it displaces a certain amount of kWh generated over the course of a year from other sources that are more polluting and more expensive.

In this regime, the dynamics of production over time don't matter yet.

At some point, when renewable generation has very high penetration, you'll reach a point where building more is uneconomical, and to then displace the remaining other power sources you'll need to overpay (ignoring externalities).

However, that's assuming no technological change on the way there, which is a whole separate topic.

svara··on AI overly affirms users asking for personal advice
The issue is it will follow your instructions. It's sycophancy one step removed.
svara··on AI overly affirms users asking for personal advice
Yeah, and if you ask it to be critical specifically to get a different perspective or just to avoid this bias, it'll go over the top in the opposite direction.

This is imo currently the top chatbot failure mode. The insidious thing is that it often feels good to read these things. Factual accuracy by contrast has gotten very good.

I think there's a deeper philosophical dimension to this though, in that it relates to alignment.

There are situations where in the grand scheme of things the right thing to do would be for the chatbot to push back hard, be harsh and dismissive. But is it the really aligned with the human then? Which human?

svara··on Epoch confirms GPT5.4 Pro solved a frontier math open problem
I think you're misreading me. My point isn't that you can't in principle state the optimization problem, but that it's much easier in some domains than in others, that this tracks with how AI has been progressing, and that progress in one area doesn't automatically mean progress in another, because current AI cost functions are less general than the cost functions that humans are working with in the world.
svara··on Epoch confirms GPT5.4 Pro solved a frontier math open problem
The capabilities of AI are determined by the cost function it's trained on.

That's a self-evident thing to say, but it's worth repeating, because there's this odd implicit notion sometimes that you train on some cost function, and then, poof, "intelligence", as if that was a mysterious other thing. Really, intelligence is minimizing a complex cost function. The leadership of the big AI companies sometimes imply something else when they talk of "generalization". But there is no mechanism to generate a model with capabilities beyond what is useful to minimize a specific cost function.

You can view the progress of AI as progress in coming up with smarter cost functions: Cleaner, larger datasets, pretraining, RLHF, RLVR.

Notably, exciting early progress in AI came in places where simple cost functions generate rich behavior (Chess, Go).

The recent impressive advances in AI are similar. Mathematics and coding are extremely structured, and properties of a coding or maths result can be verified using automatic techniques. You can set up a RLVR "game" for maths and coding. It thus seems very likely to me that this is where the big advances are going to come from in the short term.

However, it does not follow that maths ability on par with expert mathematicians will lead to superiority over human cognitive ability broadly. A lot of what humans do has social rewards which are not verifiable, or includes genuine Knightian uncertainty where a reward function can not be built without actually operating independently in the world.

To be clear, none of the above is supposed to talk down past or future progress in AI; I'm just trying to be more nuanced about where I believe progress can be fast and where it's bound to be slower.

← PreviousPage 2 of 24Next →