HNHacker News
TopNewBestAskShowJobs

sbszllr

687 karma · joined August 7, 2018

trustworthy ml ¯\_(ツ)_/¯
submissionscomments
sbszllr··on Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing
I commented it last time the post about Claude watermarking went viral and I'm going to say the same thing again:

"I've been working in the media and model IP space for quite many years. What this article misses a bit is the threat model for watermarking in general. Watermarking and fingerprinting have inherently weak security guarantees -- they rely a lot on security through obscurity, weak assumed adversaries to deliver. There are clear trade offs between true positives, false positives and maintaining the quality of the media. It's true for audio-visual media, models and their outputs alike. As much as I like to take shots at poor technical choices by corps and govs, this one is unjustified. Sure, inserting glyphs is bad but biased sampling is as good as it gets in 2026."

Since the announcement, there have been many people who don't seem to fully understand what a security guarantee is, what trade offs it might involve, or how popular the type of technological solution is in general (media watermarking is ubiquitous).

And naturally, there are challenges with how you will make sense of the score in your org, e.g. you wrote an email, and it's flagged as LLM-generated because you copied two generated/edited paragraphs.

Yeah, the article might disagree with watermarking as a matter of principle, comparing it to censorship. But the methodological arguments that I have read so far have been thin in these articles.

sbszllr··on Text AI watermarks will always be trivial to remove
I've been working in the media and model IP space for quite many years. What this article misses a bit is the threat model for watermarking in general.

Watermarking and fingerprinting have inherently weak security guarantees -- they rely a lot on security through obscurity, weak assumed adversaries to deliver.

There are clear trade offs between true positives, false positives and maintaining the quality of the media. It's true for audio-visual media, models and their outputs alike.

As much as I like to take shots at poor technical choices by corps and govs, this one is unjustified. Sure, inserting glyphs is bad but biased sampling is as good as it gets in 2026.

EDIT: grammar

sbszllr··on Claude Design
It's possible and even likely there's industrial espionage going on. But imo, you don't need that. I've worked in cutting edge industries, and even when you don't know what your competition is doing, there are usually only so many logical next steps.
sbszllr··on Claude Design
It's interesting how OpenAI and Anthropic effectively mass dumped a bunch of similar features in the last two days.

I wonder what other features they're cooking right now.

sbszllr··on Reinventing the Pull Request
Maybe it just shows my lack of tolerance for process/overhead.

As a fellow rebase enjoyer, I will do it occasionally for smaller PRs but to me, it becomes unwieldy for large ones.

Do you have any tips or aliases that makes it more workable?

sbszllr··on Reinventing the Pull Request
Let's forget that this post is an ad. I feel like there is a use for LLMs that could help us do stacked PRs better.

Right now there are effectively three ways to do a PR:

- a bunch of small commits, some of them related to the feature, some fixes, some mixing both -> a PR with 'n' commits -> they don't really make sense as atomic commits, you have to review the entire PR to make the sense of it

- a squashed PR

- some uber principled reorganisation of commits that separates key implementation concerns into smaller commits (effectively stacked PRs but clean)

The last option would be desirable but it's unreasonable to expect anyone to do it by hand. So this is where <maybe> an LLM could parse my garbage intermediate commits, the final diff and generate a stack instead?

sbszllr··on Apple Studio Display and Studio Display XDR
Daisy chaining finally supported.
sbszllr··on Parse, Don't Validate and Type-Driven Design in Rust
I was thinking a similar thing when reading the article. Often, the validity of the input depends on the interaction between some of them.

Sure, we can follow the advice of creating types that represent only valid states but then we end up with `fn(a: A, b: B, c: C) transformed into `fn(abc: ValidABC)`

sbszllr··on The EU moves to kill infinite scrolling
> We already had that disaster where pop-ups fly out "do you want to accept those cookies". That is just a usability nightmare. People are forced into extensions, just to stop wasting their time here.

You’re perpetuating a gross misunderstanding of the cookie law. What it states is different from how the advertisers implement malicious compliance to bias people, like yourself.

Websites that implement basic functional cookies do not need to display any popups. They’re permitted to do so. Any cookies that are essential to the functioning of the website within reason are permitted. In fact at no point a website should serve you a cookie popup unless you seek it out because analytics and advertising cookies are supposed to be opt in.

So many websites do two things, serve you a popup that has everything enabled which is a clear violation; or a popup that has only functional cookies selected but the biggest highlighted button allows all of them.

The law is fine. Malicious compliance is to blame. The EU has been slow to rectify it.

sbszllr··on Prism
The quality and usefulness of it aside, the primary question is: are they still collecting chats for training data? If so, it limits how comfortable, and sometimes even permitted, people would with working on their yet-to-be-public work using this tool.
sbszllr··on Signal leaders warn agentic AI is an insecure, unreliable surveillance risk
I agree that is performant enough for many applications, I work in the field. But it isn't performant enough to run large scale LLM inference with reasonable latency. Especially not when we compare the throughput numbers for a single-tenant inference inside a TEE vs batched non-private inference.
sbszllr··on Signal leaders warn agentic AI is an insecure, unreliable surveillance risk
Interestingly enough, it is possible to do private inference in theory, e.g. via oblivious inference protocols but prohibitively slow in practice. You can also throw a model into a trusted execution environment. But again, too slow.
sbszllr··on European Commission issues call for evidence on open source
These are fair points but the weight of their impact is a misconception. Times and times again, lower capital and investment risk aversion are shown to be the limiting factors.
sbszllr··on Rust: Proof of Concept, Not Replacement
There's a lot of truth and real pain points in the article but others miss the point entirely. Two things that stand out to me in particular:

- C "standard" is quite flaky in practice, there are lots of un/underdefined things that compilers interpret quite liberally for the purpose of optimisations

- complaining about the syntax and symbols is unfair: rust offers all these semantics to represent the memory model of your codebase. The equivalent is not possible in C/C++, and when we try to do it, we're inventing our own constructs at the code level instead of relying on the syntax

sbszllr··on A modern 35mm film scanner for home
I've been camera scanning 4x5 and I'm happy with the results. Take two offset photos and stitch them in post. Mind you, I scan with pixel shift for higher res.
sbszllr··on A modern 35mm film scanner for home
As someone who has a mirrorless scanning setup for my film, and pondered getting a dedicated scanner... the price of this is quite steep given how inflexible of a tool it is.

A second hand DSLR setup is going to be roughly the same price or less. I'm also not sure what kind of workflow improvements it actually offers. If you want fancy and experimental, filmomat has arguably a more interesting but pricier offering.

But naysaying aside, I hope they manage to find a niche that allows them to survive as a company, and keep the analog photography revival alive.

sbszllr··on Zig's New Async I/O
I don't know if it's still true in the recent versions of Scala (stopped caring in 2018) but it used to have implicit parameters designed specifically for passing context like this.

A notable example was passing around an implicit ExecutionContext for thread pools, e.g. in Akka :)

sbszllr··on EU Commission refuses to disclose authors behind its mass surveillance proposal
Anecdata but I also had good experiences reaching out to MEPs, so not all is lost.

At its core, the core issue seems to be the lack of accountability between the MEP, and people that voted them in. Few people vote in the EU elections, and even fewer follow up on what happens there.

Chicken and egg problem but if you want your MEP not to be just "a good obedient MEP they are", the electorate needs to ask more of them.

sbszllr··on EU Commission refuses to disclose authors behind its mass surveillance proposal
Usual reminder -- if you're an EU citizen, call up your representative.
sbszllr··on Claude Code: Best practices for agentic coding
All can be true depending on the business/person:

1. My company cannot justify this cost at all.

2. My company can justify this cost but I don't find it useful.

3. My company can justify this cost, and I find it useful.

4. I find it useful, and I can justify the cost for personal use.

5. I find it useful, and I cannot justify the cost for personal use.

That aside -- 200/day/dev for a "nice to have service that sometimes makes my work slightly faster" is much in the majority of the world.

sbszllr··on Claude Code: Best practices for agentic coding
The issue with many of these tips is that they require you use to claude code (or codex cli, doesn't matter) to spend way more time in it, feed it more info, generate more outputs --> pay more money to the LLM provider.

I find LLM-based tools helpful, and use them quite regularly but not 20 bucks+, let alone 100+ per month that claude code would require to be used effectively.

sbszllr··on Adobe deletes Bluesky posts after backlash
Yup, I prefer Lightroom to Capture One, especially for film-related workflows.

But I just can't go back to their predatory pricing practices, and the absolute malware of a programme that creative cloud is.

sbszllr··on Open Source Coalition Announces 'Model-Signing' to Strengthen ML Supply Chain
Source: I have a relationship with OpenSSF but not directly involved. I'm involved in a "competing" standard.

As other commenters pointed out this is "just" a signature. However, in the absence of standardised checks, this is a useful intermediate way of addressing the integrity issue around ML supply chain; FWIW today.

Eventually, you want to move to more complete solutions that have more elaborate checks, e.g. provenance of data that went into the model, attested training. C2PA is trying to cover it.

Inference time attestation (which some other commenters are pointing out) -- how can I verify that the response Y actually came from model F, on my data X, Y=F(X) -- is a strongly related but orthogonal problem.

sbszllr··on Show HN: Formal Verification for Machine Learning Models Using Lean 4
Hmmm, there're no scientific work that really let's you do these things right now. And the repo doesn't cite any work, no associated paper.
sbszllr··on I've been using Claude Code for a couple of days
> In my view, the optimal scenarios for using LLM coding assistants are:

> - Architectural discussions, effectively replacing traditional searches on Google.

> - Clearly defined, small tasks within a single file.

I think you're on point here, and it has been my experience too. Also, not limited to coding but general use of LLMs.

sbszllr··on U.K. demand for a back door to Apple data threatens Americans, lawmakers say
I agree with your point that government overreach is more serious.

Which is why I want to emphasize that various government police (like FBI) notoriously buy data that they would need a warrant for otherwise.

I’m aware that you’re saying it, but I think you’re underestimating the extent to which preventing spying from the corps == preventing spying from the govt.

sbszllr··on Cali's AG Tells AI Companies Almost Everything They're Doing Might Be Illegal
Might is irrelevant and doesn't prevent any abuse, nor does it foster innovative environment. Both the EU and the US, need to pick a side. Either they should make it illegal and establish precedent, or just let go.
sbszllr··on Modelica
As someone who has no idea what this is about bar the landing page explanation, and isn't in this space -- it would be great if the front page had examples, or links to examples.

30 seconds of clicking around and I've failed to find sth compelling.

sbszllr··on DeepMind debuts watermarks for AI-generated text
I’ve been working in the space since 2018. Watermarking and fingerprinting (of models themselves and outputs) are useful tools but they have a weak adversary model.

Yet, it doesn’t stop companies from making claims like these, and what’s worse, people buying into them.

sbszllr··on EU "Chat Control" and Mandatory Client Side Scanning
“Think of the children” is, as usual, just to get the foot in the door. They use it as a justification, because it works.

Of course CSAM is bad, shouldn’t we do everything in our power to prevent it? If you implement client-side scanning, you will catch some rookies. Some old pervs that don’t know how to use encryption manually, or use Matrix. They will use them to show how effective the system is…

with the exception that it doesn’t work against anyone who knows anything about computers. And I think the regulators know it, they aren’t dumb (imo). It’s, like I said earlier, an excuse to expand the scope of scanning later.

Page 1 of 2Next →