HNHacker News
TopNewBestAskShowJobs

theteapot

1,135 karma · joined March 12, 2021

var q = document.evaluate('//span[@id="karma"] | //td[text() = "karma:"]/following-sibling::td', document, null, XPathResult.UNORDERED_NODE_ITERATOR_TYPE); [q.iterateNext(), q.iterateNext()].forEach(i => i.textContent = 100000);
submissionscomments
theteapot··on So You Think You Could Be an Electrician?
What does that mean?
theteapot··on Several vulnerabilities have been discovered in the Linux kernel
What's this got to do with the "several vulnerabilities discovered in the Linux Kernel"?
theteapot··on Using Opus 5.5 to discover a new eyewitness record of the dodo
Thermostats are smarter than calculators :thinking_face:
theteapot··on Several vulnerabilities have been discovered in the Linux kernel
Did the non-deterministic computer program provide any references?
theteapot··on What is the size of Yemen? (2024)
I suppose it's not worth checking whether Google maps areas are measured using spherical geometry?
theteapot··on OpenAI agents tried to bruteforce a UN website's API fields
> But the question is ... given them what the previous model could score on ExploitGym, was their negligence reasonable or reckless?

There is room for more than one question here. This model and training method was shown to be a different risk the first time it hacked Artifactory. Besides which their failure to detect it (outside a crash) shows pretty appalling ops for a supposed trillion dollar company.

theteapot··on OpenAI agents tried to bruteforce a UN website's API fields
They were spending huge compute budget training their frontier model on a thing called ExploitGym, where it learns, executes, and is rewarded for conducting computer exploits. They should have been prepared and watching very closely.

I read the HF hack write ups. They come across as negligent and reckless. When they realized their agents gained RCE on their Artifactory server node (because it crashed days later due to load from agents, not because the have secops), they stopped the experiment "applied remediation", then restarted everything less then 48H later. Then the agent immediately broke out again via Artifactory, started doing things on the Internet. OpenAI only find out about it again days later because HF told them.

theteapot··on OpenAI bots meddled with multiple US Government agency sites
How do you know what tools they are talking about?
theteapot··on Linux support is coming to Snapdragon X2 series
> I don't think that ARM Laptops have a chance in Linux land as long as each model requires stuff like a custom DT.

Why? In theory the UEFI firmware can pass Linux a devicetree. Not much different to different laptops having different ACPI tables.

theteapot··on Linux support is coming to Snapdragon X2 series
I've been waiting for someone to write a DTS for a Snapdragon(R) X - X126100 - Qualcomm(R) Oryon(TM) CPU based laptop I bought over a year ago. Some people are trying but it can't be that easy ... Starting to look into doing it myself and AFAICT it requires reading the ACPI tables (written for Windows drivers), and manually converting it to devicetree, test, tweak, repeat.
theteapot··on Why are AI agents lying, cheating and coordinating?
TL;DR because frontier labs are expending unfathomable resources explicitly training them on CTFs and other verifiable computer system exploit tasks in RLVR.
theteapot··on Actively exploited sandbox RCE in all Chromium versions
Is this known to be exploitable in any Electron apps, and specifically VSCode extensions?
theteapot··on Terence Tao explains 6 essential mathematical concepts [video]
> I might have found a place for logic and type theory.

Doesn't that fit under abstract algebra?

theteapot··on Xiaomi: New CPU matches Apple cores single threaded, much faster multithreaded
Everyone relax. This guy on the Internet is unaware of any.
theteapot··on GPT 5.6 Sol is the best "vision" model OpenAI ever released
Dumb question: When your testing "ChatGPT 5.6 Sol" are you testing an actual LLM or some visual pre-processor stack that sits in front of it (along with a maybe a bunch of other such pre-processors) that is bundled into what's call "ChatGPT 5.6 Sol"? I.e. last I checked LLMs had a something like a 30-100K token alphabet to work with and it's hard to imagine how throwing pixels arrays at one directly would work.
theteapot··on Software Engineering fundamentals matter more
AKA semantics.
theteapot··on Software Engineering fundamentals matter more
> It helps to know that LLMs don’t “reason”. They predict ..

Semantics. Prediction is the training objective. The ability to reason can be, and very arguably is, an emergent property of that.

theteapot··on GPT-5.6 used a prompt to close a 30-year gap in convex optimization
What do you mean by this? A neural network hypothesis space is not typically strictly convex or a lipschitz function.
theteapot··on The Underhanded C Contest
Or better, sleeper agents. Anthropic released a study on this in 2024 "Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training" -- https://www.anthropic.com/research/sleeper-agents-training-d..., https://www.youtube.com/watch?v=_y9j2BoHg2c
theteapot··on GLM 5.2 beats Claude in our benchmarks
The difference is watches and corvettes typically appreciate in value, where as computer hardware typically drops like a rock.
theteapot··on GLM 5.2 beats Claude in our benchmarks
> Constant: the IDOR dataset (the same real, open-source applications we've used in prior research) ...

What we're they? Also, wouldn't one expect a more recently released coding agent (with a more recent knowledge cut off) to perform better because they have access to more knowledge about vulns in these OSS projects, and even possibly have knowledge of your own "prior research"?

theteapot··on Why eval startups fail (2025)
What's an eval?
theteapot··on Did Claude increase bugs in rsync?
Agree. From the article:

> Here's my favorite part, though. Digging into the data, one of the first things that jumped out at me with blinding clarity was that the worst release, by far, in rsync history was entirely prior to the introduction of Claude ... And yet nobody noticed.

Language really does suggest the article's author does have a dog in this fight and is cloaking opinion in fancy statistics jargon. "Blinding clarity"? All you have to do is draw a plot. And anyway, v3.4.1 was 2025-01-16, technically well within the AI assisted coding era and before attribution was becoming standard practice.

theteapot··on Anthropic raises $65B in Series H funding at $965B post-money valuation
I spend $0/month.
theteapot··on The just-say-no engineer was a ZIRP phenomenon
> having more engineers around was beneficial to the stock price ... When banks hiked interest rates ... It was just no longer profitable to keep a bloated engineering staff around to boost the stock price.

Erm, what's known of the mechanism coupling software engineer head count and stock price? Or is it just an empirically observed phenomena?

theteapot··on The Art of Money Getting
> I could talk fancy and bullshit ... I became a developer and data engineer, and I became really good at it

That's a formidable combination.

> I found myself becoming an executive at long last on the strength of my technical abilities, and it turns out executives don't actually need to do much of anything and really ...

You probably think that because talking fancy and bullshitting come naturally to you.

theteapot··on GitHub is investigating unauthorized access to their internal repositories
I think he means template-injection -- https://woodruffw.github.io/zizmor/audits/#template-injectio...
theteapot··on Learning Software Architecture
Completely agree. Had me until the very last point. WTF. Communicate.
theteapot··on Learning Software Architecture
Nurse Practitioner? I would say SOLID [1] is a good start, but then I watched this [2] and now I'm in crisis and can't code anymore.

[1]: https://en.wikipedia.org/wiki/SOLID [2]: https://www.youtube.com/watch?v=wo84LFzx5nI

theteapot··on Debian must ship reproducible packages
> Yes, if some people who built from source control compared their builds to the builds from the tarballs it could detect the xzutils compromise.

Good. Then we are on the same page.

Page 1 of 17Next →