HNHacker News
TopNewBestAskShowJobs

brap

3,404 karma · joined March 18, 2021

submissionscomments
brap··on Revealing the details of how OpenAI agents hacked Hugging Face
I think the sandboxing was truly incompetent but in their defence something like this was probably seen as very unlikely. Let’s all hope they do better in the future.
brap··on We're gonna need a lot more mathematicians
You and the other commenters are all absolutely right, but that’s not really what I was pointing at.

So far our solutions were more or less understood thru some models of reality that we’ve constructed (on our own), which may or may not reflect reality perfectly, and even if most of us never bothered thinking about these models, some people did and they understood them on a very deep level.

But we may be getting to a point where the problems we need to tackle become too difficult for humans to model, or even to notice their existence, like asking an ant how a Boeing 747 works.

Maybe this was always the case but now it seems like we might have a shot at making these solutions useful even if we have no idea what they’re even solving.

brap··on We're gonna need a lot more mathematicians
I feel like “giving up understanding” is inevitable.

There’s some hubris in thinking we can understand everything. For truly difficult problems, it’s entirely possible that humans are simply incapable of comprehending why a solution is true. But ultimately the practical value of applying that solution to the real world is going to eclipse our need to understand it.

Math is just the beginning. I see it happening in other fields too, like physics and biology. Many of us software devs have already given up on understanding parts of our own systems for the exact same reason.

Seems like a losing battle.

brap··on Jev in 25 Lines of Python
What I don’t understand is, why would you not want “reasoning” in a classifier?

Speed and cost are obvious reasons, but isn’t this a tradeoff?

brap··on GPT-6 Sol and Luna
Am I the only one who feels icky about how these 2 companies always try to one-up each other on release day? It’s fair and all but just feels gross.
brap··on ZuckOff is a free app that sees Meta glasses before they see you
Sorry to be that guy but:

>Quiet is not proof that nobody is recording, and a detection is not proof that anyone is

Is such claudespeak it’s not even annoying it’s just funny at this point

brap··on How to Write with an LLM
One thing that sort of worked for me is to write the skeleton myself, i.e. the general ideas and how they fit together, the overall “flow” of what I want to say, then have an LLM fill in the blanks (I either point it at some context.md or have it ask me questions when details are missing).

It’s kind of like designing the high level software architecture yourself and have the LLM write the code for each component.

Not bulletproof, requires some iteration, but miles better than what it would produce on its own.

brap··on How to Write with an LLM
One thing that you should absolutely never do: have an LLM review and “improve” the text repeatedly in some closed loop.

This might work well for some tasks (coding), but for writing it will absolutely take reasonable text and turn it into a pile of incoherent garbage no human would ever write.

Maybe I’m the only one keeps trying this (more often than I’m willing to admit), but I suspect it’s a common cope engineers reach for when having to deal with the not-so-fun task of writing prose.

brap··on OpenJev
Can anyone please explain this Jev thing to me?

We’ve always had output schemas for LLMs, and we’ve had small language classifiers for decades, so what’s new? Is it just some sweet spot in between in terms of quality vs speed?

brap··on There's no point at which turning your brain off will work
Interesting thought. In a way this is already the case, to some extent.

But wouldn’t it be way more economical to have some sort of AI-insurance service? i.e. protection against AI fucking things up?

brap··on Bend – A language that blocks AI mistakes via proof, on CPU and GPU
What exactly enforces that an AI follows these rules?
brap··on Gemini 3.8 Live and 3.8 Live Extended Thinking
Also: Google has like a 20% stake in Anthropic, and a very fat cloud partnership
brap··on A misalignment of AI in mathematics
Right, but this leads to my main question: is this still necessary?

By analogy with code, do we still need code to be maintainable/readable if machines write it all?

(Obviously for now the answer is yes, but I’m not sure this will be the case in 5-10 years)

brap··on We must pace the frontier
China doesn't give a fuck, next
brap··on A misalignment of AI in mathematics
I know it's incredibly presumptuous for me, a nobody, to say this to 25 Fields Medalists, but:

Perhaps you are misaligned.

Who decided the goal of math must be human insight?

First off, some mathematical truths might simply be far beyond our biological comprehension.

Second, for us non-mathematicians, the value of math isn't in understanding exactly why a result is true, it’s in how those results can be applied to actually improve our lives.

Isn't this why we have math in the first place? To solve our real problems? Over time it morphed into this pursuit of pure theoretical insight, probably out of necessity at the time, but is it still necessary?

brap··on Measuring the sloppiness of code
There’s one thing I constantly see agents tripping over, I’m not sure what the right word for it would be, but it basically boils down to “making changes in the right places”. They seem to have very poor grasp of where things are supposed to be and they have a tendency to work against the existing architecture. Even in a world where agents are the only ones touching the code you can see how this ends poorly. Unlike correctness I’m not sure there’s an easy way to verify.

I tried writing a few skills to encourage agents to spend time thinking about this but it doesn’t seem to generalize very well.

brap··on More questions about whether researchers can trust OpenAI with unpublished math
I’m not a fan of OAI to say the least, but having worked at similar companies, my guess is that it’s just too difficult to prove/disprove beyond a doubt, and they have other priorities
brap··on OpenAI Agents API
I think the line between regular LLM "endpoints" and agents/harnesses is going to become more and more blurry until it's a meaningless distinction.

When you're using ChatGPT/Claude/Gemini etc. you're basically already interacting with some backend harness with tools etc., not a raw LLM. Just give it a computer and be done with it.

I already find myself using Claude Code / Antigravity (via web) instead of Claude / Gemini, even for tasks unrelated to coding. Why use a limited version?

brap··on iPhone Duo
Folded phones are the essence of “just because you can doesn’t mean you should”
brap··on Claude, change the “Add to Cart” button to blue
How do you manage your frustration in these interactions? I often find myself getting pissed off
brap··on Gemini 3.8 Flash and 3.8 Flash Cyber
I believe the older models are being gradually phased out, newer ones have no availability issues
brap··on Gemini 3.8 Flash and 3.8 Flash Cyber
Just like Claude Code and others it has the same —-dangerously-skip-permissions flag, auto approves everything
brap··on Gemini 3.8 Flash and 3.8 Flash Cyber
Antigravity has been also rapidly improving lately, and your can also use any of the open coding harnesses. But I mostly meant “harness” as in your workflow/loop setup.
brap··on Gemini 3.8 Flash and 3.8 Flash Cyber
People have been sleeping on Gemini lately but these last few Flash releases (which were very rapid) are damn good.

These sort of fast and cheap models are great for tasks that are verifiable and can be retried infinitely (like coding), you can basically get frontier results with a good harness (at a fraction of the time and money).

brap··on “I just chose words carefully”
Incredible. This is the kind of weaponized OCD I want in my team
brap··on Serve Markdown to AI Agents with Accept Headers
Best case scenario, this ends up being abused in order to feed LLMs crap responses (or worse).
brap··on Where did all the public bathrooms go?
Surely more social workers will stop people from pissing on the floor
brap··on The entire city of San Francisco as a video game
This is incredible, I wish I could have this in my city.

Can’t help but wondering, how will this look like if we had AI try to “augment” these maps in real time (maybe using street view images?). I wonder if it would be playable in reasonable FPS, and how expensive it would be.

brap··on How Europe is killing makers and micro-entrepreneurs
>A good idea, a terrible implementation

Idea is terrible, implementation is WAI.

brap··on New MCP Roadmap
I mean… so just HTTP + OpenAPI spec?
Page 1 of 31Next →