HNHacker News
TopNewBestAskShowJobs

parhamn

4,905 karma · joined October 6, 2012

Building:

https://bearly.ai - LLM toolkit https://synth.app - Research browser

Email: pnegahdar@gmail, or parham@bearly.ai

submissionscomments
parhamn··on Hacking OpenAI
> Valuations for server-side vulnerabilities are low, because vendors don't compete for them.

Why don't they?

parhamn··on OpenAI’s Navier-Stokes release included a Lean 4 formal proof
They estimated $40M of agent costs (it was a large fleet of them). Using the number in the post its closer to ~880,000 hours × $150/hour = $132 million for the human case. Still an amazing feat not quite "four orders of magnitude". The comparison is obviously pointless because coordinating 1M hours of intellectual labor isn't easy to say the least.

Very exciting and uncertain times!

parhamn··on Mistral OCR 4.1
Mistral is bumping the price of this thing every release. I think we're at 2x now?
parhamn··on What sort of maths are LLMs good at?
Thats the issue, they're going to get flooded soon with these holy blessings, then what?
parhamn··on What sort of maths are LLMs good at?
> If they were, then their big speed advantage over us would mean that there would be much more of a flood of results.

Is this true right now? Just recently Jarred Sumner tweeted [1] that he managed to make some progress on the Riemann hypothesis while on a jog. Managed to get somewhere by encouraging the llm to “keep going” and “believe in yourself”.

This raised a few questions for me. Had no one at Anthropic thought to try this earlier? It's an interesting footnote that a software engineer there pursued this. How many people in the world can actually verify a proof? How many would we need to sit around and do the right incantations to get a proof out of it? How many would we need to verify and give those proofs value and meaning? What happens when there are more proofs than verifiers? How many will be around in 100 years?

I think it just turns out that a lot this stuff is more socially useful than anything else. The 10 proofs drop came and went in the daily news cycle. Perhaps math is already in it's chess like "for fun" period. I am interested in when we find a very high real-world utility breakthrough math/physics, some space where we've already poured our best human resources at it.

[1] https://x.com/jarredsumner/status/2086869681785500011?s=20

parhamn··on Cheap self-hosted Kubernetes on Hetzner cloud (2025)
Terraform has been such a delight in the age of LLMs.
parhamn··on Protobuf-py: Protobuf for Python, without compromises
Whats the case for protobuf these days? I loved it for a while. Don't dislike it particularly now (but I've rolled off).

- Performance? Fastest JSON lib in most languages are as fast if not faster

- Cross language schema generation? So many tools for this these days, they do one thing do them right (depending on your choice of 'right' for things like unions/enums/etc)

- The wire protocol? Seems to get in the way vs http2/3. Need special considerations for your proxies (be it nginx or cloudflare). Forced certs, etc are annoying too.

I feel like to most shops these days its mostly a schema manager? Protobuf is super bloated for that use case.

parhamn··on [dead]
"OAuth 2.0 Client ID Metadata Document (CIMD)" is a big one. As an agent-provider, prereigstering client IDs for a bunch of different services and going through each ones special hoops for org validation sucks.
parhamn··on Launch HN: Chert (YC P26) – Twilio for iMessage
> anyone who needs to depend on this service would want to hear

Are you implying you'd be cool with it if it was Apple sanctioned? That's pretty silly.

parhamn··on OpenAI Adopts Google's SynthID Watermark for AI Images with Verification Tool
might be easier to strip it?
parhamn··on Zerostack – A Unix-inspired coding agent written in pure Rust
Not daily driver, but have used it as a utility a few times.

For my daily work I like letting different harnesses compete and look over each others work (while subsidized with the subscriptions) so I use OpenADE.

parhamn··on Zerostack – A Unix-inspired coding agent written in pure Rust
I (somewhat jokingly) wrote one recently too... https://github.com/pnegahdar/nano in under 200 lines. Repl, sessions, non-interactive, approvals, etc

The smarter the models get the less the harnesses matter (outside of devx).

Maybe one day I'll run it through swebech.

parhamn··on Hackers breach JDownloader's website to serve malware-laced downloads
This sent me down memory lane, what happened to warez-bb!
parhamn··on LittleSnitch for Linux
Or even sell the whole org for say $50M and no one ever mentions anything.

I think the type of users it attracts (techies, crypto ppl, etc) makes it worth more too.

parhamn··on LittleSnitch for Linux
Okay hear me out, I use little snitch for a while. Great product. Love finding out what phones where. I make every single request (except my browser, because I'm fine with their sandbox) block until I approve.

Recently I was wondering how you really have to trust something like little snitch given its a full kernel extension effectively able to MITM your whole network stack.

So I went digging (and asked some agents to deep research), and I couldn't find much interesting about the company or its leadership at all.

All a long way to say, anyone know anything about this company?

parhamn··on Show HN: Ghost Pepper – Local hold-to-talk speech-to-text for macOS
I see a lot of whisper stuff out there. Are these the same old OpenAI whispers or have they been updated heavily?

I've been using parakeet v3 which is fantastic (and tiny). Confused why we're still seeing whisper out there, there's been a lot of development.

parhamn··on Shall I implement it? No
I added an "Ask" button my agent UI (openade.ai) specifically because of this!
parhamn··on WebMCP is available for early preview
Is the website Stripe or NYTimes?
parhamn··on Statement from Dario Amodei on our discussions with the Department of War
Now, I'm curious. How Bedrock/Azure Claude models work?

Do these rules apply to them too?

parhamn··on Banned in California
I'm not from California but this to me seems like a great case to move to California. Why not ship your externality creating activities elsewhere? Its not like they pay more for the iPhone.
parhamn··on Show HN: Price Per Ball – Site that sorts golf balls on Amazon by price per ball
gemini flash!
parhamn··on Show HN: Free alternative to Wispr Flow, Superwhisper, and Monologue
+1, my experience improved quite a bit when I switched to the parakeet model, they should definitely use that as the default.
parhamn··on Anthropic tries to hide Claude's AI actions. Devs hate it
the project just does subprocess calls to claude code (the product/cli). I think services like open code were using it to make raw requests to claude api. Have any more context I can look into?
parhamn··on Anthropic tries to hide Claude's AI actions. Devs hate it
I think my read of "hiding" was more of a "trying to hide the secret sauce" which was implied in a few places.

Otherwise it seems like a minor UI decision any other app would make and it surprising there's whole articles on it.

parhamn··on Anthropic tries to hide Claude's AI actions. Devs hate it
"Hiding" is doing some heavy lifting here. You can run --json and see everything pretty much (besides the system prompt and tool descriptions)....

I love the terminal more than the next guy but at some point it feels like you're looking at production nginx logs, just a useless stream of info that is very difficult to parse.

I vibe coded my own ADE for this called OpenADE (https://github.com/bearlyai/openade) it uses the native harnesses, has nice UIs and even comes with things like letting Claude and Codex work together on plans. Still very beta but has been my daily driver for a few weeks now.

parhamn··on AWS Adds support for nested virtualization
whats the ~ perf hit of something like this?
parhamn··on Improving 15 LLMs at Coding in One Afternoon. Only the Harness Changed
You didn't read the article it seems (or the analogy is a bad one). The differences are much more subtle than having a screwdriver or not.
parhamn··on Improving 15 LLMs at Coding in One Afternoon. Only the Harness Changed
On first principles it would seem that the "harness" is a myth. Surely a model like Opus 4.6/Codex 5.3 which can reason about complex functions and data flows across many files would trip up over top level function signatures it needs to call?

I see a lot of evidence to the contrary though. Anyone know what the underlying issue here is?

parhamn··on Claude Code is being dumbed down?
That looks great! Planning phase is really key.
parhamn··on Claude Code is being dumbed down?
We opensourced our claude code ui today: https://github.com/bearlyai/openade

I wanted a terminal feel (dense/sharp) + being able to comment directly on plans and outputs. It's MIT, no cloud, all local, etc.

It includes all the details for function runs and some other nice to haves, fully built on claude code.

Particularly we found planning + commenting up front reduces a lot of slop. Opus 4.6 class models are really good at executing an existing plan down to a T. So quality becomes a function of how much you invest in the plan.

Page 1 of 22Next →