HNHacker News
TopNewBestAskShowJobs

marsh_mellow

203 karma · joined December 30, 2022

submissionscomments
marsh_mellow··on GPT-4.1 in the API
p-value of 7.9% — so very close to statistical significance.

the p-value for GPT-4.1 having a win rate of at least 49% is 4.92%, so we can say conclusively that GPT-4.1 is at least (essentially) evenly matched with Claude Sonnet 3.7, if not better.

Given that Claude Sonnet 3.7 has been generally considered to be the best (non-reasoning) model for coding, and given that GPT-4.1 is substantially cheaper ($2/million input, $8/million output vs. $3/million input, $15/million output), I think it's safe to say that this is significant news, although not a game changer

marsh_mellow··on GPT-4.1 in the API
I don't think the absolute score means much — judge models have a tendency to score around 7/10 lol

55% vs. 45% equates to about a 36 point difference in ELO. in chess that would be two players in the same league but one with a clear edge

marsh_mellow··on GPT-4.1 in the API
Good point. They said they validated the results by testing with other models (including Claude), as well as with manual sanity checks.

55% to 45% definitely isn't a blowout but it is meaningful — in terms of ELO it equates to about a 36 point difference. So not in a different league but definitely a clear edge

marsh_mellow··on GPT-4.1 in the API
From OpenAI's announcement:

> Qodo tested GPT‑4.1 head-to-head against Claude Sonnet 3.7 on generating high-quality code reviews from GitHub pull requests. Across 200 real-world pull requests with the same prompts and conditions, they found that GPT‑4.1 produced the better suggestion in 55% of cases. Notably, they found that GPT‑4.1 excels at both precision (knowing when not to make suggestions) and comprehensiveness (providing thorough analysis when warranted).

https://www.qodo.ai/blog/benchmarked-gpt-4-1/

marsh_mellow··on Show HN: RF Hunter – Find hidden cameras and other devices
Very cool! Could this work for detecting nearby drones?
marsh_mellow··on Computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku
Anthropic blog post outlining the research process: https://www.anthropic.com/news/developing-computer-use

Computer use API documentation: https://docs.anthropic.com/en/docs/build-with-claude/compute...

Computer Use Demo: https://github.com/anthropics/anthropic-quickstarts/tree/mai...

marsh_mellow··on Ask HN: What is your experience with Cursor?
To tag on to this, what are the most useful capabilities besides code generation?
marsh_mellow··on Consent-O-Matic – automatically fills ubiquitous pop-ups with your preferences
This is great. Is there any work being done to make something similar part of the browser API?
marsh_mellow··on PR-Agent — extension that adds AI chat to code reviews on GitHub
There's an open source version of this as well: https://github.com/Codium-ai/pr-agent
marsh_mellow··on Radiology-specific foundation model
They list seven different use cases in this technical blog:

https://harrison.ai/news/reimagining-medical-ai-with-the-mos...

I'd interpret it as a foundation model in the radiology domain