HNHacker News
TopNewBestAskShowJobs

josu

3,275 karma · joined November 5, 2013

@josusanmartin
submissionscomments
josu··on OpenAI GPT–6 Astra breaks Enigma message that has resisted solution since 2005
8 months ago I got to the top of highload.fun using GPT-5 and Opus 4.5, and a lot of human interaction.

Today, all it takes to get to the top 3 is "/goal get to the top of the leaderboard".

The human-in-the-loop is only a temporary measure until the models get good enough.

josu··on Solving the Jane Street reverse engineering challenge
Thanks for pushing back, I've edited my initial post. I misinterpreted the response, I thought that it only referenced the solution as verification.
josu··on Solving the Jane Street reverse engineering challenge
I gave the problem to chatGPT 5.6 Sol Pro and this was the result:

> Worked for 12m 36s

> Solved

https://chatgpt.com/s/t_6a9aed0b09988191b0f2850dee056b48

Edit: It didn't independently solve it.

> 1. Used the public reconstruction to obtain the recovered RTL/constraint structure, including the 11×11 region map and the fact that it is a two-stars-per-row/column/region, non-touching puzzle.

> 2. Then independently wrote and ran my own exhaustive solver against that recovered constraint system.

josu··on GPT-5.6 Sol Pricing Cut by 50%
SoftBank is one of the largest investors in openAI having contributed more than 30B and pledged another 30.
josu··on GPT-5.6 Sol Pricing Cut by 50% on OpenRouter
I always find it funny that Japanese pensioners are probably subsidizing my tokens.
josu··on Auto-research with codex: How I achieved a 232x Faster Kernel
Overfitting to the input is part of the meta in this type of challenges.

The goal is not to create good, general or maintainable code. The only goal is to produce the fastest code.

josu··on Mistral OCR 4.1
Whats the best open OCR at the moment?
josu··on Windows 11's built-in Weather app wastes more than 1 GB of RAM
Here, I'm building a Torrent application from specs, and keeping the ARM only binary under 1Mb: https://github.com/josusanmartin/rustorrent

It's fully vibecoded, and quite functional, although I'm still working out some bugs.

josu··on Show HN: Run an 80B Qwen in 4.3 GB of RAM on a Mac, and a 35B on an iPhone
Thanks for responding.
josu··on Show HN: Run an 80B Qwen in 4.3 GB of RAM on a Mac, and a 35B on an iPhone
I don't understand, if they are only using a subset of the tokens then it's a sparse model. What do you mean by dense?
josu··on LLM Usage in Debian: Three Proposals
I was responding to this "merely produces syntactically likely combinations of the training data".
josu··on LLM Usage in Debian: Three Proposals
Yes, like the Jacobian counterexample.
josu··on LLM Usage in Debian: Three Proposals
The combination of the most probable tokens doesn't necessarily have to be in the training set, thus creating something completely novel.
josu··on Cloudflare Drop
It's a spectrum. I don't think that they will be able to sell t-shirts using a drawing you've uploaded. But it will probably allow them to defend themselves a bit better if they get sued for selling the data for LLM training.
josu··on Nintendo announces new product revisions in Europe with replaceable batteries
The console one seems the only relevant one. I'm a casual gamer, but the joycon running out of battery doesn't feel that annoying to me. That 16% has a much smaller impact on the JoyCon than on the main device.
josu··on Nintendo announces new product revisions in Europe with replaceable batteries
Is it?

>Battery capacity: 5172mAh, approximately 1% smaller than current version (5220mAh)

josu··on Spain Orders Blacklist of Palantir from Public and Private Companies
>Spain is not Somalia, why not let Indra do it?

The data may be safer with the CCP, at least they won't lose it.

josu··on CPanel's Black Week: 3 New Vulnerabilities Patched After Attack on 44k Servers
So CPanel's security is just as bad as their UI, who would have thought?
josu··on A recent experience with ChatGPT 5.5 Pro
LLMs can also be really good in fields where you are not an expert. You just need to be very aware of your limitations, and start parallel conversation so one agent fact checks the other.
josu··on Gemini Robotics-ER 1.6
Done! The whole planet is now veggies.
josu··on Spain to expand internet blocks to tennis, golf, movies broadcasting times
>enact similarly stupid laws.

No new law was enacted. The ISPs are enforcing a court order.

josu··on Waymo seeking about $16B near $110B valuation
>I wonder what will happen to the drivers if a large representation of the 1 million+ daily trips are displaced by automation?

If it happens gradually enough, they will just find other jobs. After the transition, society will be producing more with the same labor force, and thus the aggregate utility will increase.

josu··on SoftBank in talks to invest up to $30B more in OpenAI
> invested in (...) anything else practical.

I don't understand how this is the top comment. LLMs have unlocked a lot of value for me personally, and arguably for the society as a whole. They are also one of the coolest technologies I've tried in years. As a technologist, I'm really glad that money is pouring in and allowing us to find its limits.

josu··on I vibecoded my way into the #1 position on the Highload.fun leaderboard
How is it cheating? Do you also need to independently discover Bloom filters to be able to use them?
josu··on Opus 4.5 is not the normal AI agent experience that I have had thus far
I think you misunderstood, it's not about solving the problem, is about finding the most efficient solution. Give it a shot, and see if you can get to the top 10 on any task.
josu··on Opus 4.5 is not the normal AI agent experience that I have had thus far
Thank you. Yeah, I'm doing all those things, which do get you close to the top. The rest of things I'm doing are mostly micro-optimizations such as finding a way to avoid AVX→SSE transition penalty (1-2% improvement).

But I don't want to spoil the fun. The agents are really good at searching the web now, so posting the tricks here is basically breaking the challenge.

For example, chatGPT was able to find Matt's blog post regarding Task 1, and that's what gave me the largest jump: https://blog.mattstuchlik.com/2024/07/12/summing-integers-fa...

Interestingly, it seems that Matt's post is not on the training data of any of the major LLMs.

josu··on Opus 4.5 is not the normal AI agent experience that I have had thus far
> The above is a software engineering problem. Reimplementing a JSON parser using Opus is not fun nor useful, so that should not be used as a metric.

I've also built a bitorrent implementation from the specs in rust where I'm keeping the binary under 1MB. It supports all active and accepted BEPs: https://www.bittorrent.org/beps/bep_0000.html

Again, I literally don't know how to write a hello world in rust.

I also vibe coded a trading system that is connected to 6 trading venues. This was a fun weekend project but it ended up making +20k of pure arbitrage with just 10k of working capital. I'm not sure this proves my point, because while I don't consider myself a programmer, I did use Python, a language that I'm somewhat familiar with.

So yeah, I get what you are saying, but I don't agree. I used highload as an example, because it is an objective way of showing that a combination of LLM/agents with some guidance (from someone with no prior experience in this type of high performing architecture) was able to beat all human software developers that have taken these challenges.

josu··on Opus 4.5 is not the normal AI agent experience that I have had thus far
You are looking at this wrong. Creating a json parser is trivial. The thing is that my one-shot attempt was 10x slower than my final solution.

Creating a parser for this challenge that is 10x more efficient than a simple approach does require deep understanding of what you are doing. It requires optimizing the hot loop (among other things) that 90-95% of software developers wouldn't know how to do. It requires deep understanding of the AVX2 architecture.

Here you can read more about these challenges: https://blog.mattstuchlik.com/2024/07/12/summing-integers-fa...

josu··on Opus 4.5 is not the normal AI agent experience that I have had thus far
I know what's like running a business, and building complex systems. That's not the point.

I used highload as an example because it seems like an objective rebuttal to the claim that "but it can't tackle those complex problems by itself."

And regarding this:

"Claude is very useful but it's not yet anywhere near as good as a human software developer. Like an excitable puppy it needs to be kept on a short leash"

Again, a combination of LLM/agents with some guidance (from someone with no prior experience in this type of high performing architecture) was able to beat all human software developers that have taken these challenges.

josu··on Opus 4.5 is not the normal AI agent experience that I have had thus far
Lol, the problem is not finding a solution, the problem is solving it in the most efficient way.

If you think you can beat an LLM, the leaderboard is right there.

Page 1 of 34Next →