HNHacker News
TopNewBestAskShowJobs

henriquegodoy

264 karma · joined January 30, 2024

henriquegodoy.com
submissionscomments
henriquegodoy··on Launch HN: Design Arena (YC S25) – Head-to-head AI benchmark for aesthetics
This is actually really needed, current ai design tools are so predictable and formulaic, like every output feels like the same purple gradients with rounded corners and that one specific sans serif font that every model seems obsessed with, it's gotten to the point where you can spot ai-generated designs from a mile away because they all have this weird sterile aesthetic that screams "made by a model"
henriquegodoy··on Show HN: Omnara – Run Claude Code from anywhere
This is pretty cool and feels like we're heading in the right direction, the whole idea of being able to hop between devices while claude code is thinking through problems is neat, but honestly what excites me more is the broader pattern here, like we're moving toward a world where coding isn't really about sitting down and grinding out syntax for hours, it's becoming more about organizing tasks and letting ai agents figure out the implementation details.

I can already see how this evolves into something where you're basically managing a team of specialized agents rather than doing the actual coding, you set up some high-level goals, maybe break them down into chunks, and then different agents pick up different pieces and coordinate with each other, the human becomes more like a project manager making decisions when the agents get stuck or need direction, imho tools like omnara are just the first step toward that, right now it's one agent that needs your input occasionally, but eventually it'll probably be orchestrating multiple agents working in parallel, way better than sitting there watching progress bars for 10 minutes.

henriquegodoy··on Evaluating LLMs playing text adventures
Looking at this evaluation it's pretty fascinating how badly these models perform even on decades old games that almost certainly have walkthroughs scattered all over their training data. Like, you'd think they'd at least brute force their way through the early game mechanics by now, but honestly this kinda validates something I've been thinking about like real intelligence isn't just about having seen the answers before, it's about being good at games and specifically new situations where you can't just pattern match your way out

This is exactly why something like arc-agi-3 feels so important right now. Instead of static benchmarks that these models can basically brute force with enough training data, like designing around interactive environments where you actually need to perceive, decide, and act over multiple steps without prior instructions, that shift from "can you reproduce known patterns" to "can you figure out new patterns" seems like the real test of intelligence.

What's clever about the game environment approach is that it captures something fundamental about human intelligence that static benchmarks miss entirely, like, when humans encounter a new game, we explore, form plans, remember what worked, adjust our strategy all that interactive reasoning over time that these text adventure results show llms are terrible at, we need systems that can actually understand and adapt to new situations, not just really good autocomplete engines that happen to know a lot of trivia.

henriquegodoy··on Claude Sonnet 4 now supports 1M tokens of context
Thats incredible to see how ai models are improving, i'm really happy with this news. (imo it's more impactful than the release of gpt5) now, we need more tokens per second, and then the self-improvement of the model will accelerate.
henriquegodoy··on GPT-5
That SWE-bench chart with the mismatched bars (52.8% somehow appearing larger than 69.1%) was emblematic of the entire presentation - rushed and underwhelming. It's the kind of error that would get flagged in any internal review, yet here it is in a billion-dollar product launch. Combined with the Bernoulli effect demo confidently explaining how airplane wings work incorrectly (the equal transit time fallacy that NASA explicitly debunks), it doesn't inspire confidence in either the model's capabilities or OpenAI's quality control.

The actual benchmark improvements are marginal at best - we're talking single-digit percentage gains over o3 on most metrics, which hardly justifies a major version bump. What we're seeing looks more like the plateau of an S-curve than a breakthrough. The pricing is competitive ($1.25/1M input tokens vs Claude's $15), but that's about optimization and economics, not the fundamental leap forward that "GPT-5" implies. Even their "unified system" turns out to be multiple models with a router, essentially admitting that the end-to-end training approach has hit diminishing returns.

The irony is that while OpenAI maintains their secretive culture (remember when they claimed o1 used tree search instead of RL?), their competitors are catching up or surpassing them. Claude has been consistently better for coding tasks, Gemini 2.5 Pro has more recent training data, and everyone seems to be converging on similar performance levels. This launch feels less like a victory lap and more like OpenAI trying to maintain relevance while the rest of the field has caught up. Looking forward to seeing what Gemini 3.0 brings to the table.

henriquegodoy··on GPT-5 for Developers
I dont think there's so much difference from opus 4.1 and gpt-5, probably just the context size, waiting for the gemini 3.0
henriquegodoy··on I gave the AI arms and legs then it rejected me
I think this blog post was the best way to get into Anthropic, and it was well-deserved. That's the reality of hiring in tech: there are many non-technical people judging whether technical people are competent or not. Escaping that matrix through things like blog posts, cold emails, and Twitter threads can be great ways to break in and get noticed by these companies.
henriquegodoy··on Open models by OpenAI
Seeing a 20B model competing with o3's performance is mind blowing like just a year ago, most of us would've called this impossible - not just the intelligence leap, but getting this level of capability in such a compact size.

I think that the point that makes me more excited is that we can train trillion-parameter giants and distill them down to just billions without losing the magic. Imagine coding with Claude 4 Opus-level intelligence packed into a 10B model running locally at 2000 tokens/sec - like instant AI collaboration. That would fundamentally change how we develop software.

henriquegodoy··on Vibe code is legacy code
I'm seeing a real-world example of Jevons paradox playing out here. When AI coding tools first emerged, everyone predicted mass developer unemployment. Instead, I'm watching demand for skilled developers actually increase.

What's happening is that all this "vibe coded" software needs someone to fix it when it breaks. I've been getting more requests than ever to debug AI-generated codebases where the original "developer" can't explain what any of it does. The security audit work alone is keeping me busy - these AI-generated apps often have vulnerabilities that would never pass a human code review. It reminds me of when WordPress democratized web development. Suddenly everyone could build a website, but that just created a massive market for developers who could fix broken WordPress sites, migrate databases, and patch security holes. The difference now is the scale and complexity. At least with WordPress, there was some underlying structure you could reason about. With vibe coding, you get these sprawling codebases where the AI has reinvented the wheel five different ways in the same project, used deprecated libraries because they were in its training data, and created bizarre architectural decisions that only make sense if you don't understand the problem domain.

So yeah, the jobs aren't disappearing - they're just shifting from "build new features" to "fix the mess the PM made last weekend when they tried to ship their own feature."

henriquegodoy··on I launched 17 side projects. Result? I'm rich in expired domains
Ever thinked on automating this process of creating this side projects? i think that more and more future feels like a lot of ones having really big swarms of "agents" that can like research about ideas on the internet (like finding problems on twitter, reddit, ... that a saas can solve it) and a team implementing and deploying since from code to marketing in a frenetic rhythm
henriquegodoy··on Launch HN: Lucidic (YC W25) – Debug, test, and evaluate AI agents in production
my vision is that the market is not really prepared for that right now, the best way is this guys is solving a really niche problem with their plataform and then expanding trough more areas
henriquegodoy··on Launch HN: Lucidic (YC W25) – Debug, test, and evaluate AI agents in production
Nice, i think that yall are on the correct path betting on evals, but please make your ui less "generic"
henriquegodoy··on Fast
Will apply this for the next interfaces that im going to build
henriquegodoy··on Show HN: A GitHub Action that quizzes you on a pull request
can i automate the process of answering this pr questions too?
henriquegodoy··on Supervised fine tuning on curated data is reinforcement learning
It's cool to see the perspective that many problems (somekinda communication problems, look at lawyers, compliance and etc...) can be solved by treating AI less as agents and more as modular components within a larger system. Once we build a working process—monitored through evals—we can then reduce costs by distilling these modules. That means starting with superintelligent models and later distilling them down to just a few billion parameters, instead of needing hundreds of billions.
henriquegodoy··on Study mode
The point is that you can have a highly advanced teacher with infinite patience, available 24/7—even when you have a question at 3 a.m is game changer and people that know how to use that will have a extremaly leverage in their life.
henriquegodoy··on Principles for production AI agents
I've been tinkering with agentic systems for a while now, and this post nails some key pain points that hit close to home. The emphasis on splitting context and designing tight feedback loops feels spot on—I've seen agents go off the rails without them, hallucinating solutions because the prompt was too bloated or the validation was half-baked. It's like building a machine where every part needs to click just right, or else you're debugging forever.

What really resonates is the bit about frustrating behaviors signaling deeper system issues, not just model quirks. In my own experiments, I've had agents stubbornly ignore tools because I forgot to expose the right APIs, and it made me rethink how we treat these as "intelligent" when they're really just following our flawed setups. It pushes us toward more robust orchestration, where humans handle the high-level intentions and AI fills in the execution gaps seamlessly.

This ties into broader ideas on how AI interfaces will evolve as models get smarter. I extrapolate more of this thinking and dive deeper into human–AI interfaces on my blog if anyone’s interested in checking it out: https://henriquegodoy.com/blog/stream-of-consciousness

henriquegodoy··on Viral Language
crazy how the behaviour impacts the language and language impacts the behaviour such like a loop
henriquegodoy··on Blender: Beyond Mouse and Keyboard
Speaking of interfaces, when will we have one that works just by thinking—something less intrusive than Neuralink—that lets us control not just Blender, but the entire computer? I think my productivity would increase a lot...
henriquegodoy··on GPT might be an information virus (2023)
IMO AI as a whole is just the catalyst
henriquegodoy··on Enough AI copilots, we need AI HUDs
Great post! i've been thinking along similar lines about human-AI interfaces beyond the copilot paradigm. I see two major patterns emerging:

Orchestration platforms - Evolution of tools like n8n/Make into cybernetic process design systems where each node is an intelligent agent with its own optimization criteria. The key insight: treat processes as processes, not anthropomorphize LLMs as humans. Build walls around probabilistic systems to ensure deterministic outcomes where needed. This solves massive "communication problems"

Oracle systems - AI that holds entire organizations in working memory, understanding temporal context and extracting implicit knowledge from all communications. Not just storage but active synthesis. Imagine AI digesting every email/doc/meeting to build a living organizational consciousness that identifies patterns humans miss and generates strategic insights.

just explored more about it on my personal blog https://henriquegodoy.com/blog/stream-of-consciousness

henriquegodoy··on Ask HN: Who wants to be hired? (September 2024)
Location: Brazil (Available 4AM to 9PM EST)/Remote

Remote: Yes Wiling to realocate: yes

Résumé/CV: https://drive.google.com/file/d/1VbOYflwT3crNqzCq8rQH7Mz31dC... / https://www.linkedin.com/in/henrique-godoy-879138252/

Email: henrique.godoy@sou.inteli.edu.br

Hello! I'm Henrique. I've been programming since I was 14 years old. I was approved for one of the most challenging math programs at USP (University of São Paulo) at the age of 15. At 17, I received a 500k scholarship at Inteli, working directly with companies such as Dell, Pirelli, Vivo, and more. I love mathematics too; last year, I achieved a top 200 ranking in OBMU (Brazilian Mathematical Olympiad for University Students), which is similar to the American Putnam Competition. I helped implement the first LLM in the largest Latin American Investment Bank, and I've already founded a company called Doki, where I developed a personal finance app for doctors in Brazil. I led the full development lifecycle, which achieved over R$10 million in registered funds. Right now, I am 19 years old and in my 3rd year of Computer Science at Inteli (only 3 hours a day at present). I'm looking for an opportunity to work with smart people solving cool problems; salary, language, or anything else is not a factor. :)