HNHacker News
TopNewBestAskShowJobs

joaogui1

334 karma · joined February 23, 2019

submissionscomments
joaogui1··on The Normalization of Inexplicable Failures
No one moved everything to Rust on the first year the language was released either, and you can get analogies for all of your examples. But it does look like the majority of people have jumped into agentic coding
joaogui1··on Gemini Robotics 2 brings whole body intelligence to robots
I won't comment much on this, but I'd say these articles are 2-3 years late and reduces the agency of the people, who in many cases moved to Gemini because they wanted to work on LLMs
joaogui1··on Leak Exposes Members of Peter Thiel's Secretive 'Dialog' Society
I think globally they’re all center-right to alt-right, which of them are considered leftwing in the US?
joaogui1··on Gemma 4 12B: A unified, encoder-free multimodal model
I mean Claude is multimodal on input but not output, why couldn't this also be?
joaogui1··on If America's so rich, how'd it get so sad?
I believe their justification is on the first sentence

> That sort of stuff causes pitchforks to rise up in other countries.

(Not that I agree)

joaogui1··on Qwen3.6-35B-A3B: Agentic coding power, now open to all
3.6 is model number, 35B is total number of parameters, A3B means that only 3B parameters are activated, which has some implications for serving (either in you you shard the model, or you can keep the total params on RAM and only road to VRAM what you need to compute the current token, which will make it slower, but at least it runs)
joaogui1··on AI overly affirms users asking for personal advice
The models get deprecated after 1-2 years, so reproducibility is pretty hard anyway (but as others pointed out the paper does list the model versions)
joaogui1··on The highest quality codebase
During pre-training the model is learning next-token prediction, which is naturally additive. Even if you added DEL as a token it would still be quite hard to change the data so that it can be used in a mext-token prediction task Hope that helps
joaogui1··on Show HN: Gemini Pro 3 imagines the HN front page 10 years from now
HN has been used to train LLMs for a while now, I think it was in the Pile even
joaogui1··on Gemini 3
Probably figured out the exact cause of the bug but not how to solve it
joaogui1··on Gemini 3
It says Gemini App, not AI Overviews, AI Mode, etc
joaogui1··on Why is Zig so cool?
Also bizarre that it got to the front page of HN while being so low quality :/
joaogui1··on Signs of introspection in large language models
Anthropic has amazing scientists and engineers, but when it comes to results that align with the narrative of LLMs being conscious, or intelligent, or similar properties, they tend to blow the results out of proportion

Edit: In my opinion at least, maybe they would say that if models are exhibiting that stuff 20% of the time nowadays then we’re a few years away from that reaching > 50%, or some other argument that I would disagree with probably

joaogui1··on Ink deformation
It's their lab notes, so it's exploring a general idea, but they're also referencing previous software they've built (like crosscut)
joaogui1··on We're Joining OpenAI
Ads on ChatGPT as a way to extract more money from users
joaogui1··on From M1 MacBook to Arch Linux: A month-long experiment that became permanenent
Not necessarily meaningless, but maybe relative, i.e. a person who generally replaces non-Apple laptops every X years would replace MacBooks every Y years, with Y > X
joaogui1··on Gemini 2.5 Deep Think
Mixture of Experts isn't using multiple models with different specialties, it's more like a sparsity technique, where you massively increase the number of parameters and use only a subset of the weights in each forward pass.
joaogui1··on Gemini with Deep Think achieves gold-medal standard at the IMO
It's not a new product/model with that name, they're just saying it's an advanced version of Gemini that's not public atm
joaogui1··on Grok 4
They used a text-only subset of HLE
joaogui1··on End of an Era
DS9 the Star Trek series?
joaogui1··on Unsupervised Elicitation of Language Models
The fact that she's a scientist communicator doesn't imply that she only did the communication part, I think
joaogui1··on Expanding on what we missed with sycophancy
Employees from OpenAI encouraged people to use ChatGPT as their therapist, so yeah, they now have to take responsibility for it
joaogui1··on Pope Francis has died
I believe you will find the majority of progressive Jews has said that Israel, and more specifically the government of Israel, does not speak for Jews worldwide. In fact many rabbis have written about the nauseating position of having Israel be considered a representative of all Jews by so many people
joaogui1··on Released Llama 4 Maverick places 32nd in LMArena
And 23 when taking Style Control into account
joaogui1··on The Llama 4 herd
I don't want to hunt the details on each of theses releases, but

* You can use less GPUs if you decrease batch size and increase number of steps, which would lead to a longer training time

* FP8 is pretty efficient, if Grok was trained with BF16 then LLama 4 should could need less GPUs because of that

* Depends also on size of the model and number of tokens used for training, unclear whether the total FLOPS for each model is the same

* MFU/Maximum Float Utilization can also vary depending on the setup, which also means that if you're use better kernels and/or better sharding you can reduce the number of GPUs needed

joaogui1··on The Llama 4 herd
I mean they're not comparing with Gemini 2.5, or the o-series of models, so not sure they're really beating the first point (and their best model is not even released yet)

Is the new license different? Or is it still failing for the same issues pointed by the second point?

I think the problem with the 3rd point is that LeCun is not leading LLama, right? So this doesn't change things, thought mostly because it wasn't a good consideration before

joaogui1··on Gemini 2.5: Our most intelligent AI model
Would be confusing for non-tech people once you did x.9 -> x.10
joaogui1··on Anthropic hires OpenAI co-founder Durk Kingma
I think the whole situation where they got some serious investment from SBF and then he got indicted pushed them into commercialising their tech so they could have more standard sources of funding
joaogui1··on AGI is far from inevitable
Are you implying that our brain learns through Machine Learning?
joaogui1··on OpenAI threatens to revoke o1 access for asking it about its chain of thought
OpenAI keeps innovating on being more closed than the other companies
Page 1 of 5Next →