HNHacker News
TopNewBestAskShowJobs

joaogui1

334 karma · joined February 23, 2019

submissionscomments
joaogui1··on Diffusion Is Spectral Autoregression
I mean you didn't mention autoregressive models anywhere in your comment, whereas the post is about the connection between diffusion and autoregressive modelling. Also it's a blog post, if it has figured out a speed-up or improved method it would probably have been a paper
joaogui1··on Brazilian court orders suspension of X
What did Canada and UK do?
joaogui1··on Creativity has left the chat: The price of debiasing language models
Depends on a ton of stuff really, like size of the model, how long do you want to train it for, what exactly do you mean by "like Hacker News or Wikipedia". Both Wikipedia and Hacker News are pretty small by current LLM training sets standards, so if you train only on for example a combination of these 2 you would likely end up with a model that lacks most capabilities we associate with large language models nowadays
joaogui1··on Elixir and Machine Learning in 2024 so far: MLIR, Arrow, structured LLM, etc.
XLA tends tends to be better optimized for TPUs, Pytorch is better with GPUs, but I believe you can choose a backend when using Nx.
joaogui1··on OpenAI's Long-Term AI Risk Team Has Disbanded
The ads are definitely coming given their pitch deck for the data partnerships https://www.adweek.com/media/openai-preferred-publisher-prog...
joaogui1··on PaliGemma: Open-Source Multimodal Model by Google
Gemini 1.5 Ultra was never announced
joaogui1··on PaliGemma: Open-Source Multimodal Model by Google
I think (iii) is about models trained using Gemma output, while the "For clarity" part says that the Output itself is not a Model Derivative
joaogui1··on Gemini Flash
I would say 2 big problems are:

1. latency, which would get worse if you have to sequentially generate more output

2. These models very roughly turn tokens -> "average meaning" on the embedding layer, followed by attention layers that combine the meanings, and feed forward layers that match the current meaning combination to some kind of learned archetype/prototype almost. When you move from word parts to characters all of that becomes more confusing (what's the average meaning of a?) and so I don't think there are good enough techniques to learn character-based models yet

joaogui1··on AlphaFold 3 predicts the structure and interactions of life's molecules
We can predict the digits of pi with a formula, to me that counts as grasping it
joaogui1··on Visualizing Attention, a Transformer's Heart [video]
Mixture of Experts is not just 16 copies of a network, it's a single network where for the feed forward layers the tokens are routed to different experts, but the attention layers are still shared. Also there are interesting choices around how the routing works and I believe the exact details of what OpenAI is doing are not public. In fact I believe someone making a visualization of that would dispell a ton of myths around what are MoEs and how they work
joaogui1··on OpenAI – transformer debugger release
In vitro fertilization too!
joaogui1··on Google search drops cache link from search results
Most of the people getting 100k are not the people making the dumb decisions. Besides, lots of cool research is still happening inside Google
joaogui1··on Something peculiar in my 2yo's bedroom led me to a revelation about our universe
Understood, thanks!
joaogui1··on Something peculiar in my 2yo's bedroom led me to a revelation about our universe
What's the definition of chaos here? I thought chaos implied the measurement error growing quickly during propagation, but here it looks like it's growing pretty slowly. In what way is this more chaotic than the movement of a single accelerating body (where the error grows linearly with time)?
joaogui1··on Math Team
He did say he's not a mathematician
joaogui1··on ChatGPT generates fake data set to support scientific hypothesis
People are probably thinking data more complex than normal distributions (though I'm also not sure if GPT-4 is the best method for that)
joaogui1··on GraphCast: AI model for weather forecasting
The issue with chaotic systems is not data, is that the error grows superlinearly with time, and since you always start with some kind of error (normally due to measurement limitations) this means that after a certain time horizon the error becomes to significant to trust the prediction. That hasn't a lot to do with data quality for ML models
joaogui1··on GraphCast: AI model for weather forecasting
I think 10 days is basically the normal term for weather, in that we can get decent predictions for that span using "classical"/non-ML methods.
joaogui1··on Llemma: An Open Language Model for Mathematics
One of the main points of Llemma is being an open-source reproduction of Minerva, so basically that's the only comparison they "need" to make. Besides, isn't WizardMath trained on GPT-4 output? They may want to be comparing only among models that can be used commercially
joaogui1··on Tiny Language Models Come of Age
Since GPT-3 OpenAI has been filtering their pre-training data, and I believe others have done it too
joaogui1··on OpenAI's justification for why training data is fair use, not infringement [pdf]
Where does this 3% figure come from?
joaogui1··on Proofs based on diagonalization help reveal the limits of algorithms
So you are listing all TMs + inputs and creating a machine which does the opposite (from a halting perspective), which is like when Cantor lists the reals and creates a new one that has the opposite digit from the other reals
joaogui1··on Falcon 180B
MOE models are trained all at once, they're not simply ensembling already trained models. Also data quality and quantity matter considerably, and how OpenAI gets their data is not public
joaogui1··on A Gentle Introduction to Liquid Types
APL started at Harvard (though it developed more in IBM) and even nowadays there's the ARRAY workshop co-located with one of the main PL conferences.

The thing is that while array programming is amazing for some specific problems it's not going to help you make sure that you don't have memory errors, race conditions, or wrong states in your program

joaogui1··on Transformers as Support Vector Machines
Programmers excel at it, you don't see that happening much with Physics, Chemistry, Math, Material Sciences etc Funnily enough it has also been observed before, by Alan Kay comparing programming to a pop culture "In the last 25 years or so, we actually got something like a pop culture" https://queue.acm.org/detail.cfm?id=1039523#:~:text=In%20the...
joaogui1··on Transformers as Support Vector Machines
The attention matrix is computed based on all tokens in the context, so it kind of functions non-parametrically (but over the batch instead of over the whole training dataset)
joaogui1··on Do Machine Learning Models Memorize or Generalize?
Sorry, I know what MDL and L2 regularization are, I would like the paper that connects them in the way you mentioned
joaogui1··on Do Machine Learning Models Memorize or Generalize?
That looks interesting, do you know what paper talks about the connection between MDL, regret, and weight decay?
joaogui1··on Show HN: Khoj – Chat offline with your second brain using Llama 2
I wonder if edge tpus (sold here https://coral.ai/products/) could help, though I guess the community would have to optimize for them instead of for standard hardware
joaogui1··on Attention Is Off By One
Just as a quick source to my claims:

1. The raft paper is titled "In Search of an Understandable Consensus Algorithm"

2. The abstract of this tutorial on Understanding Paxos https://www.ux.uis.no/~meling/papers/2013-paxostutorial-opod...

3. Lamport's own "Paxos made simple" https://lamport.azurewebsites.net/pubs/paxos-simple.pdf

← PreviousPage 2 of 5Next →