HNHacker News
TopNewBestAskShowJobs

ma2rten

7,364 karma · joined October 3, 2010

submissionscomments
ma2rten··on Ask HN: What was your "oh shit" moment with GenAI?
My personal "oh shit" moment was in 2015, when this paper came out: https://arxiv.org/abs/1506.05869

It showed me that a model trained only on movie subtitles data exhibited some (very primitive) reasoning. I have been working on Deep Learning and later LLMs ever since.

ma2rten··on Yann LeCun raises $1B to build AI that understands the physical world
Erm, ... OpenAI has hyped when it started and it took 6 years to take off. It's way to early to declare the SSI and Thinking Machines have failed.
ma2rten··on OpenAI declares 'code red' as Google catches up in AI race
Delaying doesn't necessarily mean they stop working on it. Also it might be a question of compute resource allocation as well.
ma2rten··on Show HN:emma019 Real-Time AI-Powered Texas Hold'em in Python and Flask
You can add Show HN to the title for your own projects. They will show up in the show tab.
ma2rten··on How Airbus took off
Europe is quite conservative, in the sense that they would not invest billions into an unproven venture. It makes sense that it would excel at an industry that requires putting safety above everything.
ma2rten··on BERT is just a single text diffusion step
It's actually true on many levels, if you think about is needed for generating syntactically and grammatically correct sentences, coherent text and working code.
ma2rten··on BERT is just a single text diffusion step
Interpretability research has found that Autoregressive LLMs also plan ahead what they are going to say.
ma2rten··on Boeing has started working on a 737 MAX replacement
Your use of the phrase makes no sense. It's the "no parking" that proofs the rule and not the exception.
ma2rten··on Are OpenAI and Anthropic losing money on inference?
You can also look at the price of opensource models on openrouter, which are a fraction of the cost of closed source models. This is a market that is heavily commoditized, so I would expect it reflect the true cost with a small margin.
ma2rten··on Curious about the training data of OpenAI's new GPT-OSS models? I was too
Presumably the model is trained in post-training to produce a response to a prompt, but not to reproduce the prompt itself. So if you prompt it with an empty prompt it's going to be out of distribution.
ma2rten··on MIT study explains why laws are written in an incomprehensible style
The study seemed not very convincing to me, at least the way it was described in the article. To summarize: they asked crowdworkers to write a law who used legalese, but not when writing news stories about it or when explaining the law. From that the researchers concluded that people use legalese to convey authority.

But what if people just imitated the writing style of existing laws, but not with the intention to make it authoritative but because that is what they understood their task to be?

ma2rten··on Ask HN: How does Alexa avoid interrupting itself when saying its own name?
This is the same problem as echo cancellation on calls. This is something that built into a lot of software and hardware.
ma2rten··on Maxtext: A simple, performant and scalable Jax LLM
t5x was used to train PaLM 1.
ma2rten··on Travelling with Tailscale
I have an upcoming trip to Europe, which I am quite excited about. I wanted to set up a Tailscale exit node to ensure that critical apps I depend on, such as banking portals continue working from outside the country.

I've never had an issue accessing banking portals from Europe.

ma2rten··on Apple cuts off Beeper Mini's access
Apples cares about the privacy and security of iPhones as a differentiator.
ma2rten··on Gemini AI
Noam.
ma2rten··on Gemini AI
No this is not correct. Arguably OpenAI invented LLMs with GPT3 and the preceding scaling laws paper. I worked on LAMDA, it came after GPT4 and was not as capable. Google did invent the transformer, but all the authors of the paper have left since.
ma2rten··on OpenAI is exploring making its own AI chips
Both Amazon and Google already do this, there are reports that Microsoft does as well.
ma2rten··on How Transformers Work
Yes, I think that is a reasonable way to think about it, in my opinion. However, with the language modeling objective it predicts the next token and because of the residual connections each intermediate layer is in the same space. So, maybe it would be more accurate to say that it is an increasingly accurate representation of the next token.
ma2rten··on How Transformers Work
Attention takes in all tokens in the sequence and outputs a new representation of the current token in context. Each layer of the transformer adds more context to the token.

I haven't read this explanation in detail and although they have some nice animations, I wouldn't go to FT to explain machine learning concepts. Here are two well known explanations that might be better:

http://jalammar.github.io/illustrated-transformer/

http://nlp.seas.harvard.edu/annotated-transformer/.

ma2rten··on GPT-4 can't reason
I didn't have time to read this, but it is a single author paper, the author is not affiliated with a research group, it is not peer reviewed, it was published on a preprint server that I have never heard of.

LLMs can definitely perform some kinds of reasoning. For example GSM8K is a dataset of grade school math problem requiring reasoning that LLMs are typically evaluated at. We talk about one method for this in our chain of thought paper [1]

[1] https://arxiv.org/abs/2201.11903

ma2rten··on Understanding DeepMind's sorting algorithm
How so?
ma2rten··on Understanding DeepMind's sorting algorithm
Someone just asked GPT-4 and got the same result as DeepMind did:

https://twitter.com/DimitrisPapail/status/166684395282416846...

ma2rten··on OpenAI's plans according to sama
That is only relevant for serving and not for inference, unless the model is too big to fit on a single host (typically 8 GPUs).
ma2rten··on What Neeva's quiet exit tells us about the future of AI startups
https://arxiv.org/abs/2305.15717
ma2rten··on A former pilot on why autonomous vehicles are so risky
It seems very ignorant to bet against AI given all the progress that have been made in that last year (!).
ma2rten··on IKEA bomb shelter furnishings catalog by Midjourney
https://jalammar.github.io/illustrated-stable-diffusion/
ma2rten··on ChatGPT is powered by contractors making $15 an hour
It's below minimum wage in California.
ma2rten··on Google releases Bard to a limited number of users in the US and UK
As someone working in this area, I think this is possible but unlikely. Most likely the model had access to the context, but "forgot" to use it.
ma2rten··on Fork of Facebook’s LLaMa model to run on CPU
It's still useful, but you need to know how to use it.
Page 1 of 34Next →