HNHacker News
TopNewBestAskShowJobs

janalsncm

12,623 karma · joined July 18, 2022

submissionscomments
janalsncm··on With most information hidden, the game Stratego had stumped AI until now
You are confusing the inability to compute a full game tree with not knowing anything at all. In fact there are many positions in chess where we can compute the full game tree. Forced mates, and tablebases of positions with 7 pieces or less. And even if we can’t compute the full tree, errors get smaller with depth.

> There isn't any reason to think AI have more or less trouble with hidden information games.

How about the fact that a child can beat the best rock paper scissors player in the world in a game, but no human can beat the best chess engine? Same thing with poker, a novice could get lucky and win a hand against the best poker player.

janalsncm··on On social reality in China
The US is a big place. There isn’t one “American culture”. If someone in one part of the country says we should get drinks sometime they’re just saying a nice thing. In another part of the country, that means “this weekend”.
janalsncm··on On social reality in China
The purpose of the author writing this seems to be improving the effectiveness of Rationalists in persuading the Chinese on questions of AI safety.

Based on my experience, these kinds of cultural generalizations are interesting to talk about but as with many populations there is much more variance between people within the group than between groups. For example young urban Chinese do not put Mao posters up in their apartments. For older rural Chinese it’s almost universal. Therefore, it is much more effective to get to know individuals than to write to no one in particular.

Additionally, the question of how Chinese society looks in general is less important than how the decision makers think. You don’t need to convince average Chinese people. You need to convince the equivalent of a VP at Google.

Based on those two things, and especially because as far as I can tell the Rationalists have framed the entire thing as a sort of prisoner’s dilemma, the only way out is to build trust. And in my opinion the best way to do that is for Americans to meet Chinese people and for Chinese to meet Americans.

janalsncm··on With most information hidden, the game Stratego had stumped AI until now
> The algorithm also learned far faster—it played about 34 times fewer games than DeepNash, and still ended up much stronger.

Imo, this is the critical piece and what makes the AI work at all.

With hidden information games, the best move depends on information you don’t have. So a move could be good or bad, it just depends on something that’s impossible to know.

You’d like to search ahead, meaning “if I do this they will do that” but that’s impossible since you don’t even know what the opponent can do because you don’t know their hidden state.

If the possible hidden states are randomly distributed, you are screwed. It’s just like rock paper scissors: there’s no best move if your opponent is unpredictable.

However if you can quickly learn to predict their moves, it becomes possible to make informed decisions about what to do.

janalsncm··on Vote on which of Hacker News' challenges for AI have been met
A better test would be using an image to ascii converter tool to rule out bad ascii drawing from Opus.
janalsncm··on Singapore govt dating app uses Gale-Shapley stable marriage algorithm
Tinder doesn’t know if you stopped using it because you found a partner they recommended, found a partner in some other way, got frustrated and gave up trying, or got frustrated and tried a different app.

Also they won’t know who you started dating. It’s most likely not the last person you matched with or messaged. So any training data will probably be incredibly noisy, because romantic compatibility is not something that happens in the app.

janalsncm··on Clef: Open-source decision models, and new RL fine-tuning platform
The interesting part is also the easy part. The model and architecture are not hard for an experienced machine learning engineer to build.

The hard part is the data and evaluation. Sure, it’s not that hard to build a fast model with good predictive power. But fast at doing what? You probably don’t care about classifying whether a hotdog is a sandwich (which is the Jev demo).

janalsncm··on Is sandboxing sufficient to contain rogue agents?
What did you think of the author’s concerns on the thing you are suggesting?
janalsncm··on Singapore govt dating app uses Gale-Shapley stable marriage algorithm
Marriage records are often public. But usually two people don’t stop using a dating app and get married the next day. They will date for months or years after. And that info is not so easy to get.

As far as data brokers I would imagine it’s possible to get info on an individual if you want to. You can even determine someone’s location based on data brokers info. But dating apps will want info on millions of people, which will be messy and hard to do.

janalsncm··on Singapore govt dating app uses Gale-Shapley stable marriage algorithm
Right, so if your hobby is overthrowing the government it’s probably better to put that on your Tinder profile only.
janalsncm··on Singapore govt dating app uses Gale-Shapley stable marriage algorithm
In principle this can be a lot better than regular dating apps.

Tinder doesn’t know if they were successful or not. They just know you stopped opening the app. In fact they are incentivized not to make matches that lead to marriage, because then you will stop paying.

The government knows whether you got married and stayed married, and given the social cost of divorce, they are pretty incentivized to keep people married.

janalsncm··on PSSA: A non-transformer language model written from scratch in Rust
Yes, and it’s not even research anymore! Qwen uses linear attention: https://sebastianraschka.com/llms-from-scratch/ch04/08_delta...
janalsncm··on GLM-5.3 and the spread of advanced cyber capabilities
My point is that $215 per author might not be a lot for a company that thinks it’s worth $2T but it is a lot for normal American startups. So it really isn’t a level playing field.
janalsncm··on PSSA: A non-transformer language model written from scratch in Rust
Most PyTorch tensor operations are cython not python. So imo rewriting in rust is not going to have an enormous speed up. If that really was the concern we should see a throughput comparison vs PyTorch or something.
janalsncm··on PSSA: A non-transformer language model written from scratch in Rust
Rust has its place but python is default for this kind of thing. By doing it in rust, they have now changed two things: the implementation of the transformer, and this new model.
janalsncm··on PSSA: A non-transformer language model written from scratch in Rust
OP, you should not have written this in Rust. It should be in PyTorch, which is by far the most popular. We can’t tell if this architecture is good or whether there is a problem in your implementation.

You can test the whole thing for free on a GPU with Google Colab. Test both the transformer and your new architecture on a larger dataset. Something that maxes out the GPU for an hour each run.

Also, the readme mentions keeping the same optimizer schedule which sounds nice at first but they are completely different architectures. The loss is high on the transformer, did you try raising the learning rate on it?

In general I’m interested in parameter efficient architectures. I don’t think transformers are optimal, and indeed many improvements have been made to vanilla transformers. But if you have an idea for something better you need to show it.

janalsncm··on GLM-5.3 and the spread of advanced cyber capabilities
https://www.reuters.com/world/us-judge-approves-anthropics-1...

That was just one settlement but precedent is clear. If you do what Anthropic and OpenAI did, expect to be in court. This is one reason why you don’t see labs popping up out of nowhere in the US.

Also if you distill from Anthropic and OpenAI, expect to be in court. Whether you think distillation is fair game or not, the US court system is not cheap. But it turns out that Z.ai, Minimax, Moonshot, Xiaomi, Deepseek, Alibaba etc don’t need to worry about that.

janalsncm··on GLM-5.3 and the spread of advanced cyber capabilities
Maybe it’s worth asking how much we should regulate and not regulate in order to compete with Chinese models.

For instance, one regulation which really puts American AI companies at a disadvantage is IP law. It shouldn’t be a surprise that most of the best of the text-to-video models are Chinese.

Similarly, the legal grey area around model distillation gives Chinese labs a major advantage. This one I feel better about relaxing.

https://www.goodreads.com/quotes/7515521-william-roper-so-no...

janalsncm··on Jeeves. Reasoning improves Jev-like decision models
I can’t see your gist but spam classification is a textbook example of something you shouldn’t measure with accuracy. If 95% of your samples are not spam you can get 95% accuracy by always guessing not spam.

You should use precision (when your model says “spam” how often is it spam?), recall (how many of the spam emails did it catch), or f1 (balanced between those two).

janalsncm··on MicroLLM Lab – Try 7 tiny LLM's in the browser
I should preface this by saying it is a cool demo. I love SLMs and I think they will become even more popular in the future as capability per byte improves and hardware improves to support my bytes.

So I think this post performed well in spite of the annoying parts, not because of it.

> I didn't want to remove the "mostly useless" information since clearly it was useful for a lot of people

The information was not “clearly” useful. It had obviously incorrect information that no one even noticed. In my opinion that is strong evidence for the opposite conclusion, that people ignored it because it was noise.

Why would people ignore this information? Demos are a show, don’t tell thing. For example, you don’t have to tell people that the latency is low. They should be able to see it from the demo.

janalsncm··on Nvidia wants to put a watchdog chip next to every AI agent
On the very narrow question of whether existing laws are sufficient to punish the kind of bad behavior that OpenAI has already done, is that really a legal consensus?
janalsncm··on MicroLLM Lab – Try 7 tiny LLM's in the browser
It isn’t about how small the text is or how easy it is to skip. Humans are not transformers with multiheaded anttention. Human brains do not process one million tokens in parallel. Humans have caloric constraints and will be annoyed if you ask them to process irrelevant information.

Therefore, when showing something to humans, the first thing they see should be the first thing you want them to see. And since this is a demo, the first thing they see should be the demo, with sane defaults, above the fold.

Also, it is likely that no one (including the author) has read this “fine print” because it includes incorrect information. GPT-4 isn’t a frontier model anymore.

janalsncm··on Does Reddit have an astroturfing problem? What the data suggests
This is why you need to write your own content.

> If the tail shows up in buying threads only because it posts a lot, random reassignment says it should write about 7.9% of the brand mentions there, and almost always between 6.3% and 10.1%. It wrote 11.3%.

Is it just me or is this just a block of painfully hard to parse text? This is core to the author’s analysis.

Rather than pasting the analysis from Claude, the author should just explain it themselves. Just the core of their analysis. It doesn’t have to be 2000 words. Put an AI;DR at the top and give us the core of your argument. I would take a paragraph from a human any day over an essay from Claude.

janalsncm··on Turning GLM-5.3-Flash into a Jev-like decision model
> it may raise some questions about the necessity of Jev's architecture

When I hear “architecture” I am thinking number of parameters and latency.

When I hear “accuracy” I think training recipe, data, and (later) number of parameters.

So when you say that Jev’s architecture may not be necessary, the evidence I expect to see is comparable quality at comparable latency. Not equal quality at 2x latency and 4x the cost.

janalsncm··on Turning GLM-5.3-Flash into a Jev-like decision model
Right, so it is double the latency and will no longer feel real-time to the end user.
janalsncm··on Turning GLM-5.3-Flash into a Jev-like decision model
If you are using an autoregressive decoder (which glm is) it is not “jev-like”. You lose all of the speed advantages that Jev has.
janalsncm··on Ollaya – Ollama for open-source, Jev-style decision models
I think GP has a point, though. BERT models have existed for a while but OpenAI made classification via LLM convenient and accessible for regular developers.

People didn’t know they wanted classifiers until OpenAI gave them a taste.

janalsncm··on U.S. appeals court upholds designation of Anthropic as supply chain risk
Alternative explanation:

1) The designation has nothing to do with hacking, either from Anthropic or otherwise. The dispute is about Anthropic barring the DoD from using Claude to develop autonomous weapons and conduct domestic surveillance.

2) The DoD does not care about any of the hacking OpenAI does as long as the DoD can do what they want with ChatGPT.

3) Historically, the US government has been interested in hacking other nation states, friendly and not. So OpenAI being able to hack Australian systems is a selling point.

janalsncm··on What happens when you analyze your favorite college football team like the CIA?
Ok, so in an ML system you can calculate things like cross entropy loss, which penalizes your model for making confident, inaccurate predictions.

It doesn’t care whether your system is made of LLMs, decision trees, or bananas.

However, the problem is that unlike something like a neural net, you can’t exactly use backprop to improve.

On the flip side if the quality of the prediction doesn’t matter, I might as well have Claude spin up something shiny that does the same thing faster, cheaper, and with 10x the magic sparkles.

janalsncm··on FTC chair suggests AI developers should be liable for conduct of agents
Going beyond the headline there are three things in this article:

1) FTC chairman says AI agents will be treated just like any other tool (“resist anthropomorphizing”)

2) Studying enforcement of rules around personalized (“surveillance”) pricing, specifically rideshare, airplanes and food delivery

3) Public comment period open now on new rule to punish companies like Meta and Google that advertise scams.

The last one seems closest to happening but it’s still just a public comment period so it might die.

Page 1 of 34Next →