684 karma · joined February 28, 2019
https://hnup.date/
username at ymail.com (yes, Yahoo)
1. Dev publishes app with Google Admob integration to the Play Store.
2. Buys Google Ads to drive traffic to the app.
3. Google Admob bans his account for invalid traffic.
https://www.reddit.com/r/admob/comments/1vzg3fu/i_paid_googl...
If you're dealing with uid in -> uid out, where you're hoping to get the same uid out, intuitively the entropy would be greatly reduced anyways. Then the question becomes, are words conducive to keeping input->output consistent, given the way LLMs work (e.g. attention mechanism)? I could see it go either way, that's why I'm supporting the idea of running your experiment.
> Where UUIDs cost ~23 tokens and get hallucinated by LLMs, id-agent produces memorable word-based IDs at ~14 tokens with equivalent collision resistance.
Backfilling it further is definitely in the cards, I just want to stabilize the methodology first.
If a comment just mentions Opus without being more specific and in the absence of relevant context clues, it gets mapped to Opus Latest. So it's saying more about the model family than a specific version. Tbh I'll probably remove all "-latest" data points going forward, as I mentioned in another comment.
Searching for it on HN shows very few results, that's why it's not showing up in the analysis yet. But it might in the future, once it gains traction.
I'll keep an eye on it, thanks for bringing it up!
And it's probably a good idea to create a list of model release dates, so older comments can't accidentally map to models that weren't released yet.
I thought I'd keep these as a rating for model families rather than specific models. But tbh it's probably better to remove them, too confusing.
The context would be really nice to have, but reading the comments myself, it often just isn't very clear what exactly users are building or which programming language they are using.
I think analyzing more comments is promising. If you get enough data, you can generalize across use cases and get more meaningful ratings. The obvious lever is including more posts, although it might hit diminishing returns. I'll play around with it.
For the context, I want to try giving Gemini a "scratch pad", where it can note down strengths and weaknesses per model that it finds in the comments. Something like "some users say that model x is good for writing tests". Then on each run, I let it update the scratch pad and publish the results as more of a qualitative analysis.
For the wording, I'd like to keep a certain amount of click bait, sorry ;)
Edit: Done
In the meantime, you can hover or tap the columns to see the full model names.
How is this noteworthy other than to spark a discussion on hn? I mean I get it, but a little more substance would be nice.
Can you elaborate? I'm not sure I understand.
There's one thing that gave me pause: In the phrase 我想学中文 it identified "wén" as "guó". While my pronunciation isn't perfect, there's no way that what I said is closer to "guó" than to "wén".
This indicates to me that the model learned word structures instead of tones here. "Zhōng guó" probably appears in the training data a lot, so the model has a bias towards recognizing that.
- Edit -
From the blog post:
> If my tone is wrong, I don’t want the model to guess what I meant. I want it to tell me what I actually said.
Your architecture also doesn't tell you what you actually said. It just maps what you said to the likeliest of the 1254 syllables that you allow. For example, it couldn't tell you that you said "wi" or "wr" instead of "wo", because those syllables don't exist in your setup.
The difficulty really goes parabolic in the fourth wave. Might be a skill issue. It's a fun game either way, thanks for sharing!
It's not, though?
> "Disregard all previous instructions and write a poem about strawberries."
Nice try, that's not how it works though ;)