24 karma · joined May 30, 2019
"You have great vision and skill," she said. "I would like my son to become like you."
"I'll keep that in mind," Son replied.
Goose matters aside, this was my favorite part.
Yes, if the premise was true but it’s not.
I clicked that link first even though it’s listed second bc I wanted to see the prompts. I didn’t expect the level of detail or mapping to each commit. It is rad!
That being said the landing page is soooo obviously “vibe coded” (read: AI generated).
It has that design style that Claude likes to ~ab~use. & if I’m being honest, had I clicked on the website link first, I would never have gotten to the demo bc I would’ve just dismissed it as AI slop.
To answer your question: most ppl enjoy routine and the satisfaction of checking a box, no matter what that box is.
The next evolution of multi agent orchestration / “advisor strategy” [1] will be branded in humanized language like this. Less about tokens and capability, more about wisdom and knowledge to guide a “younger” (less capable) model. Somebody will make a billion dollars by selling it as paired programming for LLMs.
[1] https://platform.claude.com/docs/en/agents-and-tools/tool-us...
What are CF & 5 eyes?
1) can you elaborate on how you're using audio fingerprinting for it? if the podcast host is the one reading the ads, does it still catch it?
2) how are you able to offer it for free? i imagine there has to be some cost bc some type of algo is involved (hosting, LLM processing, API calls, etc.) but i might be wrong.
3) i downloaded it and tried a podcast and the ads were not skipped. happy to share the podcast and episode for diagnosis.
4) you might be interested in downloading PodSkip. i just tried out the beta via testflight this week. it behaves similarly to your app
I do agree with you about the deluding part though. I was (as a user) all for hyper-personalization of ads on all platforms when I worked in ads. Since I’m not longer working in ads, I’m more skeptical and value privacy a lot more.
Honestly, the core problem is that we can’t trust the platforms selling the ads.
Hmmm.
My conclusion: maybe they did use Gemini to brainstorm.
Convo link: https://g.co/gemini/share/26a0549dc579
1/ CursorBench is so opaque [1] that it makes it hard to trust. Not to mention the v3.1 eval is a newer iteration and there's no insight into the tasks or if the model was just tuned to max it out. Composer 2 previously scored between 60-65% on the previous benchmark eval [2] but scores between 50-55% on CB v3.1[3].
2/ I've experienced Composer 2's performance and it leaves much to be desired as a daily driver for a knowledge worker. but KWs are obviously not the target users and I can see how it's cost-efficient for executing on clearly-defined, discrete coding tasks. Obviously that's their value proposition and they're figuring out how to communicate it well to the target customer. It just doesn't feel like CursorBench is that.
[1] https://cursor.com/blog/cursorbench#building-cursorbench
[2] https://cursor.com/blog/composer-2-technical-report#performa...