HNHacker News
TopNewBestAskShowJobs

ykhli

307 karma · joined March 11, 2021

https://github.com/ykhli/
submissionscomments
ykhli··on Show HN: TetrisBench – Gemini Flash reaches 66% win rate on Tetris against Opus
Thanks so much for the amazing feedback!!! Will update the app to incorporate these
ykhli··on Reverse engineering Lyft Bikes for fun (and profit?)
Most amazing tech blog I’ve read this week. What a great read!
ykhli··on Show HN: TetrisBench – Gemini Flash reaches 66% win rate on Tetris against Opus
my unvalidated theory is that this comes down to the coding model’s training objective: Tetris is fundamentally an optimization problem with delayed rewards. Some models seem to aggressively over-optimize toward near term wins (clearing lines quickly), which looks good early but leads to brittle states and catastrophic failures later. Others appear to learn more stable heuristics like board smoothness, height control, long-term survivability even if that sacrifices short-term score

That difference in objective bias shows up very clearly in Tetris, but is much harder to notice in typical coding benchmarks. Just a theory though based on reviewing game results and logs

ykhli··on Show HN: TetrisBench – Gemini Flash reaches 66% win rate on Tetris against Opus
oh that is super interesting. ty for the idea!
ykhli··on Show HN: TetrisBench – Gemini Flash reaches 66% win rate on Tetris against Opus
Wow this is incredible!!
ykhli··on Show HN: TetrisBench – Gemini Flash reaches 66% win rate on Tetris against Opus
answered this in a comment above! It's not turn or visual layout based since LLMs are not trained that way. The representation is a JSON structure, but LLMs plug in algorithms and keeps optimizing it as the game state evolves
ykhli··on Show HN: TetrisBench – Gemini Flash reaches 66% win rate on Tetris against Opus
Thanks for all the questions! More details on how this works:

- Each model starts with an initial optimization function for evaluating Tetris moves.

- As the game progresses, the model sees the current board state and updates its algorithm—adapting its strategy based on how the game is evolving.

- The model continuously refines its optimizer. It decides when it needs to re-evaluate and when it should implement the next optimization function

- The model generates updated code, executes it to score all placements, and picks the best move.

- The reason I reframed this problem to a coding problem is Tetris is an optimization game in nature. At first I did try asking LLMs where to place each piece at every turn but models are just terrible at visual reasoning. What LLMs great at though is coding.

ykhli··on Show HN: New.email – Building Emails with LLMs
Super interesting - I'd actually love to write a MCP server based on this as I find myself generate react email from cursor all the time.

Also is there a way to add styles / instruct on colors?

ykhli··on Show HN: MCP server that lets Cursor agent send Morse code through your light
This MCP server allows Cursor or Claude Desktop to control Philips Hue lights and send messages through them using Morse code. Have fun!
ykhli··on A Deep Dive into MCP and the Future of AI Tooling
How do MCPs actually work, where are the real use cases, and the challenges today.
ykhli··on OpenAI operator and Anthropic performed below average on chimp memory test
I benchmarked OpenAI operator and Anthropic's computer use on the human benchmark chimp test (https://humanbenchmark.com/tests/chimp).

Before I began the test, I thought the agents would be much better at this task than most humans -- after all they should have better, more stateful memory than us. The results are intriguing.

Here are the scores from 10 attempts: OpenAI operator: 5, 5, 6, 5, 5, 4, 6, 5, 5, 5 Anthropic computer use agent: 7, 9, 6 (rate limited), 12, 9, 7, 9, 11, 12, 6 (rate limited)

ykhli··on Show HN: Inngest 1.0 – Open-source durable workflows on every platform
Congrats! I used Inngest when I wrote a video processing pipeline here https://github.com/tigrisdata-community/multi-modal-starter-...

Amazing devEx. Thanks so much for all the work and enabling a local mode too

ykhli··on Open source Wikipedia search through vector db
Oh wow this is such an interesting case. I wonder if it's because of the embedding model used here.

If I search for Leo Tolstoy it works well, but "Amor towles" doesn't lead to the right page.

ykhli··on Show HN: A cartoon intro to how the attention mechanism works
THANK YOU!!! this is amazing. Will update the cartoons over next week.

Also thanks so much for taking the time to write these down!! Can't appreciate it enough

ykhli··on Show HN: A cartoon intro to how the attention mechanism works
hey! thanks so much for the feedback. I'd actually love to keep updating / iterating these cartoons so they are more approachable. If you have time, I'd love to hear more on which pages are confusing & how I could have explained it better!

I _tried_ to give a definition to embeddings on page 11, but maybe that's not the most intuitive? Lmk! feel free to DM

ykhli··on Show HN: Multi Modal Starter kit - roast a movie with AI
Hi HN! Recently we made a starter kit for better learnings on how to get started building AI apps with multi modal models. It's been fun to build, but we also discovered there are many things we needed to take care of: caching, long video processing pipelines, model evaluation, etc

I wrote more about the technical details here. Feel free to try it out and open PRs! https://twitter.com/stuffyokodraws/status/177558959544044376...

ykhli··on Building a fair multi-tenant queuing system
Great job abstracting away so much complexity!

> each step is a code-level transaction backed by its own job in the queue. If the step fails, it retries automatically. Any data returned from the step is automatically captured into the function run's state and injected on each step call.

This is one thing I've seen so many companies spending tons of time implementing themselves, and happens _everywhere_ -- no code apps, finance software, hospitals, anything that deals with ordering system...the list goes on.

Glad I no longer need to write this from scratch!

ykhli··on Show HN: JS Local-only AI Apps starter kit: cost $0 to run and test locally
These are great suggestions!! Really appreciate it. Let me open a few PRs and add you to review :D