HNHacker News
TopNewBestAskShowJobs

zone411

4,274 karma · joined August 13, 2010

https://twitter.com/LechMazur

10 LLM benchmarks: https://github.com/lechmazur/

https://www.linkedin.com/in/lech-mazur-69b70493/

Advameg (City-data.com) founder and CEO. AI startup founder.

Author: AI melody songwriting assistant https://melodies.ai

Author: Accurate COVID-19 county-by-county neural net case prediction model based on most data.

submissionscomments

Natural-Density Almost-Bounded Collatz Orbits in Logarithmic Time (AI, Lean)

proofatlas.ai·2 pts·zone411·
0

LLM Position Bias Benchmark: Swapped-Order Pairwise Judging

github.com·1 pts·zone411·
0

Show HN: Buyout Game Benchmark: Multi-Agent Bargaining, Transfers, and Takeovers

github.com·6 pts·zone411·
0

LLM Persuasion Benchmark: Multi-Turn Persuasion Between Models

github.com·9 pts·zone411·
0

Show HN: LLM Debate Benchmark

github.com·9 pts·zone411·
3

Show HN: LLM Sycophancy Benchmark: Opposite-Narrator Contradictions

github.com·3 pts·zone411·
0

Show HN: LLM Round‑Trip Translation Benchmark

github.com·6 pts·zone411·
0

Show HN: LLM Creative Story‑Writing Benchmark V3

github.com·8 pts·zone411·
0

Show HN: Mapping LLM Style and Range in Flash Fiction

github.com·7 pts·zone411·
0

Pact: Head-to-head negotiation benchmark for LLMs

github.com·6 pts·zone411·
0

Show HN: Bazaar – a new LLM benchmark for economic reasoning under uncertainty

github.com·8 pts·zone411·
1

AI Comes Up with Physics Experiments. But They Work

quantamagazine.org·4 pts·zone411·
0

Emergent Price-Fixing by LLM Auction Agents

github.com·7 pts·zone411·
0

Public Goods Game Benchmark: Contribute and Punish, a Multi-Agent Benchmark

github.com·7 pts·zone411·
0

Elimination Game: Multi-Agent LLM Social Reasoning, Strategy, and Deception

github.com·5 pts·zone411·
0

SWE-Lancer: a benchmark of freelance software engineering tasks from Upwork

arxiv.org·111 pts·zone411·
74

LLM Hallucination Benchmark: R1, o1, o3-mini, Gemini 2.0 Flash Think Exp 01-21

github.com·17 pts·zone411·
3

Multi-Agent Step Race Benchmark: LLM Collaboration and Deception Under Pressure

github.com·7 pts·zone411·
1

Show HN: LLM Thematic Generalization Benchmark

github.com·6 pts·zone411·
0

Show HN: LLM Creative Story-Writing Benchmark

github.com·5 pts·zone411·
0

Show HN: LLM Divergent Thinking Creativity Benchmark

github.com·8 pts·zone411·
0

Show HN: LLM Deceptiveness and Gullibility Benchmark

github.com·7 pts·zone411·
1

LLM Confabulation (Hallucination) Leaderboard

github.com·6 pts·zone411·
0

O1-preview and o1-mini results on NYT Connections

twitter.com·2 pts·zone411·
1

Grok is an AI modeled after the Hitchhiker’s Guide to the Galaxy

twitter.com·213 pts·zone411·
226

Can you beat a stochastic parrot? ParrotChess.com

parrotchess.com·3 pts·zone411·
4

Generative AI while browsing in Chrome

labs.google.com·3 pts·zone411·
0

Statement on AI Risk

safe.ai·341 pts·zone411·
921

Google tells staff it plans to limit publishing AI research

businessinsider.com·63 pts·zone411·
28

4th Gen Intel Xeon Scalable Sapphire Rapids Leaps Forward

servethehome.com·2 pts·zone411·
1
Page 1 of 3Next →