HNHacker News
TopNewBestAskShowJobs

luke14free

296 karma · joined December 4, 2010

Luca Giacomel building stuff @ https://plexagon.io
submissionscomments
luke14free··on Which LLMs fold under pressure? We made 6 LLMs argue 300 hard cases to find out
Hey HN,

We built a benchmark that tests how LLMs behave when they have to hold a difficult debate position against an adversary.

We took 6 frontier models, paired them in structured disputes (business conflicts, ethics dilemmas, property disputes, family disagreements), and forced them to argue opposing sides before a third LLM mediator. Each model gets a position to defend and a fixed number of turns. A separate judge panel scores the outcome.

The interesting part isn't who "wins" but rather what the disputes reveal about post-training behavior. Some models fold almost immediately, conceding points they shouldn't. Others hold firm on weak positions when a smarter move would be strategic compromise.

We ran this as a Swiss tournament (like chess) - 10 rounds, ~300 matches total, every case played twice with sides swapped to cancel position bias. Three independent frontier judge LLMs score each ruling, majority vote decides the outcome.

A couple of things we noticed: - models tuned hardest to be agreeable are the ones that lose most, they tend to concede points mid-argument even when holding a strong position - some models argue much better when they're on the "sympathetic" or "morally comfortable" side of a dispute than when they're assigned the harsher position. E.g., a model might crush it defending a tenant against eviction but argue poorly when it has to defend the landlord's right to evict

P.S. For every match read the full argument transcript.

luke14free··on I want everything local – Building my offline AI workspace
you might want to check out what we built -> https://inference.sh supports most major open source/weight models from wan 2.2 video, qwen image, flux, most llms, hunyan 3d etc.. works in a containerized way locally by allowing you to bring your own gpu as an engine (fully free) or allows you to rent remote gpu/pool from a common cloud in case you want to run more complex models. for each model we tried to add quantized/ggufs versions to even wan2.2/qwen image/gemma become possible to execute with as little as 8gb vram gpus. mcp support coming soon in our chat interface so it can access other apps from the ecosystem.
luke14free··on Iconhunt: Search 150k free and open source icons
search with a - instead of a space, I get more results (i.e. F1-car returns one icon, and so does thumbs-up)
luke14free··on Show HN: Mailbrew – Automated Email Digests from HN, RSS, Reddit, Twitter
Cool product, could be good for people that want to move from constantly disturbing notifications to daily digests. What is the roadmap for supporting more sources and for creating customized digests (i.e. create digests from custom sources that other people can follow)?
luke14free··on Oh shit, my weekend project turned into an App Store Best New App (2018)
There is a public ranking of apps ranked by downloads in the app store. If you know how many downloads correspond to what position in the chart (e.g. number 100 makes ~1000 downloads), you can fit a log regression to determine the number of downloads given the ranking in the category. I am oversimplifying because you also need to take into account seasonality and time trends (e.g. the app store is always expanding), but overall fits are very very precise (unless you get near the top positions, which are those for which they have more problems in estimating). Ah, and if you wonder where the data comes from, App Annie and Sensor Tower get data for downloads to fit their models directly from tens of thousands of developers that share it with them in exchange for free analytics,
luke14free··on Oh shit, my weekend project turned into an App Store Best New App (2018)
I work on models like those of Sensor Tower and App Annie use to estimate downloads. They are extremely precise, the cross-validated error is << 10% for apps that are this small, and in the US (it increases with much bigger apps in countries where little data is available).
luke14free··on Oh shit, my weekend project turned into an App Store Best New App (2018)
Just to be clear: the app makes ~20 downloads per day (checked on App Annie premium and Sensor Tower premium) except for the fact that you change the price from paid to free multiple times, causing multiple spikes in downloads generated by sites like https://www.148apps.com/ and https://appadvice.com/apps-gone-free which basically list all apps that have become free from paid.

So claiming `about 10,000 downloads each week` is not very correct.

luke14free··on Simple type checked objects in Python
Fair point. I am editing the entry name to reflect your suggestions. Static typed -> type checked at runtime