I was thinking the cautiousness was intentionally bred into LLMs, but it could very well be an artifact of the training data, too.
1,808 karma · joined August 24, 2010
leftium.com
Leftium: The Element of Creativity!
I was thinking the cautiousness was intentionally bred into LLMs, but it could very well be an artifact of the training data, too.
When I use coding agents to help with prose, they default to bending over backwards to avoid any absolutes or otherwise risky text that might offend someone. (For example, an LLM would add a lot of hedging to my previous sentence to clarify I haven't tried all LLMs, soften the language, etc...)
While I understand the intent, it often makes the text verbose and awkward. So usually the suggested text is safer, but at the same time harder to read.
Another example is always planning/implementing a path for backwards compatibility/migration even when the project is still a prototype where the only user is myself.
Conclusion: Elixir was the best (had the highest problem solve rate).
Reasons (Theo's interpretation):
- code collocation, where documentation is integrated directly within the source code
- design philosophy of a language (readability, clear idioms, and strict expectation management)
---
Sources from video:
- https://github.com/Tencent-Hunyuan/AutoCodeBenchmark/tree/ma...
- https://martinalderson.com/posts/which-programming-languages...
- https://simonwillison.net/2026/Jan/19/nanolang/#atom-everyth...
Some things I noticed:
- The comments are not in HN order; "vibecoded prediction market–style" is at the top in OrangeCrumbs, but near the bottom on actual HN; actual top comment ("I'm struggling to understand why I'd ever use this instead") is buried in the middle on OrangeCrumbs.
- Other comment sort options are interesting
- It is difficult to read a deep thread. Requires clicking once for each level. (I just found keyboard shortcuts that kind of help, like [)
- The left and right sidebars are kind of distracting. So I find the mobile view easier to read. I do think sidebar can be better use of wide screens; perhaps try muting the sidebars and/or allow them to be toggled? I find the right sidebar more useful than the left one.
- Keyboard navigation: conflicts with Vimium. The only "natural" keyboard shortcuts I could discover were arrow keys + SPACE. (`?` was also natural, but unfortunately Vimium had priority.) Personally, I prefer skimming HN threads with the mouse scroll wheel and a few occasional clicks. (It seems it would require a lot of keyboard shortcut keys for a similar reading experience)
---
Here's my take on rendering comment-heavy HN threads: https://hn.leftium.com/i/48736605
- Render comments at 3 levels of detail: Full (top level), one-line excerpt (2nd level), grouped into single color strip (everything else)
- Make threads expandable at three levels of detail:
1. Expand all direct replies only
2. Ungroup: expand collapsed comments as one-line excerpts
3. Expand entire comment sub-tree.
This is the way I found myself wanting to read long threads: skim the top level or two for threads I found interesting. Keep recursively skimming like this for the interesting comment sub-threads.
Some other features I plan to add:
1. Expand/highlight all comments by specific user(s). Currently OP is highlighted. (But not auto-expanded, yet.)
2. Filter for certain words in thread: typing a query automatically expands matching comments; parent comments may be expanded as one-line summaries for context.
So my reading process is scroll with the mouse wheel until I find an interesting comment thread; occasionally click to toggle those threads between collapsed and higher levels of detail.
I couldn't tell you exactly what I was working on one/two years ago. But I could tell you after seeing my search history.
That's how I came to the realization search history is a zero-effort journal. I started making a feature that is similar to search history, but focuses more on the journal aspect. You can add explicit journal entries. Search/browsing are focused on surfacing interesting patterns/entries vs. finding specific entries.
Very early WIP: https://whiz.leftium.com/journal
Lots of ideas seem cool. Far fewer ideas are are easy to sell. And often the ideas that are easy to sell are not the same as the cool ideas.
That's why I think building/marketing ideas "backwards" can be more effective. Certainly less risky. First identify the "starving crowd." Then build a hamburger for them[1].
If you identify a painful problem, people will ask how to pay for it even before you start development. I've witnessed this in person. If you describe the problem better than the prospects themselves, they will just assume you have the solution.
---
[1]: "Starving crowd" story by Gary Halbert, one of the most successful marketers: https://www.thegaryhalbertletter.com/newsletters/direct_mark...
Theoretically, if you could specify a seed and the exact version of the model the output should always be the same. I wonder if this is possible with any open-weight models today?
---
On a more practical level, scripts (small programs) are deterministic so having the coding agent write (and possibly reuse) scripts might help.
As long as usage is not excessive, feel free to use the deployed version, as well.
- classless CSS library: https://leftium.github.io/nimble.css
- HN client: https://hn.leftium.com
- local realtime streaming transcription prototype: https://rift-transcription.vercel.app
---
These projects were started without AI, but heavily augmented with coding agents:
- https://weather-sense.leftium.com
- console.log replacement: https://github.com/Leftium/gg
- Thin layer over Google forms/sheets: https://veneer.leftium.com
The monitor with the specs I wanted was $201[2]
Old-school graphics in modern TS.
Several years ago, it was not possible to blit an entire screen of random pixels to the screen at a decent frame rate without something like shaders.
Even though the screen is now even higher resolution, the CPU can now blast 2560x1440 random pixels to the screen at 90 FPS. Must be advancements in hardware and/or JS runtime. (The bottleneck seems to be generating the random numbers...)
I figured out how to make my TV static effect look more realistic:
- Mostly: TV "pixels" had wide aspect ratios[1]
- Larger "grains" (see info in corner)
- Also added subtle CRT scan line effect. ('C' to toggle)
- Looks different when animated (click to toggle pause; probably should emulate 60FPS).
---
Started revisiting this rabbit hole while thinking about programming prompts from the new Recurse Center application[2]. They suggest about six different prompts; I figured out how to combine all the prompts together.
[1]: https://github.com/Leftium/fx/blob/33405b25dc7caeb48e6c563a3...
The defaults are tuned to aesthetics of my personal logo, but it's quite configurable. (You can even copy your own SVG into the icon input)
Example logos:
- https://leftium.github.io/nimble.css
Video of the editor: https://www.davepagurek.com/content/images/2025/12/Component...
By employing you, the US company must comply with all Australian tax and labor laws (in addition to the US laws). This is a huge burden. (Like the company must calculate and report how much revenue was generated through your work and pay Australian corporate income tax.)
Your best chance is to apply to US companies that already operate in Australia: they will have the necessary legal/HR infra set up. (For example Google probably has an office in Australia.)
---
Another method may be to work as a contractor through a service like https://www.toptal.com (or even on your own if you can find contracts).
I made a classless CSS library, then migrated most of my projects from PicoCSS.
I also made a quick logo generator: https://logo.leftium.com/logo
This is not new. Many Korean mobile plans actually offer even higher unlimited throttled speeds (up to 10 Mbps!)
- You can filter plans by the unlimited throttled speed on this site. The plans are usually titled by `{data amount} + {throttled speed}`: https://www.moyoplan.com/plans
- Even if not throttled, I think data overage charges were capped at about $13 (20K KRW)
So perhaps unlimited 400 kbps will become standard: i.e. no plans will ever charge data overage fees?
---
The linked statement didn't seem to specifically mention the 400 kbps thing at all.
- This was pointed out in a thread: https://hw.leftium.com/#/item/47504047
There was no [OP] label when I first reloaded this page, and now after replying, I am marked as the [OP].
edit: it seems the [OP] relies on a URL hash with the ID of the OP. However, this doesn't work for me because I don't navigate to HN posts from the HN site.
I usually come from https://hn.leftium.com. (Or a page like https://hw.leftium.com/#/item/47694036)
It's MIT-licensed open source.
I've been using a fork (also MIT): https://github.com/alexferrari88/refined-hacker-news
One thing I miss is the orange mark identifying the OP of a post.
If the resellers down the chain were purchasing your shoes for less then your cost, would you still be happy?
Say the resellers were abusing an 80% discount coupon. Anthropic is basically closing a 95% discount coupon that was being abused.
If OpenClaw users were paying the API rate, your strategy could make more sense.
The reason Anthropic is subsidizing inference is because they are trying to capture users (marketshare). However the acquisition costs for a single OpenClaw user is much higher. And OpenClaw users are less likely to convert into profitable users later.
---
In addition, there is a supply bottleneck. Currently Anthropic is having trouble servicing all the demand due to a shortage of GPU's. And in the current market it is impossible to get more GPU's (or at least prohibitively expensive).
Anthropic (and all other AI companies) also need GPU's to stay competitive: GPU's are needed to train better models. So you could view it as Anthropic has decided instead of subsidizing nonprofitable OpenClaw users, it is better to repurpose that GPU for internal R&D instead.
The colors + space simply help you understand the numbers better.
(Weather forecast precision is artificial "because weather forecasts fundamentally have very high uncertainty and error bands"[1])
- You can see the weekly high/low temperature trends by scanning down vertically along the left.
- Redder color means warmer; bluer means cooler.
- The gradient is constant for all data plots, so you can visually compare the temperature across days and hours.
- The gradient block for each day goes from the high to the low temp for that day.
- Even the hourly temperature plot line is calibrated to the same gradient.
---
The sky background gradient is slightly superfluous, but it's very subtle and meant to emulate (a more vibrant) version of the actual sky.
For anyone who wants more gradients: there's a setting here: https://weather-sense.leftium.com/wmo-codes
I disabled those by default because they were distracting and didn't serve a purpose.
I just checked, and the responsive layout seems to render correctly on Android Firefox/Chrome and iOS Safari.
You can even save WeatherSense to your home screen as a simple progressive web app.
- You can also tap any unit to toggle.
- But the main point of WeatherSense is to transcend units ^^
Explained in more detail: https://youtu.be/1WFgIjAvMDw?t=882
TL;DR:
- Cursor Ultra
- OpenAI Codex
- OpenCode Black (currently not accepting new subs)