HNHacker News
TopNewBestAskShowJobs

time0ut

2,072 karma · joined April 24, 2015

submissionscomments
time0ut··on If we do not stop to help each other, what do we become?
The early days of stack overflow were amazing, but 2013 was a long time ago. A few bad experiences when trying to help people was enough of a turn off to stop me from logging a decade ago. I do agree that we lost something and LLMs are a poor replacement.
time0ut··on Jev Can't Be Calibrated
I have been running a series of experiments on Jev since its release targeted at understanding it, seeing how it handles real use cases I have, and maybe figure out what it is inside.

Some of my tests do point towards what this post says. I was not successful in getting it's score to align with an existing rubric I had. It 'worked' but it was off and compressed from where I wanted it to be. Not a bad starting point, but I couldn't get it to move to where I intended the rubric to be. It wasn't the most robust test and I didn't spend a lot of time trying, but it wasn't just instantly magical.

However, it does seem genuinely useful just by being fast and cheap and good enough, so I am still a bit hyped.

time0ut··on Strands Harness
The sales pressure from them on their agent core stuff has been really shocking over the last six months. Never seen anything like it.
time0ut··on OpenAI is well positioned to fast-follow Jev
Yes! I really hope they release some papers on their techniques. I am very curious.

I ran it through MMLU a few days ago and it scored ~90% so seems to have a lot of general world knowledge trained in. Makes me think your speculation is right. I have some credits left, might try and think of an experiment. I saw a gist where someone was asking it which model it was and it was picking qwen a lot, but who knows...

Anyway, thank you for the interesting discussion!

time0ut··on OpenAI is well positioned to fast-follow Jev
Yes, agreed. I was speaking in general, of course. This particular topic is of interest to me, so thinking of the edge cases and confounds vs Jev.

In your example, I would expect an LLM to do fine and if you have access to the raw logits you can measure whether or not it was confused and assign a confidence to the answer it gave.

I do think that Jev handles more than this though and, in my early testing, does things that are not easily accomplished with guided decoding techniques.

time0ut··on OpenAI is well positioned to fast-follow Jev
What I mean is that, in general, constrained decoding can push model output off into less probable regimes. This is well studied; see for example https://arxiv.org/pdf/2606.21619. The mask may only retain very improbable logits. In pathological cases, the constrained output may be little better than noise filtered through the constraint. When using existing structured output APIs, it may not be possible to even know.
time0ut··on OpenAI is well positioned to fast-follow Jev
Grammars do risk pushing models off distribution in a way that impacts their output quality in a way Jev allegedly does not suffer from. Additionally, Jev's ability to answer questions independently is also exciting. Using an LLM to answer multiple questions in one generation has the property of earlier answers influencing later ones. TBD how many of TypeSafe's claims stand up, but my testing so far is promising. I hope they author some papers on their methods as well, but that might destroy their moat.
time0ut··on Bend 2 and the Vibe-Coding Trap
I am a total outsider when it comes to this topic, but it reads like the author is doing the thing they are accusing the Bend 2 guy of. Good juicy reading while I sip my coffee!
time0ut··on Gemini 3.8 Live and 3.8 Live Extended Thinking
Playing a legal game of chess without using a guided decoding technique is a massive achievement. Ref https://aclanthology.org/2025.mathnlp-main.11/.
time0ut··on Ask HN: What default model do you use and why?
If I had to pick just one, these days it’s Composer. It is just a cheap, fast, surgical workhorse. My workflow uses other models as well for various phases: Opus for planning, Sonnet for review, etc.

I prefer Cursor at this point just because of Composer. Claude Code is passable but the lack of a good, fast, cheap workhorse sucks. Sonnet and Haiku aren’t it.

I also do like Codex and Sol, Terra, and Luna. They are decent but I don’t find they stand out enough to use over the others.

Additionally, I have not tried Astra and found Fable to really not worth the cost for the tasks I do.

Finally, the latest Grok is actually a beast of a model, but expensive enough to not be a stand out.

time0ut··on Disruption with Some GitHub Services
Notice odd behavior on GitHub. Get gaslit by a green status page. Notice more odd behavior on GitHub. Think it must be me this time. See unusual action queuing. Ah, an incident on the status page. Go for a walk and check HN on my phone. The AI SDLC.
time0ut··on The entire city of San Francisco as a video game
This is awesome! I tried to make it to the Golden Gate, but my MBP got too hot to hold before I got there.
time0ut··on The Amazon tax
Wow. Amazon Haul looks like the latest way to speed run heavy metal poisoning.
time0ut··on Ruff v0.16.0 – Significant new updates – 413 default rules up from 59
Silicon Valley (the TV show) memed about it with its tabs vs spaces bit. It used to be a thing for sure. It has been a good 10 or 15 years since I had such a discussion. Automated linting and formatting tools largely killed it in my experience.
time0ut··on Ask HN: Is anyone using the A2A protocol?
Ya, we tried it a bit, but have stayed with agent-behind-MCP style patterns for now. I think A2A or something like it will become a big thing as everything matures. It just felt over complicated for our use case. One misconception I had was we would just slap A2A on our existing agents and they’d work well together. Kind of dumb in hindsight.
time0ut··on Claude: Elevated errors across many models [resolved]
Breakglass ChatGPT subscription
time0ut··on Anti-social: It's fads, not friends, which now dominate social media feeds
I have never been interested in the “normal” social media apps like Instagram, Twitter, Tiktok, and the like. The content never appealed to me as a consumer enough to get started. Occasionally something would go viral enough that a friend would eventually link it to me and that was the whole experience.

Recently, I made a dumb little app for my kids and decided to try marketing it on social media just to see what it is like. It is fascinating in a sense and disheartening as well. I have been very unsuccessful, but the most signal tends to come from the dumbest content I have tried.

In doing this, I have come into contact with the social media feeds I never felt the need to look at and man… they are like a drug. I find myself mesmerized by random IG reels. It is one thing to understand what they are on an intellectual level and a totally different to feel it first hand.

I miss MySpace.

time0ut··on 2026 HIPAA Security Rule Update
Interesting. I haven’t fully read through the rule change, but seems like HHS is directly adopting the controls required by HITRUST? I have been out of the industry for a while. Always interesting how the industry shapes regulation and vice versa.
time0ut··on Throwing AI-generated walls of text into conversations
The best are the Jira tickets with a huge wall of AI slop requirements. Usually full of nonsense of course including implementation recommendations in the wrong language or framework. Questions for clarification met with blank stares from the author. Ah well, copy/paste into claude code and say “do this. make no mistakes” and get back to browsing HN…
time0ut··on Gemini 3.5 Flash
I ran through the eval loop for a side project’s task (personalization of a micro video game, no thinking) last night. Head to head with Gemini 3 Flash Preview, results came out at basically a wash on my rubric. The output quality was good, well grounded, and reliable across 144 runs. But not noticeably better. It isn’t a traditional coding task, so can’t infer anything there. The amazing part was how fast it is. It was consistently about 2x faster than 3 Flash Preview and slightly faster than 3.1 Flash Lite Preview which is amazing. For my task, the price difference doesn’t matter, so easy upgrade. I plan to write up a quick blog post with the results over the weekend.
time0ut··on Trade Dollars with other startups. Book it as revenue
Pre-legal. That is gold.
time0ut··on AWS stops billing Middle East cloud customers as repairs to war damage drag on
Some data centers are more valuable as targets than others. For example, those comprising us-gov-east-1 and us-gov-west-1 or, god forbid, us-east-1. I don’t expect it is a difficult task to find them and other critical infrastructure for a state, but probably more involved than popping open google maps.
time0ut··on AWS stops billing Middle East cloud customers as repairs to war damage drag on
Data centers are such great targets in modern warfare. A few cheap drones can inflict billions in damage with low direct casualties (if the attacker even cares). I have heard AWS in particular is secretive about the exact location of their data centers, but no doubt every major country knows exactly where they are.
time0ut··on An update on recent Claude Code quality reports
Opus 4.7 via code has been inconsistent for me. Sometimes, it feels like working with a brilliant collaborator and is as good as 4.5 and 4.6 were. Other times, it takes dumb and lazy short cuts. It can be quite frustrating. Its response when I tell it it did something wrong is often to write a memory... which is then does not always read. The inconsistency isn't due to session length or age either. These are all new sessions. I feel like sometimes, I get routed do a dumber model or some other hidden setting is applied.
time0ut··on €54k spike in 13h from unrestricted Firebase browser key accessing Gemini APIs
It is scary building on the public cloud as a solo dev or small team. No real safety net, possibly unbounded costs, etc. A large portion of each personal project I do is spent thinking about how to prevent unexpected costs, detect and limit them, and react to them. I used to just chuck everything onto a droplet or VPS, but a lot of the projects I am doing lately need services from Google or AWS. I tend to prefer GCP at this point because at least I can programmatically disconnect the billing account when they get around to tripping the alert.
time0ut··on Show HN: Zerobox – Sandbox any command with file, network, credential controls
Very interesting. I just started researching this topic yesterday to build something for adjacent use cases (sandboxing LLM authored programs). My initial prototype is using a wasm based sandbox, but I want something more robust and flexible.

Some of my use cases are very latency sensitive. What sort of overhead are you seeing?

time0ut··on I am definitely missing the pre-AI writing era
I certainly miss the pre-AI reading era.

So much content is just straight copy/pasted from the LLM now. Articles, blog posts, linked in posts, reddit comments, etc. Even just using the LLM for 'editing' tends to shift the voice to an obvious LLM voice when used naively. It is getting worse too. Last week a co-worker sent me a screenshot of Claude for me to review their "work", which was just whatever Claude made up.

Usually, if something is very obviously unfiltered LLM output, I just stop reading.

I do use LLMs for writing myself. They are useful, but are poor authors.

time0ut··on What young workers are doing to AI-proof themselves
Optimistically, I hope it filters out the people who were only interested in it for the money.

When I was in school, decades ago now, very few people went into CS compared to other majors. Everyone I knew going into it did it because they loved it. I would have done it regardless of the career opportunities because I want to build stuff.

Interviewing candidates over the years since then, my experience has been there are still very few of those passionate nerds and a lot of people who did it for other reasons, like the money or similar. There is nothing inherently wrong with this. I don’t fault people for it.

Maybe if we get very lucky, it will go back to a relatively few passionate people building stuff because it is cool?

time0ut··on Astral to Join OpenAI
I love uv and the other tooling Astral has built. It really helped reinvigorate my love for Python over the last year.

Something like this was always inevitable. I just hope it doesn’t ruin a good thing.

time0ut··on GPT-5.4
Lowest common denominator.
Page 1 of 17Next →