HNHacker News
TopNewBestAskShowJobs

unshavedyak

2,402 karma · joined December 7, 2022

submissionscomments
unshavedyak··on Our $445M Series D
> It’s interesting to me that 9 hours of interview time is considered excessive by some people. Over my career I’ve done interviews that required plane flights, hotels, and multiple separate days on site.

I think part of the reason is that interviews have gotten more adversarial as time has gone on. Ie less respect for the interviewee, more of a power dynamic, more careless wasted time in lazy take home assignments, etc.

I wouldn't mind 9 hour interviews to 3 companies if i knew i'd land the job. It's the 1,000 companies that many people experience that gives them pause after so much disrespect through time.

unshavedyak··on Claude Haiku 5.5
this comment[1] seems clear that it will stay with subscription. Assuming they're authoritative of course.

[1]: https://news.ycombinator.com/item?id=49999702

unshavedyak··on Livenerf: Has Opus 5.5 been nerfed yet?
That's kinda me with respect to Claude. Generally i've had no issues and just kept pluggin' along.

The first real issue where i wanted to leave was the Claudish nonsense. If not for 5.5 i'd be on OpenAI by now.

unshavedyak··on OpenAI still doesn't seem to have a handle on all of its rogue AI activity
I agree with you, but i also fear it's accurate to what users an expect at home.

Is OpenAI worse at airgap/etc than Anthropic/etc or are its models worse at this rogue behavior?

Do we have special benchmarks for agentic systems going rogue in this manner? Sure it's OpenAI atm.. but it could be my grandmas PC next week. Concerning honestly.

unshavedyak··on When did Google get so weird?
Heck I kinda wonder if LLMs can do math as well as they can reason, “think”, etc. ie its all just probabilistic lunacy that somehow works great, so why are we so concerned about math being wrong? It could be wrong about the color of the sky, the size of a basket ball, how much oranges weigh, etc etc.

The nice thing about math is it can easily plug into a tool, making it even less of a concern.

unshavedyak··on Grok 4.7
> You become better at expressing your thoughts, but harder to understand.

This happens most though when the speaker doesn't (or care to) understand their audience.

Eg i find effective communication requires expertise in both the subject matter domain but also the reference of the listener. Eg in ELI5 framing, if you don't know what information 5yr olds are expected to know you'll do a poor job at an ELI5.

It often feels like Claude does poorly at both framing the response relative to what it "thinks" the listener knows, but also the prose is... sideways, just weird as you said.

unshavedyak··on Jev-Leftpad
It's funny, i actually worked with this guy (same small startup, not closely), and recognized him from that unique username on this post. It's been so long i wouldn't have had any idea if not for the oddness of the single letter username.
unshavedyak··on Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra
> and buys you the ability to correct from them

Or at the very least, make more mistakes.

unshavedyak··on Jellyfin 12.0
> There is a Caddy extension for doing DNS based validation, but that puts you back to needing a DNS provider with an API.

Yea, for clarity this is what i was describing. I use Porkbun's API and the Caddy instance isn't reachable from the internet, only local and over tailscale

unshavedyak··on Jellyfin 12.0
I've got a Caddy instance setup which proxies my traffic and has dead easy integration for domain DNS registration. Super handy
unshavedyak··on The asteroid currently hitting front end web development
Yea, i still deal with a ton of code in AI heavy workflows. If anything the frustrating part is absorbing the code quickly enough. AI (Claude for me) writes in cryptic text and the code flows can often be non-obvious.

I need (and am exploring) custom review tooling to improve this AI->Human code flow. Reviewing PRs were always the hardest part for me in programming. They were often full of the developers decisions and you have to rediscover those as you're reading code for it to make sense[1]. However i find this even more difficult to discover these decisions from AI.

However unlike human PRs we can ask more of AI. Rarely have i had a developer put on a presentation for a PR - but AI could right? AI could produce a guided walkthrough of the code. Not sure if it will help of course, but my thought is we're all stuck in the old "PR review flow" but instead of PRs it's AI - and the volume of them is far greater than anything prior. So i expect we need to tweak how we review, how we get information from LLMs.

[1]: I'm speaking generally, and about larger PRs. Not some small func where you can easily see what it does. Business logic and complex code can be difficult to decipher in PRs, imo.

unshavedyak··on Claude Fable 5.1 and Claude Mythos 5.1
I wouldn't mind it either. But the prose is obtuse atm. It doesn't feel like a writing style, it feels like an encryption.
unshavedyak··on Claude Fable 5.1 and Claude Mythos 5.1
And word on Opus 5.1 for writing style? I am on the edge of switching to OpenAI due to this horrid writing style. If Fable is better, great - but i can't even use that at work.
unshavedyak··on Fastpotify
The sad thing is that may not be true. If you use LLMs to write docs and actually read it, you can spend far longer trying to convince the LLM to write in a sane way than to just write it yourself.

Semi related, but i've caught myself nitpicking small things from the code of an LLM only to realize i'm putting in more work trying to convince the LLM to make the code the way i want rather than just modifying it myself lol. It's as if i'm trying to educate it in the ideal way, as if it would learn and not repeat the mistake.. but it rarely does of course.

It seems to me that interacting with LLMs has a large pile of cognitive traps for humans.

unshavedyak··on Fastpotify
Now that the public anger is growing i really hope Anthropics next models seriously address this. Or for that matter, i hope everyone (OpenAI/etc) starts focusing on this too.

Intelligence is important but it feels as if the LLM is only capable of conveying information while drenching it in vomit.

unshavedyak··on Notes on Private Trackers
> .. I think the incentives should instead encourage variety and availability like a distributed, highly-available, P2P archive.

Which is exactly what it is like for many trackers these days. They have an "economy" where some type of bonus point system allows you to buy upload credits. You often acquire this by either simply seeding over time, or by uploading special torrents with rewards, eg double upload credits, etc.

Some trackers even double down on this concept and basically only care about long term seeding. Ratio doesn't matter as much if you're seeding 24/7.

unshavedyak··on Walgit – a Git server that is one binary in front of an object store
> There's Git LFS, that hooks up an object store (like S3) to git for large files.

As an aside i keep meaning to review Git LFS to see if there's some fundamental reason that it requires a server. Eg could Git LFS write directly to an Object Storage?

I run a lfs proxy at home and it bothers me that it exists heh. Though i don't run ObjectStorage (minio/etc), so i guess swapping out my lfs server for OS wouldn't really net me anything - but still, feels like it was designed first and foremost for the Github API rather than generically for Git users with Git principles.

unshavedyak··on Fable and the end of the free lunch
Those "present state" comments are the bane of my existence. It was present in 4.7/etc but i put in a ton of guards against that into my global memory and it worked quite well. Fable and Opus 5 regressed badly in this space though and i can't keep it from making those types of comments again.

Really frustrating.

unshavedyak··on Vomit: Clean up Claude 5's token output with a separate LLM
It's such a great example. It writes so well compared to Claude.

I'd love to know what the hell Antrhopic has done to make Claude's writing so, so bad.

unshavedyak··on Vomit: Clean up Claude 5's token output with a separate LLM
I'm on my last straw with them. I've been around for a year now and for many months i've just stuck with Claude because it was plenty good and i didn't care to provider-hop to constantly compare. Previously though my UX wasn't actually affected that much, despite growing complaints/etc, generally everything was fine for me.

Opus/Fable output these days though is... not enjoyable. It's just really bad. The code quality is fine, but i want information from claude and it's just awful to read.

My biggest problem honestly is that i can't move my day job.. we're using enterprise claude and i'm not sure how much effort it would be to get access to another provider. I should inquire though, claude is really frustrating these days.

unshavedyak··on Claude Opus 5
> I wonder if there's a correlation between me refusing to use LLMs and me being happy to read a novella-sized PDF about them.

Semi related, but i would hate to read that PDF but i also hate reading what LLMs write lol.

LLMs are pretty terrible at being concise. Using an LLM these days means putting up with bizarre and often confusing phrasing, wordy explanations, etc. It's kinda crazy to me how good they are but how bad their writing style is for me personally. Even though i use an LLM constantly i can't stand reading its responses.

unshavedyak··on Making
I have a slightly different approach. I've done this as both a profession and a hobby for around 18 years now, and even these days i spend quite a bit of time "making" things i take pride in. Be it decisions (decision fatigue is a battle), code written, architectural improvements, etc. I don't generally feel lacking in this area.

Where i've felt lacking for the last ~5 years though is output. Specifically blocked by having the energy to create all these damn ideas.

I find LLMs neat, albeit their own type of exhausting, as they open up a lot of possibilities for things i want. Ie software i've wanted but have never cared to find the time for or didn't have the time for.

LLMs feel like a software equivalent to a 3D Printer. I don't have to carve it out of wood now. But also like a 3D printer there are some tasks that LLMs are just horrible for. So it's a niche, a skill even, to find software you want that also are a good for for unsupervised development. It's also of course massively more risky than a 3D printed doohickey depending on its internet access/etc.

With that said i still don't really take pride in something vibe coded. I just take enjoyment out of using a thing that i've wanted but never had time for.

unshavedyak··on What's the deal with all the random weekly quota resets for agents lately?
Stop, what, exactly? This is simply a new tool, not an AI girlfriend. We've had ML based autocomplete for 10 years now, why is it suddenly a new chicken-little for ya'll?
unshavedyak··on What's the deal with all the random weekly quota resets for agents lately?
Yup, or we'll find out there's nothing to fear for folks who are using this responsibly.

It seems a huge "the sky is falling" to think all LLM use is bad for fear of some "addiction" to me. Even many skeptics (eg Hashimoto) have come around to the idea of using LLMs.

One thing is clear in my mind. VCs are burning cash, and so if you have an LLM flow you find useful, take some free cash. The sky is not falling in this respect imo.

Which is not a defense of AI, to be clear. AI may very well be a big societal problem, but in this context i don't see Hashimoto/etc becoming heroin addicts like ya'll are so concerned about. It's just a fancy autocomplete. The fear seems a bit over hyped.

unshavedyak··on What's the deal with all the random weekly quota resets for agents lately?
Which is funny because I’ll take the free crack in this case lol. Am I addicted to this new workflow? No idea but there are many providers. I feel like I’m being given free VC money so yea, I’ll use it.

My expectation is that this cash handout is going to stop soon, so take while the giving is good right?

I should note that for my workflow /loop is eating the credits, so it’s basically no skin off my back. I just queue up some more work and let it go.

unshavedyak··on DEA to Temporarily Schedule 7-Oh and Related Substances to Protect Public Safety
Huh, don't think i had any of those. Though arguably i had "Lack of focus", difficult to say how much is due to the lack of caffeine or due to undiagnosed ADHD though.

Generally i felt fine. I'll keep it in mind, thanks

unshavedyak··on DEA to Temporarily Schedule 7-Oh and Related Substances to Protect Public Safety
Huh, i should look at this. I've been an aggressive drinker for most of my adult life (2 pots a day at my height), but for kicks i decided to cut all caffeine for about 9 months. No real issues aside from very short term headaches, though even those i mitigated by gradually moving down in quantity.

Aside from the headaches what addictive effects are you referencing?

unshavedyak··on Blender 5.2 LTS
For those of you who get value out of Blender please consider donating: https://fund.blender.org/ :)

It's a wonderful project. Even small amounts help. Thanks!

unshavedyak··on I used to love Claude, but the latest models are slowly ruining it
I'm tempted to try it out. I'm not keep to move but Fable rejected some work i was working on recently and frankly it's infuriating lol. I've been on Claude x20 for like 8 months now and now i'm tempted to switch out of spite.

It is surprisingly offensive.

The only friction for me is the general expensiveness of trying out top tier models, eg OpenAI's Fable equivalent (Sol?) to run for a trial period. I'd like to see a like-like comparison, eg buy x20 on OpenAI and see how much Sol i can use, how well it works, etc.

edit: Though surprisingly Claude seems to think Codex doesn't have hooks? That'll be tough, i use Claude hooks quite a bit.

unshavedyak··on SWE-1.7 Reach Near GPT 5.5 and Opus Intelligence
It's a lossy conversion though. "Mistake" is relative to the stated goals and specifications which are often heavily lacking. So unless you write with a high degree of architectural and implementation specificity then it might make very high quality code that is still not what you wanted.
Page 1 of 34Next →