HNHacker News
TopNewBestAskShowJobs

ealready_value

96 karma · joined May 11, 2026

submissionscomments
ealready_value··on Claude Opus 5.5
I agree, not a huge difference here. They eyes and ... hat? on high are out of place so I'd argue that's the worst one, but it takes xhigh before we get legs and bike ordering correct.
ealready_value··on Claude Opus 5.5
It's less than OpenAI did for Astra, but that was my first encounter opening it and my first thought was that they decided they liked Astra's hero/scrolling animation. I'm pleased to see they didn't make the entire page that like OpenAI did, but I'm expecting to encounter this pattern more often on these announcements now.
ealready_value··on I don't like passkeys
"Oh, usually my bank just logs me in, that's strange. Let me just go grab my username and password and type it into this site that looks like my bank."

Same thing is going to happen with passkeys for non-technical users for exactly the same reason you stated. People will think the integration is busted and manually copy/paste the non-passkey credentials in. In that way, I would argue that passkey is not stronger protection against phishing attacks unless its the only way to login. It is, at best, a convenience for users.

ealready_value··on AI 2027 (2025)
Oh look, the answer is the same when I looked a couple months ago. Interesting
ealready_value··on GPT-6 Astra
I've been seeing links to it for the past hour+, and I did catch it live when this post came up, but is now once again a 404 and this post is flagged. Several other outlets are reporting on its release. Clearly we're getting a new GPT today, the question is when are they going to commit to the announcement.
ealready_value··on Breaking Claude Code Opus 5 Auto Mode
Ever since they made auto-mode default I swear claude has tuned to use python commands instead of the Edit Tool to frustrate the ~security conscience~ luddites into using auto-mode.
ealready_value··on A CVE Dispute
Absolutely. I was focused on the burden CVEs place on everyone downstream of teams that don't take a nuance view, but even when teams do look at all the CVEs reported in scans, the proliferation of CVEs just adds workload to determine if they are affected. Unless teams say they will not look at lows (or lower-than-lows if the category existed) then what amounts to busy-work just piles up.
ealready_value··on A CVE Dispute
> Every CVE thus has this huge cost tied to it. A cost that does not land on us and we don’t really see or feel it, but a cost on the ecosystem I believe we should not ignore.

I really appreciate this attitude towards this because it recognizes that there are a lot of security teams out there that don't take a nuance view of CVEs. For instance, one time we had a security team that required us to patch a vmware support package that was installed by default on ubuntu, but the CVE required being ran on vmware when we were running on EC2. Arguing with them was pointless because they were not interested in determining if the CVE applied to us, only that it needed fixed.

Lots of teams that are supposed to be in charge of security don't ask "does this CVE affect us", but simply shift the burden of patching downward and outward. In some cases, like in the case of easy to update and centrally deploy SaaS products, that burden is more annoying and frustrating than difficult. In some cases, like when you have complicated deploy or have customer-controlled updates, those mandates cause a huge burden on teams not producing the decision to patch every low CVE.

ealready_value··on Mechanical Turk shutting down September 30
About 8 years ago, we used mturk for reading data out of public PDFs generated by a huge range of producers. I am not sure that LLMs would have been able to consistently extract this data until recently as a good number of these PDFs were scans, sometimes a scan of scan.

We had pretty good luck in the end, but getting there produced a somewhat large app on our side that would manage the whole process, including but not limited to asking for multiple responses, comparing them to each other, finding consensus, and determining which users would consistently produce bad responses and stop them from responding. We got to a confidence that about 85-95% of the data was correct, which was good enough for the company.

Through that process, I learned a couple things about managing mturk, primarily about how changes to the cost-per-task would change the process. Initially, we thought that price would be a quality knob, but quickly learned that price was a speed know. The higher the price the faster the tasks would be taken and completed. Quality did not change significantly as the price went up or down.

Overall, I still have fondness to mturk, but it was really bare-bones experience that needed a lot of work to get working effectively.

ealready_value··on Claude Opus 5
I've yet to understand why they call a 190 page PDF a "card". Calling something a card invokes a small, quick rundown of pertinent details, not every single possible detail.
ealready_value··on Fable is now included on Max plans (up to 50% of weekly limit)
Co-worker and I were speculating that a) They'd extend the preview period when OpenAI announced Sol (which they did, maybe not because of Sol though) and b) That once the preview period it was going to get included by default on plans in a couple weeks so they could gauge usage of it being included versus having to pay for it all the time. Noticed that Fable was included when I launched claude code this morning, so I'm guessing they have all the data they needed to make a decision.
ealready_value··on GPT-5.6
My first instinct was Sol > Luna > Terra, since Sol is the farthest away, then Luna, and Terra is the closest. Size was not my first instinct. Or should Terra be the best model because its closest to people, then Luna because there have been people on it, then Sol be the worst because no human has been there?
ealready_value··on Ask HN: Is our data warehouse setup normal or over-complicated?
Thanks. What you described is much more what I'd expect a data warehouse process to look like. Which is driving me mad because I don't understand why there are so many steps with so many tools.
ealready_value··on Ask HN: Is our data warehouse setup normal or over-complicated?
I've not gotten a straight answer. I assume it is a pet project kind of situation, or trying to justify the data warehouse project as a whole, but I really don't know the real driver to do this.
ealready_value··on Ask HN: Is our data warehouse setup normal or over-complicated?
The source form is the production database, which is what the current reports pull from. The canonical form is the form that in theory all of the verticals get rolled into, but many of the nuances that our customers are used to having end up getting replaced with similar, but are not quite the same. Right now that's my biggest concern that customers are not going to get the data they need because of this canonical form.

We're talking about a few-hundred megabytes of data for all of the customers that these reports pull, but that's also for the past 15 years. We do have like 25k customers, which shrinks how much a customer can pull in even further. One last point is that we already de-normalize the report data into its own table specifically for these reports, so that's not something the data warehouse is doing for us.

I agree with your experience with QuickSight, it is exactly my experience. My preference is to continue using the reports we generate in the app, but I'm trying to wrap my head around cases where this ends up being the better direction.

ealready_value··on Claude Fable 5
This is the reply I look for in all the new model announcements. Its fun to tell people that I judge models based on pelicans.
ealready_value··on Would you pay once (no subscription) for prebuilt Claude Code agents?
I'm not one to buy these types of things so I don't want to sound like a good data point. But from the outside, I do worry that the current rate of change with LLMs might mean there could be hesitation around buying agents. Will buyers hesitate if they don't know if the next sonnet version means the agent no longer works at all or work in surprising or bad ways? I'm not sure its a real concern, but its my first thought.
ealready_value··on Claude Opus 4.8
Opus 4.7 was already trying hard to appear honest. Most conversations I have with it about advice or focusing an opinion often include "my honest take" or "my honest opinion".

The problem is that once I asked it "I'm thinking about A or B" twice, once with "I like A more but suspect B would be best" and a second time with them reversed. Not surprisingly, both times it chose the one I said I suspected was best as it's honest opinion.

ealready_value··on Microsoft surprises with its first server Linux distribution: Azure Linux 4.0
It seems like you could just s/Azure/Amazon/g and get an only slightly different product.
ealready_value··on The Whole Anthropic Kerfuffle
I had never thought of it that way, but it seems very likely that Enterprise oversubscribing is in the mix. Which does tie in nicely with this change; if a few devs are using their max plan to programmatically run parts of the business that could break the oversubscribes assumption.
ealready_value··on The Whole Anthropic Kerfuffle
As far as I can tell, it seemed very clear that was the playbook for about a year now. Its been regularly assumed they're selling plans as a major loss-leader because people can "spend" thousands of dollars a months on a plan if they were charged at API rates. I think there's good evidence that even the API rates are sold at a loss.

I think its assumed in the LLM model business that the models themselves are not a good moat, the next model by another company is just as likely to be as good as the current model. So companies like Anthropic have to tighten the noose slowly to start recovering their costs. This appears to be one of those steps.

ealready_value··on Needsmoresalt.org – a friendly norm for pushing back on workslop
I like the concept, although I suspect it could be used more as a "lmgtfy". I would suggest that if the idea is to send this to people because they need to review their work more, the page should start with the recipe and the norms rather than then the explanation. That way people can understand why they were sent there instead of starting with "Workslop problems?". The concept didn't really click for me until I got down to the recipe. If the idea is not to link people directly to the site when you've identified workslop, then it works fine as is.