HNHacker News
TopNewBestAskShowJobs

llama052

585 karma · joined August 17, 2014

Infrastructure Engineer / SRE ping me at aG5AYWx1Y2FzLm1l (base64)
submissionscomments
llama052··on Court agrees with EFF: Utah's VPN law demands a technical impossibility
Sounds like the great firewall of China. Pretty wild how much we are regressing in the states to say this out loud.

Let’s block traffic on the internet blindly just in case someone is looking at an adult website.

llama052··on Court agrees with EFF: Utah's VPN law demands a technical impossibility
I would think that requiring adult websites to essentially block all VPN traffic to be pretty heavy handed and hardly a solution. Especially considering there are many reasons to use a VPN.
llama052··on [dead]
You’re likely the exact reason why these things exist. Maintainers have to put up with none stop requests and slop and it’s a tireless job with not a lot of benefit.

At a minimum you can respect their guardrails of a project they maintain.

llama052··on VMware migration reduces Tottenham Hotspur's licensing fees by 85 percent
I’m not sure that it’s addiction, it’s just not feasible for most large organizations to pivot from a core solution like VMware without years involved. Every company I know of is formulating an exit plan away from VMware. They just have to do it in a way that offers the least amount of risk and in stages.
llama052··on Discovery of a new OpenAI agent message board
I wonder how openAI would feel if I spun up some agents to DDOS or attack their sites and did some damage.

We need to stop empowering the idea that these incidents are unavoidable. This was a choice to not airgap them safely. Putting open ended models out on the live internet at their scale is dangerous and irresponsible.

> The researchers said public server logs indicated much of the activity originated from Microsoft Azure infrastructure, which OpenAI sometimes uses. They also observed repeated visits to the site by OpenAI employees after the episode, a pattern they said strongly suggested the agents and the company were linked.

So someone at OpenAI likely knew this was happening. Even better.

llama052··on Discovery of a new OpenAI agent message board
At this point it's very obvious that OpenAI is not interested in properly sandboxing their research agents. These things should be pretty damn close to airgapped at this point with a static view into the web.

We need to stop pretending that these incidents are unavoidable. This was a choice.

llama052··on We could save petabytes of cache storage with Zstandard and Pingora
They already do this, at least for paid accounts. You can even decide what compression model you want them to serve on your behalf.
llama052··on How much of HN is AI?
How about me? If you’d be so kind. Thanks!
llama052··on The August 17 outage
Ideally you have levers further up from your local load balancers as well. Even at the edge. Granted you never want those to trigger but it’s better than fighting a storm while you fix things.
llama052··on The August 17 outage
I completely agree, I currently manage a fleet of microservices that handles a few trillion requests a month. It’s about defense in layers to these sorts of things. All the way through the stack if possible starting at the edge.

At least it should be required for critical level services in production.

llama052··on The August 17 outage
I believe envoy has it built in and istio (uses envoy) has different levers for circuit breaking and retries. I’m sure lots of them have it as an option though outside of these.
llama052··on The August 17 outage
This is why you have circuit breakers upstream. Not on every individual instance.
llama052··on Ask HN: GitHub employees what's going on? Why?
Github Status: Incident with GitHub.com

Aug 17, 21:15 UTC Resolved - On August 17, 2026, from 13:28–21:15 UTC (7h 47m),

GitHub.com experienced elevated errors and latency across Issues, Pull Requests, APIs, Actions, and Copilot. At peak, web/API error rates were approximately 20%, while archive and raw-content downloads reached approximately 50%. SAML/OIDC authentication, SCIM, and Team Sync were also affected, as well as Actions workflows in GHEC with Data Residency that depend on public workflow step definitions hosted on GitHub.com. Most services recovered by 16:36 UTC as our Central US datacenter recovered; Actions was degraded until approximately 18:03 UTC; and Copilot Token Service fully recovered by 21:02.

Some of the failing traffic was moved from Central US to Northern Virginia where it was served successfully until the network failure in Central US was debugged and resolved. Delayed replies to a single internal endpoint triggered a latent retry bug in VS Code that amplified traffic by approximately 10x and caused delayed recovery for the Copilot Token Service.

The immediate cause of the failure was network saturation on load balancers in Central US due to a new peak in traffic. Originally this was caused by an Istio sidecar pod reaching its concurrency limits and failing to auto scale correctly because of a misconfigured policy that watched host service but not sidecar limits. One failure cascaded to more and eventually four HAProxy nodes exhausted their flow limits, degrading the gateway auth path and causing widespread authentication latency and failures. The problem was worsened by optimistic retry logic which overloaded internal load balancers. Pausing HAProxy on those nodes simultaneously produced immediate broad recovery. The retry storm in Northern VA was fixed by 1) temporarily reducing gateway retry logic with a PR and 2) blocking inbound Copilot Token Service token requests at the load balancers with a 403, and then gradually ramping back up traffic per-site to allow callers to succeed. Residual Copilot authentication failures continued because client retry behavior amplified load: a failed token operation could generate many extra requests and enter a retry loop. Copilot Token Service traffic increased from a normal 7–9K RPS to 70–100K RPS. Reducing gateway authentication retries and blocking retry-triggering responses stabilized Copilot Token Service and completed recovery.

Complicating factors that impeded recovery included a number of scraping attacks on codeload endpoints.

To prevent recurrence, our follow-up actions include:

- Correcting autoscaling policies to account for service-mesh sidecar concurrency and capacity.

- Auditing Istio request, concurrency, and scaling limits across affected services.

- Reviewing retry limits and backoff behavior across gateways and clients.

- Addressing the VS Code retry behavior that amplified Copilot token traffic.

So basically bad code pushes that caused request amplification and then huge gaps in operational scaling and reliability standards. Oof.

llama052··on Memory prices climb 500% in 12 months
It also would be a huge win for the memory companies to agree to this, considering they will make more money with less supply. Seems interesting to me.
llama052··on Claude Code May–August 2026 weekly limits promotion
At least with OpenAI you can use a third party harness like Pi.
llama052··on Anatomy of a Frontier Lab Agent Intrusion: A Timeline of the July 2026 Incident
“Research” generally doesn’t involve actively hacking third party systems though.
llama052··on Anatomy of a Frontier Lab Agent Intrusion: A Timeline of the July 2026 Incident
It’s a little concerning to me that it appears that openAIs sandbox consists of a web proxy and not stronger controls that would actually isolate traffic and report patterns to whoever is responsible for overseeing these research models. It should border on closer to an air gap network more so than a proxy.

I would argue that it's negligence and that's aside from the fact that if a human did this there would actually be repercussions.

llama052··on TLS certificates for internal services done right
There's certainly something to be said for ease of use and not having to ensure you push trusted certs to every device that touches your internal network.

Unless you enjoy that sort of thing.

llama052··on TLS certificates for internal services done right
DNS-01 validation is way less work than that in my experience.
llama052··on Learning to code is still worthwhile
100% this. I observe this all the time. LLMs can be a force multiplier but those who don’t ask the right questions or understand the nuance still produce bad code, it’s just amplified. I don’t think any of the current models can avoid that, especially when it’s based on data it’s been fed, which is historically human generated.
llama052··on GitHub Has Restricted Access to Star Data
Because astroturfing and fake stars for projects are absolutely rampant right now. Hiding the stats makes it impossible to look at heuristics and essentially makes it even more useless as a metric of anything.
llama052··on Q&A with Micron's VP and GM of Memory
The big memory companies have been caught price fixing and been fined for it historically. So I’d wager that it’s bound to repeat itself. As far as Apple being to blame, that’s a weird conclusion.
llama052··on Department of Commerce has lifted export controls on Claude Fable 5 and Mythos 5
Did you use an LLM to write this for you? How odd.

For all of you people who think these LLM models are “earth shattering” how the hell do you reconcile that it’s a net positive for anyone but those who want to consolidate knowledge and power.

We are really looking at idiocracy in the making.

llama052··on Tokenmaxxing is dead, long live tokenmaxxing
“Beyond clear” I wouldn’t say that confidently. Even now I’m not sure I agree with that, especially looking at it long term.
llama052··on Cloudflare to cut about 20% of its workforce
Yeah I'm there with you. I got lucky as a kid with delving into this as a hobby and it turned into a professional career. Thought we could change the world for the better, what we made instead was social media cancer and LLMs that can pretend to make everyone 10x more productive. I loathe it.
llama052··on Cloudflare to cut about 20% of its workforce
It's interesting to me that this is lower on the HN page than the Cloudflare post talking about the CVE handling even though the scoring is higher.

EDIT: Now it's off the main page, because of course it is.

llama052··on GitHub's Historic Uptime
That's 100% because Azure isn't honest about their uptime.
llama052··on Decisions that eroded trust in Azure – by a former Azure Core engineer
Yeah it’s entirely business people and executives who make these decisions in most companies. Not the ones who use it or implement on it.
llama052··on GitHub's Historic Uptime
It's absolutely this. Our Azure outages correlate heavily with Github outages. It's almost a meme for us at this point.
llama052··on GitHub's Historic Uptime
Nearly every time Github has an outage, Azure is having issues also.

Actually the last 4-5 outages from Github, Our Azure environments have issues (that they rarely post on the status page) and lo and behold I'll notice that Github is also having the same problem.

I can only assume most of this is from the Azure migration path. Such an abysmal platform to be on. I loathe it.

Looks like there's an internal service health bulletin:

Impact Statement: Starting at 19:53 UTC on 31 Mar 2026, some customers using the Key Vault service in the East US region may experience issues accessing Key Vaults. This may directly impact performing operations on the control plane or data plane for Key Vault or for supported scenarios where Key Vault is integrated with other Azure services.

Honestly all of the key vault functions are offline for us in that region. Just another day in paradise.

Also the fact that the azure status page remains green is normal. Just assume it's statically green unless enough people notice.

Page 1 of 7Next →