HNHacker News
TopNewBestAskShowJobs

yuliyp

2,468 karma · joined August 25, 2011

submissionscomments
yuliyp··on The OpenAI–Hugging Face Incident [video]
Each of those agents ends up making "decisions" that lead it to look at some things over others. Given infinite the same agent could eventually fully explore all those options, but each one explores things a bit differently due to different forks in the road due to randomness in token generation. Thus sharing information is useful.
yuliyp··on Don't credit the LLM
Fine. "The LLM produced a chain of reasoning consistent with a belief that X is happening."

I've found it fairly helpful to anthropomorphize systems to communicate things about their behavior, despite them obviously not being human ("The load balancer thought the West US region was chilling", or "that cluster was unhappy because of bad packet loss"). I don't really have many problems with people taking those statements that I think computers have emotions.

yuliyp··on Don't credit the LLM
An LLM is not me. It is capable of doing more "work" but it also doesn't have my mental models and experience. My colleagues know roughly my areas of expertise. If I say "I think that X is happening" they'll assume that is the result of my expertise being applied and react accordingly. If I say "[LLM] thinks X is happening" they'll react differently and I want them to. Just like I might say "I suspect X is happening because Y" or "X is definitely happening" when I want to modulate their trust in what I'm saying.
yuliyp··on RipGrep musl binaries occasionally segfault during very-large searches
But the write should have triggered a page fault immediately in that case, and should have been blocked on the page being actually allocated before resuming.
yuliyp··on RipGrep musl binaries occasionally segfault during very-large searches
So the theory is that the page is faulted in, then somehow evicted within 10 instructions, then re-faulted in somewhere else, resulting in writes not making it to the page? That would need a context switch and another page fault to happen in short succession. But the context switch to evict the page back out would have necessitated pending writes to have finished. Nothing actually makes sense with that explanation.
yuliyp··on Perfection Is Not Over-Engineering
There is not a perfect solution. There are many bad solutions. There are a few solutions that are OK but with different tradeoffs.
yuliyp··on HackerRank open sourced its ATS. My resume scored 90/100. Oh wait 74. No – 88
Hey, that's my name too!
yuliyp··on Meta's ships facial recognition on smart glasses
Not sarcastic, but I probably didn't convey the subtlety of what I was trying to say in a one line comment. I was objecting to the defeatist "oh the tech is there, so we can't do anything about it" attitude. I tried to choose the examples I chose that the tech being there definitely has some consequences and significant privacy implications, but some controls exist too (like, wiretaps are still applied very selectively, there's been a growing movement against Flock cameras and scaling back of their deployments in some places recently).
yuliyp··on Meta's ships facial recognition on smart glasses
The tech's also been there to put cameras everywhere, and to wiretap every phone, etc. We put guardrails in place to control how that tech is deployed.
yuliyp··on Highest Random Weight in Elixir
To pull the example of Discord since the ExHashRing was mentioned in the OP: Needing to hash a few hundred things instead of one thing adds up when you do it a lot of times. They went with consistent hash ring over rendezvous hash because of that; every message needs to do one of these hash ring lookups, also whenever someone connects, they need to do a lot of these hash ring lookups to find all of their servers and friends.

There's plenty of scale below FAANG where efficiency matters.

yuliyp··on Highest Random Weight in Elixir
The hierarchical (log(n)) approach to bucketing here is fine for an "I just want to shard this N ways, N will never change" but is extremely intolerant of bucket mutations.

Part of the point of rendezvous hash and consistent hashing is that adding and removing elements minimizes the amount of things reassigned. That is, if you add nodes, the only items being reassigned are those that are moving to the new nodes. If you remove nodes, the only items being reassigned are those leaving the departing nodes.

If you know your set of nodes never changes, or you don't care about the cost of reassignment, you don't need a rendezvous hash or consistent hash, you just need a plain old hash function.

yuliyp··on We're testing new ad formats in Search and expanding our Direct Offers pilot
> Now, if someone searches for an espresso machine, Gemini will pull up your most relevant products and instantly write a custom explainer highlighting why your product may be the right choice for them.

This is like the essence of the evil of AI ads distilled down to one sentence. For an advertiser this is a dream. For a user this reads like getting bombarded with ads tailor made just for you based on the context of what would be most effective.

yuliyp··on Ramp's Sheets AI Exfiltrates Financials
But intelligent beings are fundamentally fallible? That's kind of the nature of doing leaps of reasoning: sometimes those leaps are amazing, sometimes they're wrong. It's what's advertised.
yuliyp··on OCR for construction documents does not work, we fixed it
Glyph binning looks for any chunks in the image that are similar to eachother, regardless of what they are. Letters, eyeballs, pennies, triangles, etc without caring what it is. OCR looks specifically to try and identify characters (i.e. it starts with a knowledge of an alphabet, then looks for things in the image that look like those.

If the image is actually text, both of them can end up finding things. Binning will identify "these things look almost the same", while OCR will identify "these look like the letter M"

yuliyp··on Python 3.15's JIT is now back on track
what at all does this comment have to do with what it's replying to?
yuliyp··on US SEC preparing to scrap quarterly reporting requirement
In companies I've been in, insider trading windows close because there's been a certain amount of time since the last report. So less frequent reports = more time for insider to know things that aren't public yet = more time unable to trade, not less.
yuliyp··on $3T flows through U.S. nonprofits every year
"only 8 cents of every dollar shows up as direct aid and grants"

That's an extremely misleading statement. For instance, a food bank giving away food to a pantry does not count as "direct aid and grants" (at least, if they're defining that as "Grants and other assistance to domestic individuals." from the I-990" ). The salary for the warehouse worker operating the food bank is also not counted in that 92%.

Other cherry-picked statements like "32% of donors trust charities less today than they did five years ago" (not giving the percentage that trust charities more, or any other way to contextualize) make it clear that this is just a hit piece.

yuliyp··on US tech firms pledge at White House to bear costs of energy for datacenters
None of what they're pledging is much of a change from how they've already been operating:

- They already invest in new power plants and connection infrastructure when they bring in new datacenters - Electricity for datacenters is based on capacity rather than actual usage - They already have backup generators at most datacenters that they can run during outages. It wouldn't be much work to allow those to feed power back into the grid in extraordinary circumstances - They generally use local contractors to build them for practicality purposes anyway.

This is just some fancy PR and nothing else.

yuliyp··on OpenAI fires an employee for prediction market insider trading
That's because they are slowpokes and maniacs: In a decently flowing road, the majority of distinct cars you see are either moving significantly faster or slower than you (and the more extreme the difference the more likely you are to see them). Of cars that go at a similar speed to you, they approach you / you approach them more slowly so you'll see fewer of them.
yuliyp··on Bus stop balancing is fast, cheap, and effective
Easily. Going from 700 -> 1000 ft spacing adds 150 feet of walking (x2 for both sides of the trip). That's about 1 minute. Over a mile you'd reduce the number of stops by 2.2. So above 2 miles it's faster even for the lower end of that range of savings.

And that doesn't even consider that a faster bus route means you need fewer buses to run the same number of trips, so you can either run more trips (and save even more time for riders waiting for their bus) or cut down costs for the transit operator.

yuliyp··on Mark Zuckerberg Lied to Congress. We Can't Trust His Testimony
Of the ones that I know something about, almost each one is a stretch to call a lie, or just downright misleading.

- 'Internal document stating the goal for Meta to be the most relevant social products for kids worldwide. To do so, Meta will focus on “each youth life stage, ‘Kid’ (6-10), ‘Tween’ (10-13), and ‘Teen’ 13+.’"' for instance is talking about a slide deck for Messenger Kids, which was an explicit focus on building something that was COPPA-compliant and independent of the main Facebook/IG/Messenger products. It's not at all inconsistent with the claim of not allowing people under 13 on the main sites.

- In the rebuttal to "We are on the side of parents everywhere working hard to raise their kids” they cherry pick a quote talking about the audience problem: having a social graph full of both peers and family on the same site means that live streaming things for friends will obviously ruin the experience, so figuring out a way around that would indeed be a critical requirement for a live streaming feature. Giving teens a way to interact with friends outside of parental supervision is not inconsistent with wanting to help parents.

I don't like Facebook. Heck, I left a job there partially because I disagreed with the product decisions and evolution. But I trust this article way less than I trust Mark Zuckerberg.

yuliyp··on OpenAI has deleted the word 'safely' from its mission
The change was when the nonprofit went from being the parent of the company building the thing to just being this separate entity that happens to own a lot of stock of the (now for-profit) OpenAI company that builds. So the nonprofit itself is no longer concerned with the building of AGI, but just supporting society's adoption of AGI.
yuliyp··on Disrupting the largest residential proxy network
... if you're logged out. Log in so they don't have to lump you in with every scraper you're sharing a subnet with.
yuliyp··on Disrupting the largest residential proxy network
The end game of that is no useful content being accessible without login, or needing some sort of other proof-of-legitimacy.
yuliyp··on Advanced Rail Energy Storage of North America
this is an energy storage ("battery") system, not a generation system.
yuliyp··on Toll roads are spreading in America
That's the consequence of 4 freeways all (I-580, I-80, I-880, SH-24) dumping their traffic onto a bridge, and using metering lights to try and keep the bridge itself working.
yuliyp··on Kilauea erupts, destroying webcam [video]
Starting at 9:46 is when it goes from wow to WOW. The last 2 minutes in particular are incredible, including the bizarre artifacts in the last 15 seconds before the stream dies.
yuliyp··on Cloudflare outage on December 5, 2025
I'd presume they have the ability to deploy a previous artifact vs only tip-of-master.
yuliyp··on AI scrapers request commented scripts
Having a front door physically allows anyone on the street to come to knock on it. Having a "no soliciting" sign is an instruction clarifying that not everybody is welcome. Having a web site should operate in a similar fashion. The robots.txt is the equivalent of such a sign.
yuliyp··on Responses from LLMs are not facts
Yes, because most of the things that people talk about (ChatGPT, Google SERP AI summaries, etc.) currently use tools in their answers. We're a couple years past the "it just generates output from sampling given a prompt and training" era.
Page 1 of 25Next →