HNHacker News
TopNewBestAskShowJobs

landhar

839 karma · joined December 11, 2009

submissionscomments
landhar··on I were 17, I'd learn how to build LLMs from scratch
> Current models are clearly nowhere near the efficiency limit (the brain does vastly more with far less power).

I think this is disingenuous. One could say that drones are nowhere the efficiency limit either: a bee can fly for hours on the energy contained in just a few milligrams of honey, while our best battery-powered drones can't stay airborne for more than 30 minutes. But comparing energy efficiency of electric/mechanical devices to their biological counterparts is not an apples-to-apples comparison. There's a world of difference between the energy storage and delivery mechanisms.

And as many have pointed out already in the siblings, it's not just about the compute but the access to petabytes of training data.

landhar··on Why are anime catgirls blocking my access to the Linux kernel?
> But here we're talking about a system where legitimate users (human browsers) and scrapers get the same value for every application of the work function. The cost:value ratio is unchanged; it's just that everything is more expensive for everybody. You're getting the worst of both worlds: user-visible costs and a system that favors large centralized well-capitalized clients.

Based on my own experience fighting these AI scrappers, I feel that the way they are actually implemented makes it that in practice there is asymmetry in the work scrappers have to do vs humans.

The pattern these scrappers follow is that they are highly distributed. I’ll see a given {ip, UA} pair make a request to /foo immediately followed by _hundreds_ of requests from completely different {ip, UA} pairs to all the links from that page (ie: /foo/a, /foo/b, /foo/c, etc..).

This is a big part of what makes these AI crawlers such a challenge for us admins. There isn’t a whole lot we can do to apply regular rate limiting techniques: the IPs are always changing and are no longer limited to corporate ASN (I’m now seeing IPs belonging to consumer ISPs and even cell phone companies), and the User Agents all look genuine. But when looking through the logs you can see the pattern that all these unrelated requests are actually working together to perform a BFS traversal of your site.

Given this pattern, I believe that’s what makes the Anubis approach actually work well in practice. For a given user, they will encounter the challenge once when accessing the site the first time, then they’ll be able to navigate through it without incurring any cost. While the AI scrappers would need to solve the challenge for every single one of their “nodes” (or whatever it is they would call their {ip, UA} pairs). From a site reliability perspective, I don’t even care if the crawlers manage to solve the challenge or not. That it manages to slow them down enough to rate limit them as a network is enough.

To be clear: I don’t disagree with you that the cost incurred by regular human users is still high. But I don’t think it’s fair to say that this is not a situation in which the cost to the adversary is not asymmetrical. It wouldn’t be if the AI crawlers hadn’t converged towards an implementation that behaves as a DDOS botnet.

landhar··on WhatsApp introduces ads in its app
June 18, 2012 -> https://blog.whatsapp.com/why-we-don-t-sell-ads

Almost 13 years to the day!

I find it really frustrating that I am not able to avoid using whatsapp due to how popular it is to the point that it’s become the go-to communication channel for most things :/

landhar··on No Longer Posting to Pinboard
Hi, linkhut dev here.

Just wanted to call out that linkhut does have IFTTT integration [0]. Although you will need to pay for an IFTTT developer account if you want to enable that on your self hosted instance. If you need help setting that up let me know, I’ll be happy to walk you through the process (and that way I can write documentation on how to do it).

Edit: I also was under the impression that one could use the webhooks integration [1] to make bespoke integrations (as the documentation says they support passing arbitrary request headers) but haven’t tried it myself, have you tried it and run into any issues? I’d also be happy to help improving support for that workflow.

[0] https://ifttt.com/linkhut

[1] https://help.ifttt.com/hc/en-us/articles/115010230347-Webhoo...

landhar··on North Yorkshire Council to phase out apostrophe use on street signs
Do you have pointers to this? As a Spaniard I recall that even though I was originally told the spanish alphabet treated the LL as its own letter, it always felt quite inconsistent. And I always assumed it’s removal was more about simplifying things than having to do with computers
landhar··on Show HN: Free Plain-Text Bookmarking
Curious, what do you mean by “autocompletion of tags”?
landhar··on The likelihood of unilateral solar geoengineering
You’re right.. and yet I don’t believe everyone got to vote on whether we should burn all those fossil fuels in the first place.
landhar··on Windows Copilot's is showing third-party Ads to Windows users
Indeed:

> The term "soap opera" originated from radio dramas originally being sponsored by soap manufacturers.

https://en.m.wikipedia.org/wiki/Soap_opera

landhar··on The quest for a simple smartwatch
I’ve been using a Steel HR for a few years now. I’m extremely satisfied.

Long battery life is an understatement! Even when the battery drops to near 0 the watch is still useable as a plain analog watch for at least a week giving me enough time to find where I left the damn charger (given that I don’t get to use it often at all).

What I don’t get is how this model (analog watch with just enough smarts) hasn’t gained more traction. It seems to me there’s a lot of opportunities in this space, not everything requires a high resolution display.

landhar··on Mark Zuckerberg on Apple’s Vision Pro headset
You might be right that it won’t matter in the long run. But I’d argue that targeting gamers and tech enthusiasts is going to make it harder to reach mass market adoption.

For instance, with computers, laptops and cell phones being first and foremost business tools it was common in the early days for those to be provided by the employers. That really played a big role in normalizing their presence around the home before they became mainstream products.

I have a hard time seeing a similar thing happening with these goggles, but I’m far from being an expert in these things.

landhar··on Mark Zuckerberg on Apple’s Vision Pro headset
Yes, that was what I was trying to articulate. Thank you for phrasing it more clearly.
landhar··on Mark Zuckerberg on Apple’s Vision Pro headset
> If we go further back in time, mobile phones were bricks used only by business people [...]

But they were widely used by business people, and I’m assuming that was an important aspect of permeating into the other demographics.

To me, what happened with cellphones is similar to what happened with personal computers. It first entered households of people that needed them for work, and as the tech evolved, so did their recreational capabilities making them more appealing to the mass market.

AR/VR devices are attempting this the other way around: their current use cases seem to be mostly recreational, and we are hoping they will eventually come up with some productivity use cases. Someone in this thread talked about how they are surprised that tablets have sort of flopped, and I feel it’s because of similar reasons. Sure there are now more and more ways to be productive on a tablet, but it’s still more of a consume-centric experience and I believe that’s the reason they haven’t replaced laptops. Time will tell.

landhar··on Notes on Vision Pro
> could anyone have imagined walking into a living room 30 years ago and seeing everyone hunched over a phone?

True, but I don't recall smartphones ever being unveiled with promotional videos of people all being hunched over them at dinner.

Smartphones were promoted by enabling us to do things on the go: making and receiving phone calls, getting directions, finding restaurants in the area. This is why this feels different: the experience that is promoted is one that feels already slightly dystopian.

landhar··on Ask HN: Where are laid off employees gathering?
Not parent and not arguing that this model is common by any means. But a noteable example is Motion Twin [0]: a game company that had a couple of mainstream hits (most recently Dead Cells).

[0] https://en.m.wikipedia.org/wiki/Motion_Twin

landhar··on Ask HN: Bookmarking with _working_ full text search?
Shameless plug:

I’m working on an open-source social bookmarking site in Elixir that is API compatible with delicious/pinboard.

It’s named linkhut and it’s currently able to import your bookmarks from pinboard and browser exports. The flagship instance is: https://ln.ht

I’m still working on the archiving and full text search feature. I’ve been experimenting with different approaches and there’s still a few things I want to explore before settling on a solution.

I get that this is not really useful to you as of now, but if you still haven’t found anything in a couple of months, think about checking it out I might just have launched that feature by then.

landhar··on The Mystery of the Dune Font
The flourishes in the D of that cover are exactly how I was taught to write the uppercase D in cursive at school. I remember (as a kid) thinking it was odd that to write a `D` one had to start by writing an `I`, but never questioned it.

But from looking at examples of cursive on google images, it seems like that form is no longer as prominent.

landhar··on Sourcehut will blacklist the Go module mirror
But since it looks like the proxy fleet is made of a large set of nodes and each one of them only keeps track of their own state, if sourcehut were to only keep one counter for the whole user-agent it would lead to a large number of nodes not being able to refresh their cache for potentially long periods of time (there's nothing that would guarantee for example that it's always a new node that gets to make the one request per hour).

That's precisely why it's so unfair to ask the sourcehut team to try to come up with a solution to this problem: it's the very design of the proxy that Google put together that is causing this issue in the first place. And the sourcehut team has no control over their design. And to add insult to injury: when the sourcehut team offers recommendations on how to improve their implementation, Google responds with "it's too much work".

landhar··on ChatGPT: How many letters has the string “djsjcnnrjfkalcr”?
Still, it makes for a great example of the difference between GPT-3 and an AGI. We would expect the latter to have enough self-awareness to recognize when it is being asked to do something beyond its abilities.
landhar··on Chunking strings in Elixir: how difficult can it be?
Thanks for clarifying. And thanks for the article, it was a great read!
landhar··on Chunking strings in Elixir: how difficult can it be?
> During development, I noticed that the unicode_string library raises an exception when trying to get the next word if there is invalid UTF-8 anywhere in the string. That seemed weird to me, because I would expect the library to only scan the part of the string that is relevant to get the next word. If the library is checking the whole string for UTF-8 validity every time you want to get the length of the next word, that will lead to non-linear algorithmic complexity.

I wonder: could the algorithm implemented in unicode_string be improved? Or were the algorithms implemented as Rust NIFs also prone to this non-linear complexity?

landhar··on Ask HN: What Happened to Pinboard (Dec '22 edition)?
Here they are:

Linkhut’s API: https://docs.linkhut.org/overview.html

Pinboard’s API: https://pinboard.in/api/

Delicious’ API: https://github.com/domainersuitedev/delicious-api

landhar··on Ask HN: What Happened to Pinboard (Dec '22 edition)?
I’m working on an open-source social bookmarking site in Elixir that is API compatible with delicious/pinboard.

It’s named linkhut and it’s currently able to import your bookmarks from pinboard and browser exports. The flagship instance is: https://ln.ht

The source code is hosted here: https://sr.ht/~mlb/linkhut/

The documentation: https://docs.linkhut.org/introduction.html

The two things that are important to me in a bookmarking app (and why I‘ve been working on making one of my own on and off for a while): open source and offering an API for other tools to build upon.

The current API on linkhut aims to be bug for bug compatible with pinboard while being more OAuth-y.

I’m still working on an archiving feature similar to pinboard’s. But it’s probably going to take a while, I’ve been experimenting with different approaches and there’s still a few things I want to explore before settling on a solution. Once that’s ready I think I’ll call 1.0 complete.

In the meantime I’m still working through a laundry list of quality of life and small features (e.g.: recently I pushed support for “unread” links).

landhar··on Pinboard vs. Raindrop
Glad to hear! Let me know if you run into any issues either here or directly by cutting a ticket: https://todo.sr.ht/~mlb/linkhut
landhar··on Pinboard vs. Raindrop
I’m working on an open-source social bookmarking site in Elixir that is API compatible with delicious/pinboard. Its named linkhut and it’s currently able to import your bookmarks from pinboard and browser exports.

The flagship instance is: https://ln.ht

The source code is hosted here: https://sr.ht/~mlb/linkhut/

The documentation: https://docs.linkhut.org/introduction.html

Two things that I think are very important in a bookmarking app (and why I‘ve been working on making one of my own on and off for a while): open source and offering an API for other tools to build upon. The current API on linkhut aims to be bug for bug compatible with pinboard while being more OAuth-y.

I’m still working on a snapshotting feature similar to pinboard’s, once that’s ready I think I’ll call 1.0 complete. Perhaps I’ll do a Show HN then.

landhar··on What’s wrong with medieval pigs in videogames
Although I agree with the problem you’ve framed. I still think the idea of building VR experiences to learn and educate about ancient times has a lot of potential.

Since, to a great extent, the problem you’ve outlined extends to every medium. And I still believe that movies and books that try to capture the state of the art in our understanding of previous civilizations play a major role in getting people interested and to care about digging deeper into the topics.

landhar··on Ask HN: Is there a spiritual successor to del.icio.us?
This is definitely something I’ve been giving a lot of thought. I still haven’t made my mind about this but what I do know is that IF I do something like this I would:

- make sure that the data ingested in such a way is tagged in such a way that it is obvious it isn’t organic

- wait until I get a few more features that I really care about implemented (at the very least the archival and indexing of the page bookmarked)

But yeah, I agree that - for me - a huge part of the appeal of using delicious was to see the tags the community had already applied to a bookmark I would submit, and with only a handful of users at the moment we’re nowhere near having that experience :)

landhar··on Ask HN: Is there a spiritual successor to del.icio.us?
I’m working on an open-source social bookmarking site in Elixir that is API compatible with delicious/pinboard. Its named linkhut and it’s currently able to import your bookmarks from pinboard and browser exports.

The flagship instance is: https://ln.ht

The source code is hosted here: https://sr.ht/~mlb/linkhut/

The documentation: https://docs.linkhut.org/introduction.html

The one thing that I’m working on before releasing 1.0 is taking a snapshot at time of bookmark and index its contents to make it searchable (similar to pinboard’s feature).

landhar··on GitHub is adding web cookies for enterprise users
The 2020 announcement for reference: https://github.blog/2020-12-17-no-cookie-for-you/

This is very reminiscent of the “why whatsapp doesn’t sell ads” [1]. A good reminder that we should never trust any long term promises from companies.

[1] https://blog.whatsapp.com/why-we-don-t-sell-ads/?lang=en

landhar··on RESH: Rich Enhanced Shell History
One thing I’d like these to do (I tried a couple and settled for zsh-histdb with the fzf integration) is to be able to annotate past commands, with a description, tags or, even better, both.

The use case I have in mind is when I end up crafting some complicated and inscrutable incantation and being able to “earmark” it for future reference with a little bit of context that FutureMe might have forgotten by the time he thinks to reach out for the history.

Put another way, I do tend to use my shell history as a scrapbook of sorts, and I wish I was able to easily write on the margins.

Please, let me know if this is already a feature of any of these tools that I’ve completely missed.

landhar··on Facebook/Meta asks: “Wouldn't it be nice to die?”
And the irony is really strong here since the term “metaverse” was coined in the context of a dystopian version of VR.
Page 1 of 5Next →