HNHacker News
TopNewBestAskShowJobs

MatthiasPortzel

2,206 karma · joined July 22, 2018

MatthiasPortzel.com
submissionscomments
MatthiasPortzel··on OpenAI slams court order to save all ChatGPT logs, including deleted chats
I occasionally use ChatGTP and I strongly object to the court forcing the collection of my data, in a lawsuit I am not named in, due merely to the possibility of copyright infringement. If I’m interested in petitioning the court to keep my data private, as you say is possible, how would I go about that?

Of course I haven’t sent anything actually sensitive to ChatGTP, but the use of copyright law in order to enforce a stricter surveillance regime is giving very strong “Right to Read” vibes.

> each book had a copyright monitor that reported when and where it was read, and by whom, to Central Licensing. (They used this information to catch reading pirates, but also to sell personal interest profiles to retailers.)

> It didn’t matter whether you did anything harmful—the offense was making it hard for the administrators to check on you. They assumed this meant you were doing something else forbidden, and they did not need to know what it was.

=> https://www.gnu.org/philosophy/right-to-read.en.html

MatthiasPortzel··on Watching AI drive Microsoft employees insane
Yep. I heard someone at Microsoft venting about management constantly pleading with them to use AI so that they could tell investors their employees love AI, while senior (7+ year) team members were being “randomly” fired.
MatthiasPortzel··on Ask HN: How can I load test PostgreSQL but avoid changing actual data?
> Does rollback introduce any performance overhead that would skew my results?

I would expect it to be the other way around—since the transactions are rolled back and not committed, they would have significantly less performance impact. But I’m working from an academic model of the database.

MatthiasPortzel··on US vs. Google amicus curiae brief of Y Combinator in support of plaintiffs [pdf]
> The remedy order should also prevent Google from entering into exclusive agreements to access AI training data…

Google, for example, bought exclusive access to Reddit's data. No one else can train on Reddit unless you have more money than Google (you don't). So one of the asks is that that sort of exclusive deal be prevented. If everyone is allowed to buy Reddit's data, and Google makes the best model, that wouldn't be a problem.

MatthiasPortzel··on Google Play sees 47% decline in apps since start of last year
This requirement is the result of EU regulation.
MatthiasPortzel··on Writing "/etc/hosts" breaks the Substack editor
> someone who doesn’t know or care how a system works shouldn’t be prescribing what to do to make it secure

The part that’s not said outloud is that a lot of “computer security” people aren’t concerned with understanding the system. If they were, they’d be engineers. They’re trying to secure it without understanding it.

MatthiasPortzel··on Sapphire: Rust based package manager for macOS (Homebrew replacement)
Yes, this is only a replacement for the Homebrew CLI. It doesn’t have its own package repository and moreover it doesn’t have the ability to build packages (yet)—it’s just downloading and installing the binaries built by Homebrew.
MatthiasPortzel··on WEIRD – a way to be on the web
> I want Weird to be my digital ‘soul pod’. It’s the digitized me; the source from which all of my other digital being springs forth.

I recommend working on your elevator pitch, because while strangely eloquent, that is utterly incomprehensible.

MatthiasPortzel··on Faster interpreters in Go: Catching up with C++
It's crazy that this post seems to have stumbled across an equivalent to the Copy-and-Patch technique[0] used to create a Lua interpreter faster than LuaJit[1]

[0]: https://sillycross.github.io/2023/05/12/2023-05-12/ [1]: https://sillycross.github.io/2022/11/22/2022-11-22/

The major difference is that LuaJIT Remake's Copy-and-Patch requires "manually" copying blocks of assembly code and patching values, while this post relies on the Go compiler's closures to create copies of the functions with runtime-known values.

I think there's fascinating processing being made in this area—I think in the future this technique (in some form) will be the go-to way to create new interpreted languages, and AST interpreters, switch-based bytecode VMs, and JIT compilation will be things of the past.

MatthiasPortzel··on An image of an archeologist adventurer who wears a hat and uses a bullwhip
The claim is that these models are training on data which include the problems and explanations. The fact that the first model trained after the public release of the questions (and crowdsourced answers) performs best is not a counter example, but is expected and supported by the claim.
MatthiasPortzel··on Tell HN: Announcing tomhow as a public moderator
I am firm believer in the right to free speech and the importance of expressing ideas that are contrary to the general cultural attitude.

That’s why I turn on the Show Dead setting on HN, and I love that HN has that feature.

Save your upvotes for people positively contributing.

MatthiasPortzel··on RubyLLM: A delightful Ruby way to work with AI
`binding.irb` and `show_source` have been magical in my Ruby debugging experience. `binding.irb` to trigger a breakpoint, and `show_source` will find the source code for a method name, even in generated code somehow.
MatthiasPortzel··on Mistral OCR
Slightly unrelated, but I once used Apple’s built-in OCR feature LiveText to copy a short string out of an image. It appeared to work, but I later realized it had copied “M” as U+041C (Cyrillic Capital Letter Em), causing a regex to fail to match. OCR giving identical characters is only good enough until it’s not.
MatthiasPortzel··on Apple M3 Ultra
Apple debuted dedicated machine learning hardware in 2017 with the Neural Engine on iPhones. While I don’t think they predicted the LLM explosion in particular, they knew machine learning was important and they have been allowing that to influence hardware design.
MatthiasPortzel··on An update on Mozilla's terms of use for Firefox
> I really struggle to understand what legal team believes this language is necessary in downloaded software.

Exactly. Even if nothing is changing at Mozilla, their legal team has invented a new interpretation of copyright law. That’s a huge deal from a legal perspective—Apple, Google, Microsoft, etc need to be rushing to add corresponding terms to their applications.

Mozilla PR is dropping the ball completely by trying to sweep this under the rug as ‘standard legal boilerplate’ because it’s not a clause in any other application I’ve ever seen.

Since I use FireFox at work, I don’t even have permission to give Mozilla a license to the content I create on the clock, so I will be switching browsers.

MatthiasPortzel··on Everything about Google Translate crashing React (and other web apps)
> I still won't run my own ad-blocker for this reason.

I maintained an extension for a public website for a couple years. (It did things like, for example, adding information that was available in the API to the page, for power users.) I eventually gave up with the conclusion that the concept of a browser extension was fundamentally unsound. So I also don’t use an-blocker.

MatthiasPortzel··on The OBS Project is threatening Fedora Linux with legal action
Also worth pointing out, Qt versions are supported for 6 months before EOL, unless you purchase enterprise support.
MatthiasPortzel··on The OBS Project is threatening Fedora Linux with legal action
These two comments stand out to me as inappropriate (directed at OBS).

> keeping up with runtime updates is one of the most basic expectations of a maintainer, and I suspect it's a sign there may be other problems as well.

> I won't mince words: allowing the runtime to go EOL is unacceptable and indicates terrible maintainership.

I don't use Fedora but I do use OBS… on Mac, because OBS is hands-down the most popular application for streaming on any platform. It's crazy that OBS works great on Mac, works great on Windows, works great on Linux if using the OBS Flatpak, and when the Fedora-packaged-flatpak breaks and this Fedora guy starts saying that this is indicative that "there may be other problems".

If OBS isn't good enough for Fedora to ship a working version, then show me the streaming software that is.

MatthiasPortzel··on Tell HN: Cloudflare is blocking Pale Moon and other non-mainstream browsers
Why not just ignore the bots? I have a Linode VPS, cheapest tier, and I get 1TB of network transfer a month. The bots that you're concerned about use a tiny fraction of that (<1%). I'm not behind a CDN and I've never put effort into banning at the IP level or setting up fail2ban.

I get that there might be some feeling of righteous justice that comes from removing these entries from your Nginx logs, but it also seems like there's a lot of self-induced stress that comes from monitoring failed Nginx and ssh logs.

MatthiasPortzel··on Beej's Guide to Git
> Stashes are more like commits — they even appear in the reflog! But they also aren't real commits either, in the sense that you can't check them out, or rebase them, or manipulate them directly.

They are real commits, you can check them out (it detaches your HEAD, the syntax to reference them is `stash{N}`). Although I think this furthers, rather than undermining, your point that there are an unnecessary number of other commands to work with the stash.

I think this is a failure of the git CLI as much as the internal data-structures. I think the idea of a commit tree is very good and a lot of people recognize that. The commands that git exposes to work with that tree can sometimes be miserable.

MatthiasPortzel··on Zig's comptime is bonkers good
I'm a pretty big fan of Zig--I've been following it and writing it on-and-off for a couple of years. I think that comptime has a couple of use-cases where it is very cool. Generics, initializing complex data-structures at compile-time, and target-specific code-generation are the big three where comptime shines.

However, in other situations seeing "comptime" in Zig code has makes me go "oh no" because, like Lisp macros, it's very easy to use comptime to avoid a problem that doesn't exist or wouldn't exist if you structured other parts of your code better. For example, the OP's example of iterating the fields of a struct to sum the values is unfortunately characteristic of how people use comptime in the wild--when they would often be better served by using a data-structure that is actually iterable (e.g. std.enums.EnumArray).

MatthiasPortzel··on Stimulation Clicker
Cookie Clicker has received updates almost continuously for the last decade. I don’t have the commitment to ascend but I understand there’s quite a bit of content to be unlocked there even after you’ve maxed out your first “run.”
MatthiasPortzel··on I wrote a Game Boy Advance game in Zig
Andrew’s point is that marking the memory as volatile should prevent this optimization. (i.e. the compiler isn’t allowed to split a 16-bit volatile write into two 8-but writes.) This is idiomatic because volatile is intended for MMIO.

There’s a proposal for a way to disable LLVM generating builtin calls like memcpy. But I’d argue volatile would still be more appropriate in this case.

https://github.com/ziglang/zig/issues/22110

Edit: this comment from later in the thread is also relevant:

https://news.ycombinator.com/item?id=42556624

MatthiasPortzel··on Git is bad. Here's how I'd make it better
I agree with the initial two bullet points. The git data structures are good, the git CLI is bad.

However, the issue with these attempts to create a user-friendly wrapper over the git data-structures, is that that’s what the git CLI already is. The ideal form of git that you explain to beginners is `git switch -c branch-name`, `git commit`, `git push`. This is not any more complicated than OP’s `sg create`, `sg save`, `sg submit`.

The issue with the git CLI is that it doesn’t expose commands for working with raw git data structures. This prohibits people developing an understanding of what the git commands are actually doing.

I’m at the point where I understand git well enough that I’m very rarely in the situation that OP describes, of guessing whether a git command will work, or asking ChatGTP for help. However I’ve gotten here by slowly memorizing what every git CLI command does to the underlying git data structures.

`git commit` creates a commit and updates the current branch pointer. `git commit --amend` created a new commit with a parent that is the commit before the commit at the HEAD pointer and updates the branch pointer. `git reset --hard` updates the current branch pointer and the files to match. `git push` runs `git fetch` and then `git merge`. `git merge` creates a new commit with two (or more) parent commits and then updates both branch pointers (if applicable) to point to the new commit. etc. etc. etc.

So I hate the git CLI because it tries to be beginner friendly by supporting a ‘normal coding workflow’ and it ends up being more complicated than just understanding the raw git data-structures.

MatthiasPortzel··on Tog's Paradox
It can also be a result of the XY problem. Person wants to do Y, they imagine a software that does part of the hardest parts of Y (call that part X), and they commission software that does X. They then commission a ton of other small parts of Y to be added to the software of the course of years. Whereas an all-inclusive software to do Y from the beginning would have been simpler.

This issue can be avoided by product leads with vision for the entire problem.

MatthiasPortzel··on CRLF is obsolete and should be abolished
They acted on these words, updating their HTTP server to serve just \n.

=> https://sqlite.org/althttpd/info/8d917cb10df3ad28 Send bare \n instead of \r\n for all HTTP reply headers.

While browser aren't effected, this broke compatibility with at least Zig's HTTP client.

=> https://github.com/ziglang/zig/issues/21674 zig fetch does not work with sqlite.org

MatthiasPortzel··on Internet Archive: Security breach alert
Cloudflare isn't even that big. They're 1/100th the size of Google or MS. They're not even the biggest CDN—Akamai has twice the revenue, but it depends on what you measure. Cloudflare gets brought up disproportionately often on HN because they have generous free tiers and cater to indie hackers more. So it feels a little ironic that they're perceived as "the big dog" by the indie hackers.
MatthiasPortzel··on The Remarkable Life of Ibelin
Why the title change?
MatthiasPortzel··on The best browser bookmarking system is files
On paper, tagging is objectively better for the reasons you describe. But in my experience, the human brain has an intuition for location and object-permanence which is confused by having the same thing in multiple places.
MatthiasPortzel··on How Discord stores trillions of messages (2023)
The other reply goes to airplanes but there are much more common ways to get disconnected. Locking my phone or closing my laptop lid disconnects me from IRC. A lot of Discord users have desktops that are always on (since Discord originally advertised to gamers), but a lot of Discord users don’t.

Discord is fundamentally a very versatile platform. If you lose one seemingly unimportant, you lose a lot of versatility. Maybe I’ll write a blog post just with examples of how I’ve used it. It replaces IRC, but it also replaces Facebook groups, Skype, a lot of group texts, and a lot of email for me.

← PreviousPage 2 of 16Next →