HNHacker News
TopNewBestAskShowJobs

curioussquirrel

243 karma · joined September 5, 2025

Multilingual LLM evals
submissionscomments
curioussquirrel··on We're losing our voice to LLMs
+1 on this one! I only use LLMs once I'm done with writing, and basically using them as my editor.

In case it helps anyone, here is my prompt:

"You are a professional writer and editor with many years of experience. Your task is to provide writing feedback, point out issues and suggest corrections. You do not use flattery. You are matter of fact. You don't completely rewrite the text unless it is absolutely necessary - instead you try to retain the original voice and style. You focus on grammar, flow and naturalness. You are welcome to provide advice changing the content, but only do that in important cases.

If the text is longer, you provide your feedback in chunks by paragraph or other logical elements.

Do not provide false praise, be honest and feel free to point out any issues."

(Yes, you kind of need to repeat you're actively not looking for a pat on the back, otherwise it keeps telling you how brilliant your writing is instead of giving useful advice.)

curioussquirrel··on GitHub: Git operation failures
Same, even started adding new ssh keys to no avail... (I was getting some nondescript user error first, then unhealthy upstream)
curioussquirrel··on Beets: The music geek’s media organizer
A very similar workflow on my end, both beets as the main tagger/organizer and Picard to pick up whatever can't be processed through beets. Beets is amazing!
curioussquirrel··on Wharton AI 2025 Adoption Report
Surprised this did not get any attention whatsoever. Some really surprising findings in it: 75% of firms already have a positive return on investment from AI, less than 5% negative return. Also 46% of businesses leaders now use AI daily themselves.
curioussquirrel··on Samsung Family Hub for 2025 Update Elevates the Smart Home Ecosystem
Although I loathe ads, I think that for new products where the presence of ads is disclosed clearly upfront, this is acceptable. Especially if this comes with a discount. We have Kindles with and without ads and people are generally fine with it.

But the fact that this gets retrofitted to fridges that people already bought, without any way of opting out or other mitigation, is criminal. Is this a lawsuit in the making? Am I naive?

curioussquirrel··on You can't turn off Copilot in the web versions of Word, Excel, or PowerPoint
Azure copilot is really something. It can't see the context of the page it's embedded in, and the message you send is limited to 500 characters, so good luck pasting a log or configuration.
curioussquirrel··on Affinity Studio now free
Photoshop now has a bunch of features that get used in professional environments. And in the end user space, facial recognition or magic eraser are features in apps like Google Photos that people actively use and like. People probably don't care that it's AI under the hood, in fact they probably don't even realize.

There is a lot of unchecked hype, but that doesn't mean there is no substance.

curioussquirrel··on It's insulting to read AI-generated blog posts
100% agree. Using it to polish your sentences or fix small grammar/syntax issues is a great use case in my opinion. I specifically ask it not to completely rewrite or change my voice.

It can also double as a peer reviewer and point out potential counterarguments, so you can address them upfront.

curioussquirrel··on Polish top-performing language for complex AI tasks, finds study
Interesting experiment, but I'd say aggregating the scores across models is far from ideal. Gemini 1.5 Flash got close-to-perfect scores on most languages (probably boils down to small variances in temp/top_k and statistical error). Small models are generally quite bad at non-English languages and tank the overall performance.

BTW, newer generations of models seem to have made some real progress in multilingual performance.

curioussquirrel··on Google flags Immich sites as dangerous
Looking forward to Louis Rossmann's reaction. Wouldn't be surprised if this leads to a lawsuit over monopolistic behavior - this is clearly abusing their dominant position in the browser space to eliminate competitors in photos sharing.
curioussquirrel··on Dane Stuckey (OpenAI CISO) on Prompt Injection Risks for ChatGPT Atlas
I wonder how much more telemetry and behavioral data Atlas needs to collect on its users given the rapid response system. And how well is all the session data stripped of sensitive information when transferred.
curioussquirrel··on LLMs are getting better at character-level text manipulation
You're right. There was also an experiment in Meta which tokenized bytes directly and it didn't hurt performance much in very small models.
curioussquirrel··on LLMs are getting better at character-level text manipulation
True. But does that scale to less common words? Or to other languages than English?
curioussquirrel··on LLMs are getting better at character-level text manipulation
This is a very good answer and I'm commenting only to bring more attention to it apart from voting up. Well put!
curioussquirrel··on LLMs are getting better at character-level text manipulation
Yes, but it would hurt its contextual understanding and effectively reduce the context window several times.
curioussquirrel··on LLMs are getting better at character-level text manipulation
Thanks for the explanation and for the tokenizer playground link!
curioussquirrel··on LLMs are getting better at character-level text manipulation
Why test for something? I find it fascinating if something starts being good at task it is "explicitly not designed for" (which I don't necessarily agree with - it's more of a side effect of their architecture).

I also don't agree that nobody is using this for - there are real life use cases today, such as people trying to find meaning of misspelled words.

On a side note, I remember testing Claude 3.7 with the classic "R's in the word strawberry" question through their chat interface, and given that it's really good at tool calls, it actually created a website to a) count it with JavaScript, b) visualize it on a page. Other models I tested for the blog post were also giving me python code for solving the issue. This is definitely already a thing and it works well for some isolated problems.

curioussquirrel··on LLMs are getting better at character-level text manipulation
Even GPT 3.5 is okay (but far from great) at Base64, especially shorter sequences of English or JSON data. Newer models might be post-trained on Base64-specific data, but I don't believe it was the case for 3.5. My guess is that as you say, given the abundance of examples on the internet, it became one of the emergent capabilities, in spite of its design.
curioussquirrel··on LLMs are getting better at character-level text manipulation
Yep, there is still a room for improvement, but my point is that the LLMs are getting better at something they're "not supposed to be able to do".

Quartiles sound like an especially brutal game for an LLM, though! Thanks for sharing

curioussquirrel··on LLMs are getting better at character-level text manipulation
Thanks, Simon! I saw the same approach (numbering the individual characters) in GPT 4.1's answer, but not anymore in GPT 5's. It would be an interesting convergence if the models from Anthropic and OpenAI learned to do this at a similar time, especially given they're (reportedly) very different architecturally.
curioussquirrel··on Neutts-air – Open-source, on device TTS
Will check it out, thx!
curioussquirrel··on Neutts-air – Open-source, on device TTS
Could we finally get a decent opensource TTS app for Android? This project is very cool.
curioussquirrel··on I only use Google Sheets
This is true, but sadly for very large spreadsheets, it kind of stops working at a certain point. You get timeouts when loading cell history or the version of the spreadsheet version history.
curioussquirrel··on Why Tech Workers Don't Trust AI
That is only true for employees, ignoring anything but work circumstances and also reducing returns purely to take home salary. I've personally saved a lot of time because of LLMs already, and that's worth quite a lot to me.
curioussquirrel··on In Praise of RSS and Controlled Feeds of Information
I admit I do have a little bit of old web nostalgia, but I know that we live in a very different world now and web platforms and the types of interactions we have online are much more complex now.

That said, specifically for the type of information you'd typically subscribe to in an RSS reader, I still think the web 1.0 approach has its place. I do believe that if you have something to say, standing up a blog with an RSS feed, writing posts there and then potentially linking on social media is the best way. It's also been trivially easy to set up a blog for years.

Likewise with news - I don't think there's anything better to get them than reading a news site or perhaps subscribing to a video channel. All very RSS-friendly.

I'm mostly interested in longer form content, not shorter or ephemeral types of content. And even there, platforms like Mastodon support RSS feeds natively.

But yeah, as we clarified below, the angle for writing the article was the delivery channel for content and whether it is curated by someone or not. For the quality and provenance of the information, that's still on you and that's a very hard problem with no clear solution - and arguably one that will get better with more AI generated content.

curioussquirrel··on In Praise of RSS and Controlled Feeds of Information
Understood! I know the onus for how the words are interpreted are primarily on the author, so that's all fair for you to raise and I'm glad for that feedback.
curioussquirrel··on In Praise of RSS and Controlled Feeds of Information
I may need to talk to my hosting provider :) Thanks for pointing this out. The site is indeed statically generated (Hugo), so this should not be happening.
curioussquirrel··on In Praise of RSS and Controlled Feeds of Information
This is very much true, and one of the downsides of RSS is that you need to make effort to discover new sources, or make sure that what you're consuming is at least somewhat balanced. However you have no guarantees of the latter when you use algorithmic feeds.
curioussquirrel··on In Praise of RSS and Controlled Feeds of Information
I see now - your issue is with the "controlled feeds of information" part? I am not claiming they are "feeds of controlled information" (which is how you seem to be interpreting it). Of course, all the sources you subscribe to will have their own biases and issues, but you do not lose agency over what you select for consumption. That is the control I am seeking and what I like about RSS.
curioussquirrel··on In Praise of RSS and Controlled Feeds of Information
Author here. I do not necessarily think algorithmic feeds are the only thing wrong with social media, but it's certainly one of the major problems. More so if the platforms don't even allow me to revert to chronological feeds, or make it really user unfriendly.

Of course the cultural context has changed, but I think your view is quite cynical. I do believe that AI could, in theory, be a good steward and curator of news feeds (think Google News), but I haven't seen an implementation that would be open and customizable enough. I do not like the idea that someone could be manipulating what I'm being presented, or what reaches me and what doesn't.

Could you elaborate on why you think this is misguided tech-nostalgia? Most of your arguments seem to be true regardless of how you discover content (RSS, social media, link aggregators, ...)

← PreviousPage 2 of 3Next →