HNHacker News
TopNewBestAskShowJobs

creesch

1,005 karma · joined November 4, 2012

submissionscomments
creesch··on German ruling declares Google liable for false answers in AI Overviews
> compared to the results in the actual Gemini.

Even those results have a lot to be desired, it is just buried deeper in the insanely verbose research report and impressive looking amount of sources you see move past.

I recently have had a close look at the various "deep research" options the big three (Anthropic, OpenAI and Google) offer. None of them are exactly transparent about how they perform searching other than the "research plan" the present upfront the and shitload of sources they show you (which, to be frank, seems to be clever UX/marketing to make it look extra legitimate and impressive). Which is already a worrying sign to me, as you can't audit the process itself properly. But even with the lack of information available on the front-end I can still see enough that worries me. A few examples:

- "Sources" are taken at face value almost no critical look at the validity of the source, the context it is placed in, etc.

- A lot of sources I know are legitimate are rarely included while a lot of listicles, low effort "reviews", etc do make the cut.

- In multiple instances when looking closer at the research plan and the "hints" they show during searching it becomes painfully clear that often enough they start with an answer in mind based on training data and try to validate that rather than actually researching the data itself.

- Subtly different prompts that by all means should still produce the same factual outcome actually provide wildly different results. This one probably relates to the other points.

In addition to all of this, I also am 100% convinced that AI powered search is incredibly expensive[1], more so than traditional search. In my mind this increased cost eventually will need to be paid by someone, which likely is going to be the user. Since the process is non-transparant I am not confident that the results will not end up being polluted by sponsored deals, etc. There is simply no way in my mind that this is going to end up well for us users.

[1] A while ago I have experimented with creating my own deep research flow with the idea that I might be able to do something with local models. To limit costs I used a SearXNG instance for searching, setup playwright for browsing sources. Using an agentic flow with agents making all the various calls and dispatching other agents ended up eating A LOT of tokens. Even when I did switch to a non agentic flow where each step is orchestrated by code calling on LLMs with simple prompts to validate results still ate a metric ton of tokens for the simplest search query. Mind you, this was not even doing actual deep research but only a few simple search queries. Ironically, google models also did seem to have more trouble coming up with good search queries compared to other models.

creesch··on Upcoming breaking changes for npm v12
> And to be fair 2: The other package repos also suck.

If you mean other languages, then yeah a lot of similar issues and weirdness there as well. Maven dependencies in any complex project are a "fun" challenge as well.

Though the sort of recurring supply chain attacks you see within the npm ecosystem is something I haven't seen elsewhere to this degree.

creesch··on Changing how we develop Ladybird
In this case they seem to be firmly closing the path though

> There will not be a separate process for submitting patches by other means. We do not want to create a shadow contribution system through issues, comments, email, or forks. External code can of course exist under the terms of the license, but we will not treat forks or patch dumps as a review queue for upstream Ladybird.

This does raise the question on how they are going to get new maintainers. The only thing I can think of is by active outreach to people contributing to adjacent projects that are still open. But that does not seem ideal to me as that will not yield people specifically interested and caring for the project you invite them to.

creesch··on Preparing for KDE Plasma's Last X11-Supported Release
> There is no support for accessibility for the visually (or otherwise) disabled in KDE Plasma's wayland extensions (and none in core wayland at all

Can you clarify what you mean by this? In the process of KDE implementing Wayland support I also have seen several issues and blog posts dedicated to accessibility features. In fact, I am fairly sure I saw KDE explicitly funding accessibility development in relation Wayland a while ago.

I am using KDE with Wayland and just had a look in my settings and the Accessibility menu is there and the features in there also appear to be working. Including the screenreader which worked on all windows I had open at the time.

Which makes sense as none of that goes through the display server but rather a D-Bus protocol implemented by Qt and GTK as far as my understanding goes.

There is a bunch of stuff that came with X11 "for free" like access easier screen capture for magnifiers, input injection, etc but as far as my understanding goes KDE (just like GNOME) has been working on DE specific implementations of each.

I am not saying things are perfect right now as far as accessibility goes. I am not someone who depends on these features. I also know that things are in fact not perfect across the board and there is still work to be done. But the claim that there is no support for accessibility seems like a rather large hyperbole to me.

creesch··on What if remote working, not AI, is to blame for weak junior hiring?
> So what ends up happening is seniors become more heads down, getting things done, and juniors struggle to get time with more experienced coworkers.

I just replied further down ( https://news.ycombinator.com/item?id=48353154 ) about this. You are entirely right, but it is also something that can largely be mitigated if companies and teams are self aware enough. I am not going to rewrite that entire comment but in addition to what I wrote there any self respecting company over a certain size should still have a junior training process in place that spans at least a year possibly two. Letting individual teams or even individuals figure out how to handle juniors always would give you wildly different results, but being in the office this was often hidden because some juniors would organically find other people for support. If you are not physically in the office you need to make sure they have other check-in moments with each other. Allow for moments where they can meet people outside their teams (knowledge presentations, workshops, etc).

I still think working hybrid (but one day per week imho is often enough) is the sweet spot for many reasons. But overall I mostly think that the FT (as often) is making excuse for things that boil down to "no, the main reason is actually corporate cost savings and refusal to invest in core processes".

creesch··on What if remote working, not AI, is to blame for weak junior hiring?
From a different perspective your sample size is just one, your team/company.

I started in a whole new team (as a senior) remotely during Covid which also contained juniors. They did incredibly well and were able to reach anyone remotely with no issue.

What might have been different is that the entire team was new and we knew we had to focus on our communication online and think about effective ways to do so. Which also benefited the juniors in the team. Many teams and companies never really gave it that much thought to begin with and I still see teams struggling to work remote at times. But, after giving them some pointers they often manage to do a whole lot better.

Some basic things out the top of my head that have benefited teams and juniors specifically:

- Have a "working together" channel where people can start meetings and where anyone can join if they feel like it. It often ends up being used by people who either like working together or those who can use some overall input on what they are working on. - Have social online moments as well. One team had a 15 minutes social block in front of standup's an other team had just a weekly social call. - Actively check in on other team members. Which feels silly to say, but the amount of times where I have seen teams only communicate during standup is also silly. Specifically juniors. If they are given a task after a little while check in with them how it is going and ask if they want to share their process. Basically how you probably in the office would walk past and also have a little conversation with them. - Take time for questions from juniors and make it clear you will do so. Whenever you are in the office and they approach you for help it means you also often serve them on the spot. Yet online I have seen juniors being ghosted for a variety of reasons. At the very least make sure to respond to juniors with a "give me 5 minutes and I'll give you a call".

To be clear, I personally like working hybrid and I do think there are benefits to coming to office at least for one day per week (assuming it is coordinated and not a ghost town). But my main point is that juniors struggling due to remote work is often more a symptom of the company not really having a good training and coaching process/culture in place more than anything else. Which I am not blaming on individual teams either. Training people is hard, people get bachelors degrees in education and then spend a lifetime getting better at educating. It's up to companies to educate their teams in this as well, offer the resources and have people on staff who solely focus on junior training.

creesch··on What we lost the last time code got cheap
> it's not just about cost reduction, it's about solving some long-term structural deficiencies of industry.

You know, I hate that this is a world where I have to ask myself if this is LLM written because it is one of those patterns.

But that is besides the point of what I wanted to say anyway. Those deficiencies aren't going to be solved by LLMs I recon. In fact, they likely will make things worse. As you said, a lot of human devs didn't understand the context when they wrote code previously. True, but LLMs are even worse at context in many areas and still need human prompting for input.

The only thing I really see happening is that the blast radius of people not fully grasping the context and still producing something is going to be larger. More specifically, it is already larger. Previously incompetence limited the damage people could do, now that is less of a factor.

creesch··on Maybe you shouldn't install new software for a bit
Except that a lot of software likely is already broken in fun ways we currently don't know about. That is what makes it such a "fun" challenge. Supply chain attacks are one thing, but CVEs in already released software allowing other attackers are another.

As always, I know most of us work in IT, but things rarely are actually binary.

creesch··on Why does it take so long to release black fan versions?
With Noctua I highly doubt that is the case given their track record for quality overall and all other information available around their design and engineering process. As far as I know based on all the information I have seen all the design and engineering is done in Austria. They also have a track record of only releasing things once they are satisfied something performs within their standards. Something that would be next to impossible when solely relying on external fabs and process engineering.

They also utilize different manufactures afaik (historically Taiwan, but also China these days) meaning they need to have pretty solid in house knowledge and expertise to make sure different factories produces similar results. When they first started utilizing Chinese factories people noticed visual differences and were worried about that. But Noctua at the time claimed that they made sure that performance was still the same. A claim that was put to the test by various review outlets at the time (I want to say gamer nexus did a big piece about it?) and confirmed to be true.

Having said that, if you do utilize external factories you automatically are making use of their process engineering to some degree as well. But, and this is difficult for many people to understand, that isn't a binary thing either. You can entirely rely on the factory to basically do everything for you and just send feedback on iterations but you can also work closely with them and actually get involved in the process itself.

creesch··on Ask HN: What are you doing during inference?
> If you’re using agents to program, what are you doing while they work?

If I am using agents I try to do something that is closely related to the task they are on. Otherwise I am just context switching once they are done and I want to review the work, which makes it difficult to focus on that task.

I also don't try to run too many agents at the same time as that is just madness. That's just herding cats at that point.

> As a side note, having Codex review Claude’s work (or vice-versa) throws up so many show-stopper issues (even with plan, revise, implement, review loops), I feel like you’d have to be nuts to just have a bunch of agents YOLOing it

Solely relying on agents is bad regardless. It certainly is the easy route and our brains are wired to take the easy/lazy approach. But even with how good models have gotten in the past year they still do make mistakes. In fact, they are now at a level where the hallucinations aren't obvious making it even more important to keep a close eye on the result.

If you do want to lean more heavily into agents doing most of the work, try to make sure they are following proper development practices. Something they don't actually do by default but using something like the superpowers skills makes a world of difference: https://github.com/obra/superpowers

Having them follow TDD helps a lot. I've even considered adding a QA agent/skill in here expanding things further to not just unit tests and some basic manual tests but also creating proper automated tests (following the test automation pyramid principles) to create an entire test suite.

Not to actually give more control to LLMs, but because it allows me to more easily spot where things go sideways. Since more tests, including e2e tests and UI tests where possible let me review more aspects of the work they do.

Having said that, I haven't created that skill yet. As reviewing the work of multiple agents is already exhausting as is. I am not in a position where I have to use AI or else in my company so instead I have dialed down my agentic usage by quite a lot to the point where I barely use agents anymore. To me using LLMs mostly as tools outside the process still is the sweet spot.

creesch··on OpenAI Privacy Filter
> The advantage of computers was that they didn't make human errors;

Sure they do, computers repeatedly, quickly, and predictably do what they are programmed to do. Which includes any human errors in that programming.

creesch··on GitHub's Fake Star Economy
> This might have to die in the era of AI,

Sadly that is probably true.

At the very least I'd add release cadence to it and the quality of releases. Mature, good software will have hotfixes and patch releases every now and then. But not in every release and certainly not 50% of the changes. In the same sense I will often look at the effort put in changelogs. If they took the effort of putting things in category, writing about possible breaking changes, etc it is a possible indicator of some level of quality. At the very least I will have a lot more faith in software with good changelogs compared to something that is just a list of the last N commit messages.

creesch··on GitHub's fake star economy
> * last commit date. Newer is better

To be honest, these days I have more faith in an application or library with a moderate development pace where maybe the last commit wasn't 2 seconds ago co-authored by claude (in the most blatant examples).

The same is true for amount of commits, the type of commits, release cadence and the amount of fixes and hotfixes in releases. I don't feel like being a glorified alpha tester so I look for maturity in a project.

Which more often than not means that, yes there needs be activity. But, it is also fine if it was two days ago and there is a clear sign of the same pattern over a longer period. Combined with a stable release cycle, sane versioning and clear changelogs that aren't just a list of the last 10 commit messages.

On your point of stars, I think they used to be a valid metric in a similar category. Namely, community behind the software. But it has been a while since that has been true. It certainly hasn't been for a while, ever since I saw these star tracking graphs pop up on repos I knew that there was no sense in paying attention to them anymore.

creesch··on Agentic AI Tools – A directory to find and compare AI agent tools
Slop spam, basic wordpress theme and the "browse tools" menu item does not work.
creesch··on Nitrile and latex gloves may cause overestimation of microplastics
> That’s not to say that there is no microplastics pollution, the U-M researchers are quick to say. > > “We may be overestimating microplastics, but there should be none. There’s still a lot out there, and that’s the problem,”
creesch··on Nitrile and latex gloves may cause overestimation of microplastics
Good news with a note:

> That’s not to say that there is no microplastics pollution, the U-M researchers are quick to say. > > “We may be overestimating microplastics, but there should be none. There’s still a lot out there, and that’s the problem,”

creesch··on CSS is DOOMed
Be honest, did you just reply to the title and the title along even skipping the other comments?
creesch··on Wine 11 rewrites how Linux runs Windows games at kernel with massive speed gains
> There has been a somewhat fast "fsync" library built around Linux's futex

The article actually goes into that in quite a bit of detail about that.

creesch··on Apideck CLI – An AI-agent interface with much lower context consumption than MCP
I mean, that is not what they are writing buddy.
creesch··on Linux is good now
AMD has very decent drivers on Linux which are even open source. It is one of the main reasons people recommend people go with AMD cards for Linux.
creesch··on Germany: Amazon is not allowed to force customers to watch ads on Prime Video
That argument ignores the reality of the current market structure. The "competition and exit" theory only works when valid alternatives actually exist.

Right now, we are dealing with effective monopolies and duopolies. You can't just exit the App Store if you have an iPhone, and Amazon has cornered the market so hard that switching isn't really an option for most people. When competition is dead, the market can't self-correct because consumers have nowhere else to go.

Also, "monetizing their own property" is doing a lot of heavy lifting in your take. These platforms already charge: transaction fees, commissions, listing fees, higher product prices baked in, and in many cases consumers are paying directly (Prime, app purchases, ride fares). Injecting ads is basically double charging. On top of that it shifts the platform from "help me find the best match" to "whoever pays the platform wins".

Honestly, unless you are in a C-suite role, I'm not sure why you would defend a model that actively works against you as a consumer.

creesch··on Germany: Amazon is not allowed to force customers to watch ads on Prime Video
It really should be illegal. Companies aren't going to do it themselves as it is a huge potential revenue stream.

So much so that it effectively has become the main focus of some companies who we as consumers still perceive as online stores/marketplaces. Specifically sponsored search results apparently can become a bigger income stream than the one from actual sales themselves.

Which is great for these companies, terrible for us consumers.

creesch··on Nano Banana Pro
Don't forget the tradition of having to migrate to a new API after a while because this one gets deprecated for "reasons". Not just a newer version, but a complete non backwards compatible new API that also requires its own setup.

To be fair, that might have changed in recent years. But after having to deal with that a few times for a few hobby projects I simply stopped trying. Can't imagine how it is for companies making use of these APIs. I guess it provides work for teams on otherwise stable applications...

creesch··on F-Droid and Google’s developer registration decree
> One could argue whether Phones with the Google android were ever really open.

In recent years, you can argue that android has no longer been open. In the early years of Android that argument would be much harder to make. To be clear, I am not talking hardcore FOSS libre open. But meaningfully open for the end user to do what they want on their device without much restriction. Early android didn't have sandboxing, had no permission system, was easy to root, etc.

Certainly with Nexus devices you had pretty much the freedom to what you wanted.

Could it have been more open? Sure, but I feel like it is almost disingenuous to say it was never if we are comparing it to the real world situation we find ourselves in today.

creesch··on How I, a non-developer, read the tutorial you, a developer, wrote for me
> LLM can theoretically interpret and use the whole project based on this info.

That's the thing though, LLMs really can't. At least not to a degree that they are able to act on it at a same level as when trained on everything else including tutorials and such.

Languages and technologies that LLMs excel at are those that are widely spread with numerous examples.

Just plain documentation with just the api calls isn't enough to train a LLM on. They effectively learn from example.

So with just #1 and no longer #1 aimed at humans you will never get to a point where you can ask an LLM about the technology.

This is what prompted me to remark that I feel you haven't thought this through. Which you might have, but that makes me think you have a overly optimistic view of what data is enough to reliably train LLMs on.

Again, to stress the point, just documentation isn't enough. So you really do need humans adapting the technology first, widening the base of examples to train on.

creesch··on How I, a non-developer, read the tutorial you, a developer, wrote for me
Given it has been a few days it might be unlikely that you read it. But I figured I'd reply anyway in case you do.

I mean this with no hostile intend, but have you honestly stopped and thought about what you did type down here?

What I mean by that is, have you looked at the complete picture to see if what you are saying makes sense in relation to what you initially said.

You questioned the need for documentation. Now you are saying there needs to be good documentation for LLMs to train on. Good documentation for LLMs to train on is actually much more extensive than than the documentation written for humans to begin with. So, you are effectively saying there needs to be more documentation.

Secondly, how can developers ask about something when they don't have decent documentation to start with.

creesch··on How I, a non-developer, read the tutorial you, a developer, wrote for me
I am surprised nobody mentioned the curse of knowledge: https://en.wikipedia.org/wiki/Curse_of_knowledge

It is actually a fairly well known phenomenon, certainly in educational circles. Being aware of it when you are writing any form of documentation is a first step. But even then, it is very difficult to properly assess the knowledge entry level of your audience.

Having others read through your documentation and importantly work with your documentation is a good strategy.

One thing I can also highly recommend is simply start out with a list of assumed prerequisite knowledge in your intro. Specifically things like certain environments, frameworks, etc. Bonus points for not only listing those but also linking to the documentation for those.

creesch··on How I, a non-developer, read the tutorial you, a developer, wrote for me
I am not sure if you thought through the implications of your proposal. LLMs are trained on examples in the training material. If something is new and isn't accessible because it lacks tangible examples the adoption rate will be lower, so there will be less training material and therefore LLMs will not be of use here.

In fact, that entire aspect of LLMs is something that is not talked about as often. But is worth a whole discussion in itself. If I remember correctly, the availability of training material for a technology already has slightly impacted more niche corners of the tech world.

creesch··on Help us raise $200k to free JavaScript from Oracle
> I don't think they're teaching Java much in university or boot camps anymore so it doesn't matter much anyway

That might just be the bubble you are in. Java is still one of the biggest languages used in corporations across the globes for anything backend related. If it is because it is a modern COBOL or because it actually is a stable language with a solid ecosystem might be a matter of some debate.

In the circles I navigate it is still heavily featured in various bootcamps.

creesch··on Pnpm has a new setting to stave off supply chain attacks
> Does the JS ecosystem really move so fast that you can’t wait a month or two before updating your packages?

Really depends on the context and where the code is being used. As others have pointed out most js packages will use semantic versioning. For the patch releases (the last of the three numbers), for code that is exposed to the outside world you generally want to apply those rather quickly. As those will contain hotfixes including those fixing CVEs.

For the major and minor releases it really depends on what sort of dependencies you are using and how stable they are.

The issue isn't really unique to the JavaScript eco system either. A bigger java project (certainly with a lot of spring related dependencies) will also see a lot of movement.

That isn't to say that some tropes about the JavaScript ecosystem being extremely volatile aren't entirely true. But in this case I do think the context is the bigger difference.

> then again, we make client side applications with essentially no networking, so security isn’t as critical for us, stability is much more important)

By its nature, most JavaScript will be network connected in some fashion in environments with plenty of bad actors.

← PreviousPage 2 of 15Next →