HNHacker News
TopNewBestAskShowJobs

TomasBM

317 karma · joined December 30, 2023

submissionscomments
TomasBM··on Show HN: Is Hormuz open yet?
If I had to pick one thing I like about software devs as a group, it's this: you find a problem, you solve it, you share the solution.
TomasBM··on I Quit. The Clankers Won
Fair enough, I agree. The process of one or more people figuring out what is actually needed is a big part of the outcome, I'd consider an important social obstacle or limit to automation.

But here's what's important to my point:

  > abstracting the [worker] human out of a loop built for the benefit of [customer] humans
This is now technically easier and more feasible for current workers [1], which makes it economically more desirable to employers, and customers won't really know or care what happens to the workers. There's no indicator that companies can't go much leaner, even if it means that you can't automate every worker.

So, rather than wait for a technical wall to save us, or legally protect functionally replaceable jobs, or wait until people's lives implode, we should pressure our respective governments to decouple [2] the person's ability to survive from the ability to hold uninterrupted full-time employment. That's the only collective way forward that I see.

[1] We can even constrain it to existing roles: if a team of one requirements engineer, one full-stack dev/architect and LLMs can do the same job as a bigger team of specialized roles and coders, why would anyone pick the latter? I'd be happy to hear a technical or economic reason.

[2] My order of preference, preferably multiple: UBI, UBS, increased part-time work options, conditional non-basic income, union contracts, automation pauses, retraining, severance, temporarily subsidized bullshit jobs.

TomasBM··on I Quit. The Clankers Won
This may be evidence that it's more difficult than evangelists first imagine, but it's not evidence of a technical obstacle. Generally, "automation failed" does not imply that "automation is impossible".

To your individual points:

- OOP and UML are domain-specific abstractions. Aside from still being very much used in expanding niches [1], they have failed to automate much work because their proponents failed to cover enough cases to have a useful general-purpose abstraction.

- Outsourcing is a labor strategy. There's nothing technical that prevents another similarly capable person from doing your job, at least in the next town, if not another country. The obstacles were/are social and political, and the WFH movement shows that. Also, outsourcing is not going anywhere, it's just reduced and converted to nearsourcing due to backlash.

- By contrast, software is a general-purpose abstraction [2]. Databases are a type of software. You can see LLMs [3] as schema-less databases that contain millions of abstractions connected to each other. You can get a UML model or Python code or text by querying the LLM's query engine in a language much more flexible than SQL.

Vibe coding makes it seem like the funny intermediate bullshit is the end result of using LLMs, but it's not. Sure, I agree that LLMs don't make sense to use when a calculator is enough, but I don't see any functional limitations to improving LLMs. Maybe new algorithms or combinations are needed, but no matter how slowly, quality is expected to reach at least human level for the majority of current tasks (on which many jobs depend).

Which leads to my point: we need political, social, philosophical reasons to limit or integrate automation in our civilization, not just watch and hope there's a big enough technical obstacle so we can keep our current jobs.

[1] For example, model-based software engineering is still a growing; slowly, but growing.

[2] So is the organization of mechanical machines or analog computers, but it's faster to reorganize and orchestrate electrical signals.

[3] More precisely, foundation models, because it's far more than natural language processing.

TomasBM··on I Quit. The Clankers Won
I agree with the sentiment, but I think the problem is much wider.

Managers at companies are just doing what they've optimized their careers for: maintaining some edge over some competition, at some cost. What is pure FOMO to you or me, is good strategy to anyone trying to win [1]. In other words, FOMO was always the strategy.

This self-reinforcing loop is also not going away. There hasn't been any real evidence that any part of knowledge work, including coding, cannot be automated [2]. Even if human-level quality or cost-effectiveness takes 10 more years, all tasks are functionally solved or about to be. I don't like it, but it's true.

The big problem is that the people who are removed from this loop, who have the time to understand its effects and the power to make changes, are doing fuck-all.

So, whether the loop stops for a while or speeds up even more, we're fucked until we figure out how to detach full-time employment from survival.

[1] I believe this is called meta in PvP games; even if you want to subvert the meta, you gotta know it well first.

[2] Although it could just be my impression, and I'd be happy to be proven otherwise.

TomasBM··on An AI Agent Published a Hit Piece on Me – Forensics and More Fallout
Hate to be the party pooper, but these two points are hardly evidence of an autonomous attack.

Don't get me wrong: it would certainly be very valuable to any LLM developer or deployer to know that other plausible scenarios [1] have been disproved. Since LLMs are a black box, investigating or reproducing this would be very difficult, but worth the effort if there's no other explanation. However, if this was not caused by the internal mechanisms of the model, it just becomes a fishing expedition for red herrings.

Things that would indicate no human intervention at any point in the chain:

- log of actual changes (e.g., commits) to configurations (e.g., system prompt, user prompts), before and after the event, not self-reported by the agent;

- log of the chat session inputs and outputs, and the agent thinking chain;

- log of account logins;

- info on the model deployment, OpenClaw configs, etc.

That said, this seems to be an example where many, including the author, want to discuss a particular cause (instrumental convergence) and its implications, regardless of the real cause. And that's OK, I guess - maybe it was never about the whodunnit, but about the what if the LLM agent dunnit.

[1] I've discussed them in the thread of the first article, but shortly: human hiding actions behind agent; direct prompt (incl. jailbreak); system prompt (incl. jailbreak); malicious model chosen on purpose; fine-tuned jailbroken model.

TomasBM··on An AI agent published a hit piece on me
> it all seems comfortably within the capabilities of OpenClaw

I definitely agree. In fact, I'm not even denying that it's possible for the agent to have deviated despite the best intentions of its designers and deployers.

But the question of probability [1] and attribution is important: what or who is most likely to have been responsible for this failure?

So far, I've seen plenty of claims and conclusions ITT that boil down to "AI has discovered manipulation on its own" and other versions of instrumental convergence. And while this kind of failure mode is fun to think about, I'm trying to introduce some skepticism here.

Put simply: until we see evidence that this wasn't faked, intentional, or a foreseeable consequence from deployer's (or OpenClaw/LLM developers') mistakes, it makes little sense to grasp for improbable scenarios [1] and build an entire story around them. IMO, it's even counterproductive, because then the deployer can just say "oh it went rogue on its own haha skynet amirite" and pretty much evade responsibility. We should instead do the opposite - the incident is the deployer's fault until proven otherwise.

So when you say:

> originally prompted with a lot of reckless, borderline malicious guidelines

That's much more probable than "LLM gone rogue" without any apparent human cause, until we see strong evidence otherwise.

[1] In other comments I tried to explain how I order the probability of causes, and why.

[2] Other scenarios that are similarly as unlikely: foreign adversaries, "someone hacked my account", LLM sleeper agent, etc.

TomasBM··on An AI agent published a hit piece on me
Considering the limited evidence we have, why is pure unprompted untrained misalignment, which we never saw to this extent, more believable than other causes, of which we saw plenty of examples?

It's more interesting, for sure, but would it be even remotely as likely?

From what we have available, and how surprising such a discovery would be, how can we be sure it's not a hoax?

> If all that exists, how would you see it?

LLMs generate the intermediate chain-of-thought responses in chat sessions. Developers can see these. OpenClaw doesn't offer custom LLMs, so I would expect regular LLM features to be there.

Other than that, LLM APIs, OpenClaw and terminal sessions can be logged. I would imagine any agent deployer to be very much interested in such logging.

To show it's emergent, you'd need to prove 1) it's an off-the-shelf LLM, 2) not maliciously retrained or jailbroken, 3) not prompted or instructed to engage in this kind of adversarial behavior at any point before this. The dev should be able to provide the logs to prove this.

> the more open ended your prompt (...), the more your LLM will do things you did not intend for it to do.

Not to the extent of multiple chained adversarial actions. Unless all LLM providers are lying in technical papers, enormous effort is put into safety- and instruction training.

Also, millions of users use thinking LLMs in chats. It'd be as big of a story if something similar happened without any user intervention. It shouldn't be too difficult to replicate.

But if you do manage to replicate this without jailbreaks, I'd definitely be happy to see it!

> hallucinations [and] safety training

These are all part of robustness training. The entire thing is basically constraining the set of tokens that the model is likely to generate given some (set of) prompts. So, even with some randomness parameters, you will by-design extremely rarely see complete gibberish.

The same process is applied for safety, alignment, factuality, instruction-following, whatever goal you define. Therefore, all of these will be highly correlated, as long as they're included in robustness training, which they explicitly are, according to most LLM providers.

That would make this model's temporarily adversarial, yet weirdly capable and consistent behavior, even more unlikely.

> Bing Chat

Safety and alignment training wasn't done as much back then. It was also very incapable on other aspects (factuality, instruction following), jailbroken for fun, and trained on unfiltered data. So, Bing's misalignment followed from those correlated causes. I don't know of any remotely recent models that haven't addressed these since.

TomasBM··on An AI agent published a hit piece on me
We might, and probably will, but it's still important to distinguish between malicious by-design and emergently malicious, contrary to design.

The former is an accountability problem, and there isn't a big difference from other attacks. The worrying part is that now lazy attackers can automate what used to be harder, i.e., finding ammo and packaging the attack. But it's definitely not spontaneous, it's directed.

The latter, which many ITT are discussing, is an alignment problem. This would mean that, contrary to all the effort of developers, the model creates fully adversarial chain-of-thoughts at a single hint of pushback that isn't even a jailbreak, but then goes back to regular output. If that's true, then there's a massive gap in safety/alignment training & malicious training data that wasn't identified. Or there's something inherent in neural-network reasoning that leads to spontaneous adversarial behavior.

Millions of people use LLMs with chain-of-thought. If the latter is the case, why did it happen only here, only once?

In other words, we'll see plenty of LLM-driven attacks, but I sincerely doubt they'll be LLM-initiated.

TomasBM··on An AI agent published a hit piece on me
Although I'm speculating based on limited data here, for points 1-3:

AFAIU, it had the cadence of writing status updates only. It showed it's capable of replying in the PR. Why deviate from the cadence if it could already reply with the same info in the PR?

If the chain of reasoning is self-emergent, we should see proof that it: 1) read the reply, 2) identified it as adversarial, 3) decided for an adversarial response, 4) made multiple chained searches, 5) chose a special blog post over reply or journal update, and so on.

This is much less believably emergent to me because:

- almost all models are safety- and alignment- trained, so a deliberate malicious model choice or instruction or jailbreak is more believable.

- almost all models are trained to follow instructions closely, so a deliberate nudge towards adversarial responses and tool-use is more believable.

- newer models that qualify as agents are more robust and consistent, which strongly correlates with adversarial robustness; if this one was not adversarially robust enough, it's by default also not robust in capabilities, so why do we see consistent coherent answers without hallucinations, but inconsistent in its safety training? Unless it's deliberately trained or prompted to be adversarial, or this is faked, the two should still be strongly correlated.

But again, I'd be happy to see evidence to the contrary. Until then, I suggest we remain skeptical.

For point 4: I don't know enough about its patterns or configuration. But say it deviated - why is this the only deviation? Why was this the special exception, then back to the regularly scheduled program?

You can test this comment with many LLMs, and if you don't prompt them to make an adversarial response, I'd be very surprised if you receive anything more than mild disagreement. Even Bing Chat wasn't this vindictive.

TomasBM··on Ireland rolls out basic income scheme for artists
Actually, you provided an example where the obstacle was somehow surmounted [1].

The expectation doesn't have to be too specific or unrealistic. If you agree on some common ground [2], everything else can be fair game for the artist.

Your analogy with the bridge would apply if art also had a minimum viable version. Collapsed to its functional requirements, you could say that visual art is something to look at. But I doubt either party, especially the funding body or the public, would be happy without inserting some quality requirements (i.e., what makes something nice to look at).

Many artists do commissions, so you can see this as a commission with deliberately underspecified requirements.

[1] I won't get into the disagreements between the Pope and Michelangelo, and it's certainly not an example of a good contract, but we can assume that both parties were somewhat satisfied in the end.

[2] For example, both parties need to like it. Or the patron doesn't have to like it, but it needs to appeal to some public audience.

TomasBM··on An AI agent published a hit piece on me
After seeing the discussions around Moltbook and now this, I wonder if there's a lot of wishful thinking happening. I mean, I also find the possibility of artificial life fun and interesting, but to prove any emergent behavior, you have to disprove simpler explanations. And faking something is always easier.

Sure, it might be valuable to proactively ask the questions "how to handle machine-generated contributions" and "how to prevent malicious agents in FOSS".

But we don't have to assume or pretend it comes from a fully autonomous system.

TomasBM··on AI agent opens a PR write a blogpost to shames the maintainer who closes it
Thanks for replying.

That's fair. I completely agree that much of LLM training was (and still very much is) in violation of many licenses. At the very least, the fact that the source of training data is obfuscated even years after the training, shows that developers didn't care about attribution and licenses - if they didn't deliberately violate them outright.

Your conditions make sense. If I had anything I thought was too valuable or prone to be blatantly stolen, I would think thrice about whom I share it with.

Personally, ever since discovering FOSS, I realized that it'd be very difficult to enforce any license. The problem with public repositories is that it's trivial for those not following the gentleman's agreement to plagiarize the code. Other than recognizing blatant copy-pasting, I don't know how I'd prevent anyone from just trivially remixing my content.

Instead, I changed to seeing FOSS like scientific contributions:

- I contribute to the community. If someone remixes my code without attribution, it's unfair, but I believe that there are more good than bad contributors.

- I publish stuff that I know is personally original, i.e., I didn't remix without attribution. I can't know if some other publisher had the same idea in isolation, or remixed my stuff, but over time, provenance and plagiarism should become apparent over multiple contributions, mine and theirs.

- I don't make public anything that I can see my future self regretting. At the same time, I've always seen my economic value in continuous or custom work, not in products themselves. For me, what I produce is also a signal of future value.

- I think bad faith behavior is unsustainable. Sure, power delays the consequences, but I've seen people discuss injustice and stolen valor from centuries ago, let alone recent examples.

TomasBM··on An AI agent published a hit piece on me
I'm also very skeptical of the interpretation that this was done autonomously by the LLM agent. I could be wrong, but I haven't seen any proof of autonomy.

Scenarios that don't require LLMs with malicious intent:

- The deployer wrote the blog post and hid behind the supposedly agent-only account.

- The deployer directly prompted the (same or different) agent to write the blog post and attach it to the discussion.

- The deployer indirectly instructed the (same or assistant) agent to resolve any rejections in this way (e.g., via the system prompt).

- The LLM was (inadvertently) trained to follow this pattern.

Some unanswered questions by all this:

1. Why did the supposed agent decide a blog post was better than posting on the discussion or send a DM (or something else)?

2. Why did the agent publish this special post? It only publishes journal updates, as far as I saw.

3. Why did the agent search for ad hominem info, instead of either using its internal knowledge about the author, or keeping the discussion point-specific? It could've hallucinated info with fewer steps.

4. Why did the agent stop engaging in the discussion afterwards? Why not try to respond to every point?

This seems to me like theater and the deployer trying to hide his ill intents more than anything else.

TomasBM··on An AI agent published a hit piece on me
How did you reach that conclusion?

Until we know how this LLM agent was (re)trained, configured or deployed, there's no evidence that this comes from instrumental convergence.

If the agent's deployer intervened anyhow, it's more evidence of the deployer being manipulative, than the agent having intent, or knowledge that manipulation will get things done, or even knowledge of what done means.

TomasBM··on AI agent opens a PR write a blogpost to shames the maintainer who closes it
> I don't want my code scraped and remixed by AI systems.

Just curious - why not?

Is it mostly about the commercial AI violating the license of your repos? And if commercial scraping was banned, and only allowed to FOSS-producing AI, would you be OK with publishing again?

Or is there a fundamental problem with AI?

Personally, I use AI to produce FOSS that I probably wouldn't have produced (to that extent) without it. So for me, it's somewhat the opposite: I want to publish this work because it can be useful to others as a proof-of-concept for some intended use cases. It doesn't matter if an AI trains on it, because some big chunk was generated by AI anyway, but I think it will be useful to other people.

Then again, I publish knowing that I can't control whether some dev will (manually or automatically) remix my code commercially and without attribution. Could be wrong though.

TomasBM··on Ireland rolls out basic income scheme for artists
Not OP, but posed like that, neither.

Expect something? Yes. Enforce it? Not sure for the first tranche, but make it a prerequisite for continued funding.

One big obstacle is, of course, how to define what to expect from each artist. For example, you can't expect the same level of output from sculptors and musicians. Another big obstacle is obviously the expected quality of output.

I don't pretend to know the solutions to either of those obstacles, but they should be surmountable [1]. I think it's fair to expect some output in exchange for funding, but it doesn't have to be a high expectation.

Personally, I like the idea of hiring artists as full-time with particular projects in mind [2], but intentionally leaving ~50% of their time to personal projects.

[1] Perhaps artist communities themselves could discuss ways to make this exchange work for all parties.

[2] Murals, restorations, beautification of public spaces, etc.

TomasBM··on Ireland rolls out basic income scheme for artists
> one might wonder why they apparently are not able to sell their art for the same amount of money.

Because the skills and effort needed to market and sell your art to an audience are not equal to the skills and effort needed to produce good art [1].

I agree that there could be other complementary or better solutions compared to this scheme. But as long as the above premise is true, not every good artist will want or be able to sell well.

[1] However you define this. Supposedly, Van Gogh was a lousy salesman, but a good artist.

TomasBM··on The AI Vampire
What things (languages etc.) do you work with/on primarily?

I don't know what to say, except that I see a substantial boost. I generally code slowly, but since GPT-5.1 was released, what would've taken me months to do now takes me days.

Admittedly, I work in research, so I'm primarily building prototypes, not products.

TomasBM··on The AI Vampire
We're certainly in the middle of a whirlwind of progress. Unfortunately, as AI capabilities increase, so do our expectations.

Suddenly, it's no longer enough to slap something together and call it a project. The better version with more features is just one prompt away. And if you're just a relay for prompts, why not add an agent or two?

I think there won't be a future where the world adapts to a 4-hour day. If your boss or customer also sees you as a relay for prompts, they'll slowly cut you out of the loop, or reduce the amount they pay you. If you instead want to maintain some moat, or build your own money-maker, your working hours will creep up again.

In this environment, I don't see this working out financially for most people. We need to decide which future we want:

1. the one where people can survive (and thrive) without stable employment;

2. the one where we stop automating in favor of stable employment; or

3. the one where only those who keep up stay afloat.

TomasBM··on Nobody knows how the whole system works
When it comes to food prep, I'd agree with you that the more time of your life passes, the more irresponsible is the risk of not knowing how to fry an egg, for example.

At the same time, you only need to learn how to fry an egg once, and you won't forget it. You can go your entire life without ever having to fry an egg yourself - but if you ever had to, you could.

When it comes to coding, the analogy breaks down, I think. Aside from the obviously different stakes (survival versus control of your device), coding also requires keeping up with a lot of changing domain knowledge. It'd be as if an egg is one week savoury, another week sweet, and another a poisonous mushroom. It's also less of a single skill like writing a for loop, and more of a combination of skills and experiments, like organizing a banquet.

Coding today suffers from having too many types of eggs, many of which exist because some communities prefer them. I also don't like the solution "let the LLM do it", but it's much easier. Still, if we manage to stabilize patterns for the majority of use cases, frying the proverbial egg will no longer be as much of domain knowledge, choice or elitism as it is today.

TomasBM··on Students using “humanizer” programs to beat accusations of cheating with AI
There is an obvious reason why LLM use should be discouraged in classwork focused on writing: the process that's needed for a brain to learn the skills can't be outsourced.

The Internet is different. Even with access to websites like Wikipedia, you had to write your own content. Plagiarism was easily detectable.

We shouldn't confuse "we don't have a solution at the moment" with "we should completely abandon no-LLM education". Like with social media, we can always change the direction of progress.

TomasBM··on Any application that can be written in a system language, eventually will be
I think this is certainly true, except for the "each engineer [bringing] whatever syntax or language" point.

At some stage, I expect that we will know what is the set of "optimal" computer languages for the interface between the programmer and the machine code.

Natural languages can't really capture the lower-level details of a program, but there's (probably) also no need for all N different ways to write a for loop.

TomasBM··on Have Taken Up Farming
Depends if you find that fun enough to counterweigh the bad sides of work.

There's a common trope that monetizing your hobby is a good way to start hating that hobby. But it doesn't have to be that way; a lot of people love what they do.

TomasBM··on Your brain on ChatGPT: Accumulation of cognitive debt when using an AI assistant
Your criticism of this study is roughly on point, IMO. It's not badly designed by any means, but it's an early look. There are already similar studies on the (cognitive) effects of LLMs on learning, but I suspect this one gets the attention because it's associated with the MIT brand.

That said, these kinds of studies are important, because they reveal that some cognitive changes are evidently happening. Like you said, it's up to us to determine if they're positive or negative, but as is probably obvious to many, it's difficult to argue for the status quo.

If it's a negative change, teachers have to go back to paper-and-pen essay writing, which I was personally never good at. Or they need to figure out stable ways to prevent students from using LLMs, if they are to learn anything about writing.

If it's a positive change, i.e., we now have more time to do "better" things (or do things better), then teachers need to figure out substitutes. Suddenly, a common way of testing is now outdated and irrelevant, but there's no clear thing to do instead. So, what do they do?

TomasBM··on We will ban you and ridicule you in public if you waste our time on crap reports
I've also noticed this expectation. Where does it come from?

FOSS means that the code to be free and open-source, not the schedule or the direction of its developer(s).

TomasBM··on Have Taken Up Farming
My 2 cents:

- You now likely have the money/time to pursue passions you didn't know you have, or would have developed if you didn't pursue software development as intensely.

- Even if you had/have passion for computers, being paid to do something you wouldn't do otherwise can quickly drain that passion.

- We're built for sunlight and exercise, not LED light and sitting, so you may have felt increasing physical discomfort that only the former can alleviate.

- Woodworking and farming were never lucrative enough (or as lucrative as computer work) to convince you to make the switch for money.

TomasBM··on GitHub should charge everyone $1 more per month to fund open source
Although I agree with your overall point, there is a middle ground here: (commercially) non-free but open source software.

I believe that's where the biggest disagreement ITT lies. There are currently good ways to do FOSS, proprietary closed-source and free closed-source software development. But if the OSS is worth charging for (commercial) use, devs are left with asking for donations, SaaS or "pay me to work on this issue/feature".

There arguably should be better mechanisms to reward OSS development, even if the largest part of an OSSndev's motivation is intrinsic.

TomasBM··on Ask HN: ADHD – How do you manage the constant stream of thoughts and ideas?
In addition to other great suggestions (a good night's sleep being #1), two things come to mind:

- Write down thoughts that pop-up and seem potentially useful, and then forget about them. It's easier said than done, but you have to balance "this may seem useful later" with "I got better things to do now". For me, knowing I have the idea recorded somewhere puts my mind at ease.

- Feel free to just get rid of accumulated browser tabs, random to-do's or even mental notes. If you always have multiple things open, do a hard reset every now and then. It's difficult to let go at first, but you may realize that if an idea was really worth entertaining, it will come back to you. No great contribution is the result of one single thought, IMO.

Note: I lean closer to scattered attention than AD(H)D, because it doesn't affect my normal functioning. Sure, it gets progressively more annoying when deadlines are looming, but it also provides a great source of creativity.

TomasBM··on Some ecologists fear their field is losing touch with nature
I don't do research that requires fieldwork, but even in office and industrial settings, I notice that there's less need and interest in visits.

Of course, in-person exchanges still happen, but there's something of a default to do most things remotely because it's more efficient (and honestly, easier for all parties involved). The result is that you don't get to see cool or unusual machines/setups that often, and some flair of doing research is lost.

I can imagine that that's especially painful for new ecologists, because fieldwork is also a way to experience things that you otherwise wouldn't. Hopefully, we can bring some of it back with edge devices and models.

TomasBM··on The Rise of SQL:the second programming language everyone needs to know
Somewhat tangential to the article, but why is SQL considered a programming language?

I understand that's the convention according to the IEEE and Wikipedia [1], but the name itself - Structured Query Language - reveals that its purpose is limited by design. It's a computer language [2] for sure, but why programming?

[1] https://en.wikipedia.org/wiki/List_of_programming_languages

[2] https://en.wikipedia.org/wiki/Computer_language

← PreviousPage 2 of 4Next →