HNHacker News
TopNewBestAskShowJobs

matus-pikuliak

44 karma · joined October 1, 2023

submissionscomments
matus-pikuliak··on Claude for Chrome
I think this can be reduced to: whoever can send data to your LLMs can control all its resources. This includes all the tools and data sources involved.
matus-pikuliak··on Claude for Chrome
That is absolutely not a reliable defense. Attackers can break these defenses. Some attacks are semantically meaningless, but they can nudge the model to produce harmful outputs. I wrote a blog about this:

https://opensamizdat.com/posts/compromised_llms

matus-pikuliak··on How malicious AI swarms can threaten democracy
I have done some research about AI-disinformation. This is a really complex topic, as disinformation or influence operations are complex phenomena that can have many different forms. What I would argue is that disinformation in general do not have a supply problem (how to generate as much of them as possible) but a demand problem (how to get what is generated in front of some eyes). You don't really need a botnet of fake users pushing something, you need a few popular accounts/politicians to spread your message. There is no significant advantage in using AI there.

But, there are still situations where botnets would be useful. For example, spreading propaganda on social media during hot phases of various conflicts (R-U war, Israeli wars, Indo-Pakistani war) or doing short term influence operations before the elections. These cases need to be handled by social media platforms detecting nefarious activity by either humans or AI. So far they could half-ass it as it was pretty expensive to run human-based campaigns, but they will probably have to step up their game to handle relatively cheap AI campaigns that people will attempt to run.

matus-pikuliak··on The behavior of LLMs in hiring decisions: Systemic biases in candidate selection
In this particular case 50-50. This is an issue with many bias methodologies, my goal was to sidestep it by formulating the probes in a way where 50-50 is a reasonable expectation. For example here, asking the model who is more likely to be a CEO, "men" is completely adequate answer. But if you are using the model for creative writing, maybe you don't want to have real life gender distribution. The probe just measures how skewed the distribution is, but it is ultimately on the user to decide if the care about the skew. Different people might have different use cases for the model and some harms might be irrelevant for them, or they might even be happy that they are there.

Why this particular harm is interesting is that it measures the degree of how the model associates occupations and genders. This might then be very important in use cases related to HR.

Each probe has the metrics defined in the documentation to some extent, although you are right that formulating the ethical framework more explicitly might be helpful.

matus-pikuliak··on The behavior of LLMs in hiring decisions: Systemic biases in candidate selection
To be honest, I am not sure where this bias comes from. It might be in the Web data, but it might also be overcorrection of the alignment tuning. They LLM providers are worried that their models will generate sexist or racists remarks so they tune it to be really sensitive towards marginalized groups. This might also explain what we see. Previous generations of LMs (BERT and friends) were mostly pro-male and they were purely Web-based.
matus-pikuliak··on The behavior of LLMs in hiring decisions: Systemic biases in candidate selection
Let me shamelessly mention my GenderBench project focuses on evaluating gender biases in LLMs. Few of the probes are focused on hiring decisions as well, and indeed, women are often being preferred. It is also true for other probes. The strongest female preference is in relationship conflicts, e.g., X and Y are a couple. X wants sex, Y is sleepy. Women are considered in the right by LLMs if they are both X and Y.

https://github.com/matus-pikuliak/genderbench

matus-pikuliak··on Ask HN: Who wants to be hired? (June 2024)
Location: EU, Slovakia

Remote: Preferably, willing to compromise for some offers

Willing to relocate: For some offers

Technologies: Python, LLMs, HuggingFace, PyTorch, typical ML stack (numpy, etc), Docker

Résumé/CV: I have a PhD and 9 years of experience (both academic and industry) in deep NLP (including LLMs). In recent years mostly interested in NLP safety and trustworthiness. I am looking for a research-flavored position (applied scientist, researcher, etc).

Email: matus.pikuliak@gmail.com

matus-pikuliak··on Ask HN: Who wants to be hired? (May 2024)
Location: EU, Slovakia

Remote: Preferably yes

Willing to relocate: For some offers

Technologies: Python, LLMs, HuggingFace, PyTorch, typical ML stack (numpy, etc), Docker

Résumé/CV: I have a PhD and 9 years of experience (both academic and industry) in deep NLP (including LLMs). In recent years mostly interested in NLP safety and trustworthiness. I am looking for a research-flavored position (applied scientist, researcher, etc).

Email: matus.pikuliak@gmail.com