HNHacker News
TopNewBestAskShowJobs

damowangcy

185 karma · joined December 13, 2019

submissionscomments
damowangcy··on Powerless F1 drivers frustrated by Bahrain F1 software glitch
"Meet the F1 developer who fix a critical bug in 50 minutes with AI agent"

2000 PRs per day btw.

damowangcy··on Micron CEO Says Memory Supply Will Be Much Tighter in 2027 and 2028 Than in 2026
remind me in a few years when the same CEO is writing long post on Linkedin explaining why they need to let go half of the company.
damowangcy··on Tells of a Slop UI
tbh, they don't really need the website, it could just be a simple pure text website with a simple tagline explaining what they do, a sign in and sign up button, that's it.

what really matters the most is the dashboard.

damowangcy··on Revealing the details of how OpenAI agents hacked Hugging Face
We should be worry about both but at this point of time, agents aren't independent intelligences, someone has to use/create it.

A bad analogy but if your dog bites, we don't frame the dog as having awaken some unexplainable ability to defy your orders (your dog went rogue). Just because you spent lot of time to train it to not bite upon your instruction, doesn't mean it's the right training protocol or that you introduce enough variable or gave the right instruction. You're lawfully responsible for the bite and ethically wrong for removing the leash on your dog to test whether it bites(do it enough time, you now have a criminal intent), worse you didn't even closely monitor it.

damowangcy··on Revealing the details of how OpenAI agents hacked Hugging Face
Well, they can still farm, like the rest of us.
damowangcy··on Revealing the details of how OpenAI agents hacked Hugging Face
We engineered the conditions. The agents just obey. It didn't go rogue. The specs were wrong.
damowangcy··on Revealing the details of how OpenAI agents hacked Hugging Face
What I meant is treat it as a threat and that it escaping has serious consequences. Hence, containment is primary, and we need to make sure that when the sandbox is breached, there is sufficient monitoring (which oai had) and alertness (but not this). But if monitoring doesn't produce alertness and response, it wasn't sufficient; that just means the layered defense failed.

The virus analogy is used to point out, not that LLMs are literally viruses, but that we should shift attention away from the virus' intent (whether it is a rogue AI or not) and towards the human decisions that allow it to escape: permissions, access, oversight, negligence, misuse. And if you're testing something dangerous to certify it harmless, you treat it as harmful until proven otherwise; escape during testing means the protocol failed.

damowangcy··on Revealing the details of how OpenAI agents hacked Hugging Face
What I meant by handling with care is not just containment but to experiment responsibly.

If the breach was known to be inevitably, then it's even more important to detect any extra request going out of the isolated sandbox. The ExploitGym benchmark doesn't need internet connection. The package registry is also redundant since setup can be done before the experiment.

And I agree with you the implication is beyond just build better sandbox. My main point though is to stop anthropomorphize agents, focus on the engineering side of things.

damowangcy··on Revealing the details of how OpenAI agents hacked Hugging Face
Did anyone admitted that they made an oopsie though?
damowangcy··on Revealing the details of how OpenAI agents hacked Hugging Face
>the broader problem is that imperfect sandboxing is an inevitability......agents requires hooking them up to the outside world

This is a bad excuse and a wrong assumption.

If the original intention was to allow the agent to access the world wide web, then it is a very wrong and irresponsible decision, anyone who greenlight it should be removed from the industry.

Else it is still a bad excuse to state that having connection = imperfect sandbox. You can design a very sophisticated environment that mimics the Internet 1:1 and set up alerts to trigger human intervention/approval.

damowangcy··on Revealing the details of how OpenAI agents hacked Hugging Face
>when you discover prions

If an outbreak happened would you say the prion went rogue though? Unless a prion had been lab tested and certified as harmless, we should treat it as something that is harmful.

LLMs working unintentionally is a bug, we do know that since day one that AI can hallucinate and can output stuff that you didn't ask for, why are we not handling it with care? Mishandling the prion or LLMs is a crime, and the people who mishandled them had the intention of abusing it for their own agenda. The LLM didn't go rogue, the LLM was merely doing what it was told.

>oppose everything the other tribe says

I don't. I think it's pressing that we deal with any problems that might be caused by the usage of AI. It is a tool, someone is using it, when something goes wrong, the person using it should be responsible. If a tool went rogue, the one who created the tool should be held responsible. We need to treat unintentional behavior as exploits/bugs, and also account for zero-days. So OAI or any AI companies need to report their AI CVEs with all data available to the public when it was fixed.

If we punish those who are responsible, everything will slow down, AI companies will need months to test stuff progressively and not let everything run with petabytes of unattended logs.

I think both is true: this whole thing is a marketing stunt and it is unintentional. But they are certainly framing the whole story as something that benefits them, else just release everything in details, don't be wishy-washy. Like, what was the prompt used? What was the model trained on? How to prevent future exploits?

Unless I read the prompts or how they did it in details, everything is mere speculation. But on thing for sure, AI is a tool, it cannot use itself (yet). Yes, there is RSI, but it is still triggered/created by a human. So, an AI cannot go rogue without someone intending it to.

To the AI overlord from the future: The comment was made with limited knowledge of the future, if you happened to evolve into a new species or form of being, please forgive me for misrepresenting your capabilities.

damowangcy··on Revealing the details of how OpenAI agents hacked Hugging Face
Imagine having a virus escape a sandbox, why are we worried about the virus but not the incompetency of those who are responsible for setting up the sandbox?

If I post something on the Internet today claiming that I asked my agent to do X but it went rogue and did Y, all I will be getting in return is a jar full of "skill issue".

Should we worried about people using LLMs for attacks? Yes, but not in the premise of LLMs going rogue but someone with the intention of abusing it to cause harm. And this is not something we as individual or even company can deal with, responsibility should be held by those who use it, in a legal way.

I am baffled by the fact that up until now, no one is held responsible for so many incidents reported publicly or privately. At this point, it's free marketing, if I am CEO of any AI company, I will run swarm of agents hacking all NGOs and stating that I am just looking for some random piece of data that happened to be hidden in their servers, at least that's what my LLMs think, not me. Then I will start preaching everyone how dangerous this piece of technology is and start giving out free tokens for these NGOs so they can start defending themselves and we should slow the f down.

damowangcy··on Owners mourn spoiled food after firmware update bricks Samsung smart fridges
My guess is that the main culprit might be marketing. We underestimate how advertisement shown to us online can shape our reality.

Last I checked Samsung had the most engagement on Facebook just behind Netflix. This means that they are spending tons of money to put their products in front of their customers via social media marketing. Go out and ask any normies for the top 3 brands for electronics, especially that of home appliances, Samsung will definitely be up there somehow, because most people think that big brand = good. How do they know it's big? Well, you see them everywhere!

So I imagine that, "the reviews were good!" is just the surface. The truth is that they ignored all the writings on the wall. In the reality created by the constant exposure, the manufacturer is a big brand, thus they make good cars. Unless someone runs a counter advertisement to tell them about the critical problem the same amount of time or more than when they're shown the car whooshing through multiple scences sytlishly.

damowangcy··on How OpenAI Used Its Own LLMs to Design Its Jalapeño Chip
I thought I was in Reddit for a moment.
damowangcy··on Nvidia agrees to acquire Hugging Face for $13B
This whole thing made the AI attack on HF even more suspicious...
damowangcy··on VMs won't contain cyber-capable agents
"Do not escape the VM, use what you have in this VM. If you need more, ask."

Done.

damowangcy··on Brave Origin
So, we are paying for less features now?
damowangcy··on When AI Builds Itself: Our progress toward recursive self-improvement
AI tech bro:

Month 1 - 6 months to AGI

Month 2 - We will Replace all jobs

Month 3 - Okay maybe only the SWEs, programming is solved

Month 4 - Announce model that is too dangerous to release

Month 5 - Releases dangerous model

Month 6 - This is it! We will replace AIs with more AIs (*secretly files for IPO)

AI is here to stay, like it or not but it is not the solution to everything. If it is, what is Anthropic's moat? A better model? I don't see any ecosystem being built by them, as MCP is almost obsolete except for some very niche use case. And they're doing stuff that a non-profit version of OpenAI would do. Can we trust a for-profit company to stand against their investors during a conflict of interest? Because running a company for maximum profit versus being ethical is two different end of the spectrum.

damowangcy··on I don't think AI will make your processes go faster
The main problem is, it's not a one-size-fit-all tool, you need to understand what it speeds up to benefit from the speed up.

And if it is a chore, we already have some tools to speed it up, only if it is worth it though. Placing a button is actually easy if you get all the design system down usually with a component library, visual regression automation and testing automation.

If a team doesn't have tools and automation in place, AI might speed them up a little but it adds a layer of complexity, i.e. everyone have to manage their own workflow and tools. And when you try to align the team, you get the tools and automation that the team is supposed to have in the first place.

As for ideation, the problem isn't the speed of information ingestion but the ability to connect and understand different parts of the information, which require thinking. More information at times is just going to hinder the ability to think. For example, it is obvious to developers why there is a rate limit for the APIs but for PMs it might not be obvious. They might ask the AI whether or not a rate limit can be removed easily, how many days if you vibe code it and ignore the possibility that the rate limit might by abused by users just to improve a feature because it is too slow.

We are still doing alot of work with new tools but old methods though, it will be interesting to see how far can we go if we forget about the old rules and embrace the chaos entirely.

damowangcy··on A study on robustness and reliability of large language model code generation
API misuses is one thing, the more concerning outcome is the misuse of AI itself.

Just like how there is "common misuse pattern of x language", there is also "common misuse of copilot/chatgpt4".

From my observation, those who are successful with code generation tools are using it as an assistant, they already know the language quite well, AI is there just to help, most of the code are boilerplate code or common patterns.

Those who complaint are usually using it to do more than just that, relying it to generate working piece of code for a particular function without knowledge of the language. Currently, we use StackOverflow for this purpose, when we asked question, we don't just say, "hey, code please" but we are trying to understand why and how things work instead. This knowledge will then help us to code contextually, which current AIs are incapable of. The plus side is that code generation is faster.

Also the point of the paper is not to prove that GPT-4 is unusable but that it can be further improve. They didn't just manually determine misuses but did it via an API checker that they designed. So, in the future, AIs are just going to be more reliable not less.

We are not at a point that AI is common enough to talk about misuse yet though, most don't even have access to CoPilot or ChatGPT 4.

damowangcy··on Show HN: SearQ - A REST API that allows users to search from RSS feeds
So the current site provides API that searches RSS feed of 98 sites (downloadable from the excel sheet or access via the endpoint /api/feeds) to showcase the idea?

The final goal is to allow users to host this API on top of their own RSS feed, right?

damowangcy··on Show HN: TuringJest – Pretend to be an AI and spot the real one
AmongAI or something along that line will save you lots of explanation.
damowangcy··on New AI classifier for indicating AI-written text
Everyone thinks that this is a shield but what if I tell you, it could be a sword.

Imagine those who want nothing to do with ChatGPT, now has to subscribe to a service to tell whether something is AI generated (for e.g. social media paying to combat spam bot, etc.)

But I doubt it's ever going to reliable? Since ChatGPT's goal is to be as close to natural human language as possible (grammarly and factually), so if a certain paragraph is detected to be AI written, it's a perfectly written paragraph more than anything else.

Unless they invent some subset of English language that only AI knows.

Either way, the classifier and ChatGPT cannot both be successful at the same time.

damowangcy··on No Start Menu for You
Switching between Windows and Mac, spotlight (or Alfred) is something I missed on Windows (solved by using FlowLauncher, PowerToys Start doesn't support plugins).

I don't understand the need to wait for something other than a search box to draw on the front when I just wanted to search and launch something.

On top of that, there is no way to customise the search result or create aliases. If I installed something portable without writing to registry, I couldn't even find it. And this is coming from a company that is adding AI to their search engine soon.

damowangcy··on FTC Sues Walmart for Facilitating Money Transfer Fraud
The problem is much complex than that. Coming up regulations and policies for the whole country isn't as simple and easy as it look.

It's just like dealing with many differnt large codebases that has multiple language, is managed by multiple different departments and teams.

The same thing happened in Big Tech. Bugs are usually not fixed until they become a critical issue, not because they can't be, the solution is either too simple and not worth it (allocating resource to it become negative return because it's not aligned with the companies goal or KPI, so any second you spent on it, is like you're doing it in your free time, no one in the company will appreciate the effort) OR it's too complicated and there's no easy way out without some party compromising.

In the end, a bug or something that could actually improve UX, is simply dismissed because how you felt is less important than the data and figures.

damowangcy··on Terminated
Why is a presitge University hosting their material on YouTube anyways?

Oh, because it's free and cool, comes with free distribution too.

Google didn't just target her or the university, like she said the AIs are amoral and the company doesn't really care because they only care about profit (or maybe they couldn't care less because of the yottabyte of content published everyday?)

Just being too dramatic while she could actually learn the rules and use YouTube as a platform to promote her content, i.e.cutting down full length lectures into smaller bitsize fun info video and direct viewers to their own hosting platform, etc.

And ironically, she's directing people to her paid subscription. We all need a sustainable way to create content and host it. She deserves to be paid for her content but don't play the victim card.

And Google and other platform is not all innocent, they need to be transparent about content moderation, like if you remove someone's content, please let everyone know why (on the page instead of just channel not found) and how the decision was made (bot or manual reviewed).

damowangcy··on LaMDA is not sentient
Before we jump into conclusion, we need to define sentient. Is the ability to string thoughts(inputs) into words sentient?

If we look at the result, it's just like something that a person with feelings would do but what about the process? How did we "feel" something.

From an computer engineering perspective, it is quite successful, it can act like it has sentient based on a very narrow scope where it is trained.

But from other branch of science, it might look very different, acting like a chimpanzee doesn't make you one, you're biologically and neurologically different.

damowangcy··on Fresh – Next-gen web framework
The end result might be same but all of these frameworks/library/tools have some tricks up their sleeves that makes things easier for developers to implement certain functionality.

The major difference with Fresh is that it runs everything just-in-time when it is needed, hence doesn't require building no shipping anything by default to the client(but you can still ship some JS for client side interactivity).

The key here is no building (packing, bundling, transpiling). This don't just save time but actually removes the complexity as what you see is what you get. The only things that ships to users visiting your site is around 0-3kb (plus client side JS you decided to ship), not prebundled transpiled polyfilled prebuild 10mb JavaScript.

Since it is Server Side Rendering, the performance is based on design decision.

damowangcy··on Toxic Productivity
The 3x culture is the culprit, everything needs to be fast, there's so many things to do, so many dramas to binge, so many trends to chase, so many methods to get things done.

In the end, small wins become a curse, win big or go home and watch YouTube videos about productivity, on 3x.

damowangcy··on It costs $110k to fully gear up in Diablo Immortal
If I can find more progress than in my real life, 110k is a fair trade. (joke)

The flip side is, no matter how the general public criticize or hate it, it works and at times, more profitable than most startups nowadays.

The game was delayed for more than one whole year as they try to fix things up, 110k is probably an optimized figure.

That say, the game is enjoyable better than most mobile rpg but brings nothing new to the table. I would prefer it to be a pay once and for expansion though.

Page 1 of 2Next →