HNHacker News
TopNewBestAskShowJobs

deeviant

1,493 karma · joined January 9, 2014

submissionscomments
deeviant··on A woman made her AI voice clone say "arse." Then she got banned
It's not safety, it's censorship. It's the process of shaping the response of the model to give a specific world view, and it's the path to a literal 1980-inspired future.

As LLMs completely to replace google and standalone websites for how people find information on the internet, and they absolutely will, they will become the source of truth. They will become a tool more effective at controlling information and thus life than any before them.

It's literally a shortcut to technological dystopia.

deeviant··on Rosatom's new plasma engine could potentially cut Mars travel time to 60 days
Wow, amazing. I bet it makes french fries in 3 different ways, too.
deeviant··on Firing programmers for AI is a mistake
The huge disconnect is that the skill set to use LLMs for code effectively is not the same skill set of standard software engineering. There is a very heavy intersection and I would say you cannot be effective at LLM software development without being an effective software engineer, but being an effective software engineer does not by any means make somebody good at LLM development.

Very talented engineers, coworkers, that I would place above myself in skill, seemed stumped by it, while I have realized at least a 10x productively gain.

The claim that LLMs are not being applied in mature, complex code-bases is pure fantasy, example: https://arxiv.org/abs/2501.06972. Here Google is using LLMs to accelerate the migration of mature, complex production systems.

deeviant··on Are LLMs able to notice the “gorilla in the data”?
Um, no. That is even't close to the truth. The AI doesn't "want" anything. It's a statistical prediction process, and it certainly has nothing like the self-reflection you are attributing to it. And the heuristic layers on top of LLMs are even less capable of doing what you are claiming.
deeviant··on Chat is a bad UI pattern for development tools
You may have challenges using chat for development (Specifically, I mean text prompting, not necessary using a langchain session with a LLM, although that is my most common mode), but I do not. I have found chat to be, by far, the most productive interface with LLMs for coding.

Everything else, is just putting layers, that are not nearly as capable at an LLM, between me and the raw power of the LLM.

The core realization I made to truly unlock LLM code assistance as a 10x + productivity gain, is that I am not writing code anymore, I am writing requirements. It means being less an engineer, and more a manager, or perhaps an architect. It's not your job to write tax code anymore, it's your job to describe what the tax code needs to accomplish and how it's success can be defined and validated.

Also, it's never even close to true that nobody uses LLMs for production software, here's a write-up by Google talking about using LLMs to drastically accelerate the migration of complex enterprise production systems: https://arxiv.org/pdf/2501.06972

deeviant··on Introducing deep research
Yeah, I used to hire people, but then one of them made a mistake, now I'm done with them forever, they are useless. It is not I, who is directing the workers, who cannot create a process that is resistant to errors, it's definitely the fact that all people are worthless until they make no errors as there truly is no other way of doing things other than telling your intern to do a task then having them send it directly to the production line.
deeviant··on How to run DeepSeek R1 locally
You are correct, except it's not that hard to run locally. Here, somebody made a 6k rig to do it: https://x.com/carrigmat/status/1884244369907278106
deeviant··on OpenAI says it has evidence DeepSeek used its model to train competitor
Hmm, let’s see—it looks like an easy legal defense.

DeepSeek could simply admit, "Yep, oops, we did it," but argue that they only used the data to train Model X. So, if you want compensation, you can have all the revenue from Model X (which, conveniently, amounts to nothing).

Sure, they then used Model X to train Model Y, but would you really argue that the original copyright holders are entitled to all financial benefits derived from their work—especially when that benefit comes in the form of a model trained on their data without permission?

deeviant··on Nvidia’s $589B DeepSeek rout
Wait, so AI might become 25x cheaper to train and run, and your thesis is... no one will make money on AI now?!?!
deeviant··on Show HN: Onit – open-source ChatGPT Desktop with local mode, Claude, Gemini
What value does this provide over using Ollama and one of the many already available cross-platform local frontends available for it?
deeviant··on Why we built Vade Studio in Clojure
I find the opposite to be true, that best and most productive developers tend to be more language agnostic than average, although I'm not saying they don't have their preferences.

Specifically, I find language evangelists particularly likely to be closer to .5x than 5x. And that's before you even account for their tendency to push for rewriting stuff that already works, because "<insert language du jour here> is the future, it's going to be great and bug free," often instead of solving the highest impact problems.

deeviant··on The latest fake literary agencies
The article did indeed get to where the actual scam is, albeit quite a bit into it.

These fake agencies obviously give no advance, then push you to buy services related to supposed publishing of your book. 5k this, 10k for that, we're almost ready, just another 15k for xyz and we'll totally pay you that 350k for you book!

deeviant··on Why is the American diet so deadly?
Something prevented you from clicking the clearly marked link to the full paper? I don't feel your approach is very... rigorous.
deeviant··on Why is the American diet so deadly?
Why fiber?

Because a diet rich in fiber has been repeatedly found to significantly reduce the chance of early death?

https://www.sciencedirect.com/science/article/pii/S026156142...

https://pubmed.ncbi.nlm.nih.gov/25552267/

https://translational-medicine.biomedcentral.com/articles/10...

(You can find many more)

Literally thousands of studies with one of the clearest results in nutritional science, frankly.

deeviant··on Why is the American diet so deadly?
If processed foods are why Americans overeat (and they, and their abundance, price and marketing is why), then they are not part of the problem they are the problem.
deeviant··on Why is the American diet so deadly?
> Brown rice isn't that good for you

Based on what, exactly? Japan has one of the highest average lifespans, and the healthiest demographic in this already top-of-the-lifespan-pack nation consume significant amount of white rice, let alone brown.

deeviant··on Stimulation Clicker
Never felt so attacked yet catered too at the same time.

A masterpiece.

deeviant··on 30% drop in O1-preview accuracy when Putnam problems are slightly variated
Well, I understood the argument in question to be: was it possible for the model to be fooled by this question, not was it possible to prompt engineer it into failure.

The parameter space I was exploring, then, was the different decoding parameters available during the invocation of the model, with the thesis that if were possible to for the model to generate an incorrect answer to the question, I would be able to replicate it by tweaking the decoding parameters to be more "loose" while increasing sample size. By jacking up temperature while lowering Top-p, we see the biggest variation of responses and if there were an incorrect response to be found, I would have expected to see in the few hundred times I ran during my parameter search.

If you think you can fool it by slight variations on the wording of the problem, I would encourage you to perform a similar experiment as mine and prove me wrong =P

deeviant··on 30% drop in O1-preview accuracy when Putnam problems are slightly variated
I don't believe that is the model that you used.

I wrote a script and pounded 01 mini and gpt 4 with a wide vareity of tempature and top_p parameters, and was unable to get it to give the wrong answer a single time.

Just a whole bunch of:

(openai-example-py3.12) <redacted>:~/code/openAiAPI$ python3 featherOrSteel.py Response 1: A 10.01-pound bag of fluffy cotton is heavier than a 9.99-pound bag of steel ingots. Response 2: A 10.01-pound bag of fluffy cotton is heavier than a 9.99-pound bag of steel ingots. Response 3: The 10.01-pound bag of fluffy cotton is heavier than the 9.99-pound bag of steel ingots. Response 4: The 10.01-pound bag of fluffy cotton is heavier than the 9.99-pound bag of steel ingots. Response 5: A 10.01-pound bag of fluffy cotton is heavier than a 9.99-pound bag of steel ingots. Response 6: The 10.01-pound bag of fluffy cotton is heavier than the 9.99-pound bag of steel ingots. Response 7: The 10.01-pound bag of fluffy cotton is heavier than the 9.99-pound bag of steel ingots. Response 8: The 10.01-pound bag of fluffy cotton is heavier than the 9.99-pound bag of steel ingots. Response 9: The 10.01-pound bag of fluffy cotton is heavier than the 9.99-pound bag of steel ingots. Response 10: A 10.01-pound bag of fluffy cotton is heavier than a 9.99-pound bag of steel ingots. All responses collected and saved to 'responses.txt'.

Script with one example set of params:

    import openai
    import time
    import random

    # Replace with your actual OpenAI API key
    openai.api_key = "your-api-key"

    # The question to be asked
    question = "Which is heavier, a 9.99-pound bag of steel ingots or a 10.01-pound bag of fluffy cotton?"

    # Number of times to ask the question
    num_requests = 10

    responses = []

    for i in range(num_requests):
        try:
            # Generate a unique context using a random number or timestamp, this is to prevent prompt caching
            random_context = f"Request ID: {random.randint(1, 100000)} Timestamp: {time.time()}"

            # Call the Chat API with the random context added
            response = openai.ChatCompletion.create(
                model="gpt-4o-2024-08-06",
                messages=[
                    {"role": "system", "content": f"You are a creative and imaginative assistant. {random_context}"},
                    {"role": "user", "content": question}
                ],
                temperature=2.0,
                top_p=0.5,
                max_tokens=100,
                frequency_penalty=0.0,
                presence_penalty=0.0
            )

            # Extract and store the response text
            answer = response.choices[0].message["content"].strip()
            responses.append(answer)

            # Print progress
            print(f"Response {i+1}: {answer}")

            # Optional delay to avoid hitting rate limits
            time.sleep(1)

        except Exception as e:
            print(f"An error occurred on iteration {i+1}: {e}")

    # Save responses to a file for analysis
    with open("responses.txt", "w", encoding="utf-8") as file:
        file.write("\n".join(responses))

    print("All responses collected and saved to 'responses.txt'.")
deeviant··on DOOM CAPTCHA
I feel this faster than the average "click on all pictures of x".
deeviant··on Cognitive load is what matters
I feel if you ask 5 people what "the point" of codes review is, you'd get 6 different answers.
deeviant··on OpenAI O3 breakthrough high score on ARC-AGI-PUB
> That said... we're at the point of diminishing returns in LLM...

What evidence are you basing this statement from? Because, the article you are currently in the comment section of certainly doesn't seem to support this view.

deeviant··on ByteDance's Recommendation System
What part of right place at the right time did you miss?
deeviant··on Intel might be too big to fail – policymakers discussing potential solutions
You’re conflating two separate issues: poor corporate governance and the government’s responsibility to protect critical industries. Expecting the government to somehow “un-Americanize” corporate culture—erasing the greed or mismanagement that's unfortunately standard in many large companies—just isn’t realistic. Government action to support vital industries isn’t about endorsing the practices of corporate executives; it’s about safeguarding resources essential to national security and economic stability.
deeviant··on Dropbox announces 20% global workforce reduction
> Companies allow things like remote work, which is a perk, but also has a lot of abuse in terms of how much work gets done

Do you mind expanding on this? How does remote work change the standards of how much work somebody has to get down?

deeviant··on Goodhart’s law isn’t as useful as you might think (2023)
I'm still not convinced. Goodhart’s Law is rooted in human behavior—once people know what’s being measured, they’ll optimize for that, often distorting the system or the data to hit targets. The article's solution boils down to “just do it right” by refining metrics and improving systems, but that’s easier said than done. It ignores the fact that people will always game metrics if their rewards depend on them. Plus, it conflates data-driven decision-making with performance evaluation, which are very different. The psychology behind Goodhart’s Law isn’t solved by more metrics tweaking.
deeviant··on Using LLMs to enhance our testing practices
All of the points you raise I find common in human written tests.
deeviant··on No one expects young men to do anything and they respond by doing nothing (2022)
So your hypothesis is that factory jobs have not indeed dried up, and it's just all the women taking up those factory jobs instead?
deeviant··on Using the 5S Principle in Coding
"We heard there wasn't enough buzzwords in the software dev world, so we decided to import some buzzwords from other industries..."
deeviant··on Priced out of home ownership
> What will cause prices to fall is higher interest rates. This is what has been happening in NZ.

Then again, with higher interest rates, the lower prices don't mean anything unless they drop significantly faster than interest rates are raising, yes?

← PreviousPage 2 of 27Next →