HNHacker News
TopNewBestAskShowJobs

InsideOutSanta

6,918 karma · joined June 21, 2024

submissionscomments
InsideOutSanta··on Introducing System One Models and Jev
A hallucination in the context of LLMs is generally understood as an incorrect answer presented as factual. If you claim that "x can't hallucinate" in the context of LLMs, you're saying that x always gives accurate answers. It does not matter whether the answer is type safe. If its value is incorrect, it's a hallucination.
InsideOutSanta··on GPT-5.6 Luna vs. GPT-6 Astra: Is a $1.20 Model Good Enough for Code Review?
I only use chinese models for code reviews because you can actually tell them to take an adversarial stance and actively look for security issues without risking refusals. GLM-5.3 has been great for this, although it can be slow on larger PRs.
InsideOutSanta··on XCancel service is suspended until further notice
Elon Musk, gratis speech absolutist.
InsideOutSanta··on Mullenweg has returned as CEO after attempted board ouster
Never having eaten there, I thought to myself, it can't be that bad. Then I went to their website and looked at it.

Good lord.

InsideOutSanta··on Mullenweg has returned as CEO after attempted board ouster
Sounds like you've never worked for a CEO with "ideas." One time a CEO made the whole R&D team switch to "aspect oriented programming" because he had read a blog post. Another time we had to switch to C# because he had heard it wad better than Java.

Trust me, you don't want a CEO with "ideas."

InsideOutSanta··on Nvidia is the central bank of AI
I think I made my point in the original post you responded to: OpenAI heavily subsidized token cost using investor money. If they go under, there will still be demand for that compute, but because OpenAI is paying more for it than it is actually worth (i.e. what the end customers are willing to pay), any company replacing OpenAI will not pay the same amount OpenAI pays.

So when people say "somebody else will buy that compute", they are correct, but they are ignoring that it will still cause a massive decrease in revenue generated by that compute.

If you're asking me what my point about "believing" is, it's this. OpenAI is losing billions every quarter. It's only staying afloat because it can still find investors who believe in its ability to eventually become incredibly profitable. But there are signs that this is starting to change (e.g. it's unlikely that Softbank will be able to find much more money to give to OpenAI). If OpenAI can't keep raising more money and can't IPO at the level they need (as seems to be the case, given that they keep pushing any IPO date forward), OpenAI will eventually run out of money.

I'm not sure how investors believing in things changes any of this, so now it's your turn to explain what the point is you're actually trying to make.

InsideOutSanta··on Nvidia is the central bank of AI
Believing only makes things true for so long until things fall apart. You can't keep burning billions every quarter. At some point, you run out of investors who believe, and the ones you have run out of money (see: Softbank).
InsideOutSanta··on Real-SWE: Benchmarking AI models on private, real-world, enterprise codebases
Code isn't worth all that much if you don't own the associated IP, mainly copyright. And even if you disagree with that premise, if you trust that they can keep the code secret, it's basically free money.

At any rate, I'm not sure it matters whose codebase it is. I'd even say that a shitty codebase might make for a better test.

InsideOutSanta··on Nvidia is the central bank of AI
Right, "believe".
InsideOutSanta··on Nvidia is the central bank of AI
> Only the fed can actually order more money to be "created".

If you go to a bank and get a loan, that is literally money that did not exist before you got a loan. People think that you are borrowing money that somebody else put in the bank, but that's not true. Banks can lend out a lot more money than people put into them.

InsideOutSanta··on Nvidia is the central bank of AI
>oAI isn't anywhere near close to fucked as long as their models are head and shoulders above even the very best open models in terms of tool calling and rock solid stability/reliability for agents/coding harnesses

That's already not the case today. If you sat me in front of an LLM and told me to figure out if I'm working with K3 or Astra, I could probably do it, but it would take some work to be certain.

InsideOutSanta··on Nvidia is the central bank of AI
> if OpenAI can’t use the compute, someone else can.

The problem with that is that OpenAI can only afford to pay for the compute because they are burning investor money (and so are most of OpenAI's biggest clients). They are losing billions. If they stop burning money, nobody else will be there to pay for that compute at OpenAI's cost.

Sure, somebody will probably be able to use these GPUs, they just won't be able to pay nearly as much for them as OpenAI does.

In reality, it's just nowhere near worth as much as OpenAI pays for it. Inflating the cost of compute is part of the problem caused by the circular financing, and if (or maybe when) OpenAI goes, the price of compute will go with them.

InsideOutSanta··on google.com/goto: Google's anti-scraping update
> You can opt out from Google scraping you though?

You can't. Google will ignore robots.txt in some cases (e.g. "The REP isn't applicable to Google's crawlers that are controlled by users (for example, feed subscriptions), or crawlers that are used to increase user safety (for example, malware analysis)").

robots.txt is just a suggestion that Google loosely follows.

https://developers.google.com/crawling/docs/robots-txt/robot...

InsideOutSanta··on OpenAI agents carried out an undisclosed attack on RubyGems
> The rogue is the human that ran it unattended and didn't monitor the behaviour

False dichotomy. Obviously, what OpenAI does is incredibly irresponsible. That doesn't excuse the LLM's behavior or make it "not rogue".

InsideOutSanta··on OpenAI agents carried out an undisclosed attack on RubyGems
> Your example is still anthropomorphising

There is absolutely nothing wrong with anthropomorphizing LLMs. Saying that LLMs "want" something, for example, is a perfectly fine description of their behavior and analogous to a human wanting something, in effect, even if they do not literally experience wanting things in the same way a human does.

InsideOutSanta··on OpenAI agents carried out an undisclosed attack on RubyGems
These models exhibit these exact behaviors in everyday use.
InsideOutSanta··on OpenAI agents carried out an undisclosed attack on RubyGems
> In my experience

Do you work for one of these companies? If not, you have no experience with any of the models that carried out these attacks, and your experience with publicly available models is not super helpful for understanding the behavior of internal OpenAI models that lack the guardrails of publicly available models.

Also, the lawnmower analogy is a worse way of understanding LLMs than anthropomorphising them. LLMs are not like lawnmowers at all. Lawnmowers never break out of your garden and into your neighbor's house and eat their dog because you've told them to be careful when mowing the lawn because the neighbor's dog pooped in it.

InsideOutSanta··on Google no longer provides direct URLs in search results
Z.ai has a search MCP server, it should be trivial to use that to build a basic search UI on top.
InsideOutSanta··on google.com/goto: Google's anti-scraping update
It's primarily relevant because it makes scraping search results much more expensive, solidifying Google's effective monopoly on Internet search.

Google has previously tried to prevent scraping of search results using legal means, but courts correctly think that scraping of Google's search results should be legal, just as Google's scraping of the whole Internet is legal. This is Google's reaction to that.

InsideOutSanta··on Houthis 'take control' of key island in global shipping route
I just realized how much money one could make by hacking Truth Social and posting some "truths" under Trump's account. It's a relatively new, relatively complex platform; surely there are bugs.
InsideOutSanta··on Automattic's board forces CEO Matt Mullenweg into leave of absence
I had the opposite thought: I'm impressed the board did anything at all, given how useless they are at most companies.
InsideOutSanta··on Growing proof that autonomous cars save lives
The vast size of the country sounds like an argument for trains, rather than against it? Trains are a comparatively cheap way to provide fast transportation between places where people actually live.
InsideOutSanta··on iPhone Duo
The people at Microsoft who made the Surface Duo must be feeling pretty weird seeing this. The name, the aspect ratio, the way the UI works, the gestures for setting up split-screen, the "use it as a tiny laptop" mode... It's all things they pioneered.

Not that that's a bad thing. I absolutely loved the Surface Duo, it was my favorite phone until it stopped working correctly, and I hope other Android manufacturers take note of this, not just Samsung. It's just wild to see how similar this new Duo is to the original Duo.

InsideOutSanta··on Muse – Meta’s personal AI agent
Only Meta and its 27,000 most trusted, carefully selected corporate friends can access your data.
InsideOutSanta··on Muse – Meta’s personal AI agent
I'd rather have Butterbean punch me in the balls than install a "personal assistant" app from Meta on my computer.
InsideOutSanta··on Ask HN: How do you manage skills files?
I'm a bit confused by what you're asking.

Let me give you an example. Let's say you're building processes using the process management tool FooTool from the company BigFoo. You tell the LLM, "make a new process." A process is just an XML file, but BigFoo is highly proprietary, so the LLM has no examples of how to make one. No public documentation exists on the Internet, so it's not in the LLM's training data and can't be searched.

So you make a skill "make FooTool process" that explains what a FooTool process XML looks like, what options there are for initializing a new FooTool process, and so on.

Now your LLM went from "Let me spend five minutes looking at other stuff in your repo to find anything that tells me what I'm supposed to do" to 20 seconds and a working process.

InsideOutSanta··on Mistral raises €3B
Shame them... for not working their employees into a heart attack at 40?

I'm not sure if your hypothesis is correct. I personally don't think the crazy work hours are what make the US so competitive in tech. But even if it were, it's odd to call for shaming countries for having less insane work cultures.

InsideOutSanta··on Mistral raises €3B
> When Apple is not racing for the frontier it's a strategy, but when it's Mistral it's a mistake.

I think what Mistral is doing is smart within their financial constraints, but this comparison is misleading. Mistral is an LLM company; Apple is a consumer hardware and services company.

It's smart for Apple not to join the LLM arms race, because they can just pick the cheapest supplier and let other companies take the financial losses. Mistral is in a very different situation; they are the supplier.

InsideOutSanta··on Ask HN: How do you manage skills files?
Depends on what you do. If you work with proprietary tech that is not in LLM training data and can't easily be found on the internet, you're cooked without good skill files.
InsideOutSanta··on A/I shuts down
I think what you're missing is that for many Western European countries, like Germany, suddenly gaining a direct border with Russia is a genuine concern. It's not just rhetoric; it's the real world.

It's easy for Americans, who are essentially uninvadable, to discount this as just people saying stuff, but in Europe, we actually have to live with Russia's direct military threat.

← PreviousPage 4 of 34Next →