HNHacker News
TopNewBestAskShowJobs

silveraxe93

840 karma · joined January 15, 2020

submissionscomments
silveraxe93··on The new skill in AI is not prompting, it's context engineering
It's ironic how people write this without a shred of reasoning. This is just _wrong_. LLMs are not simply token prediction machines since GPT-3.

During pre-training, yeah they are. But there's a ton of RL being done on top after that.

If you want to argue that they can't reason, hey fair be my guest. But this argument keeps getting repeated as a central reason and it's just not true for years.

silveraxe93··on Quarkdown: A modern Markdown-based typesetting system
Do you know how it compares to marp? I've been using it last year and it's pretty nice. I hadn't heard about presenterm before.

- https://marp.app/

silveraxe93··on O3 beats a master-level GeoGuessr player, even with fake EXIF data
People found the original post so impressive they were saying that it had to be coming from cheating by looking at EXIF data. The point of this article was to show it doesn't. It got an unfair advantage in 1 (and say 0.5) out of 5. With the non-search rounds still doing great.

If you think this is unimpressive, that's subjective so you're entitled to believe that. I think that's awesome.

silveraxe93··on O3 beats a master-level GeoGuessr player, even with fake EXIF data
Yeah, the author does note that in the article. He also points it out in the conclusion:

> If it’s using other information to arrive at the guess, then it’s not metadata from the files, but instead web search. It seems likely that in the Austria round, the web search was meaningful, since it mentioned the website named the town itself. It appeared less meaningful in the Ireland round. It was still very capable in the rounds without search.

silveraxe93··on Overengineered Anchor Links
The way https://gwern.net/ does it is quite good.

The links open in a window, so you can still have centre aligned text with popups.

silveraxe93··on Gemini 2.5
From Gary Marcus' (notable AI skeptic) predictions of what AI won't do in 2027:

> With little or no human involvement, write Pulitzer-caliber books, fiction and non-fiction.

So, yeah. I know you made a joke, but you have the same issue as the Onion I guess.

silveraxe93··on Show HN: Letting LLMs Run a Debugger
I'd be extremely surprised if AI labs are not doing or planning on doing this already.

The same way that reasoning models are trained on chain of thoughts, why not do it with program state?

Just have a "separate" scratchpad where the AI keeps the expected state of the program. You can verify if that is correct or not. Just use RL to train the AI to always have that correct.

silveraxe93··on Benchmarking vision-language models on OCR in dynamic video environments
Honestly? I don't know how long it's been available. But I do know it's been some time already. Enough be aware of it when posting this on arxiv.

I'm not even disagreeing that it takes time to write papers, and it's "common" for this to happen. But it's just more evidence for what I said in my original comment:

> Academia is completely unprepared to deal with the speed AI develops

silveraxe93··on Benchmarking vision-language models on OCR in dynamic video environments
It was officially launched 10 days ago, but has been openly available for way longer.

Also, this is arxiv. The website that's explicitly about posting research pre peer-review.

silveraxe93··on Benchmarking vision-language models on OCR in dynamic video environments
For Google, definitely flash-2.0; It's a way better model. GPT-4o is kinda dated now. o1 is the one I'd pick for OpenAI. It's basically their "main" model now.

I'm not that familiar with Claude for vision. I don't think Anthropic focusses on that. But the 3.5 family of models is way better. If 3.5 Sonnet supports vision that's what I'd use

silveraxe93··on Benchmarking vision-language models on OCR in dynamic video environments
Sure, but they posted this 4 days ago. The minimum I'd expect for quality research is for them to skim the abstract before posting and change that line to:

"Models from leading AI labs" or similar. Leaving it like now signals either sloppiness or dishonesty

silveraxe93··on Benchmarking vision-language models on OCR in dynamic video environments
Posted 4 days ago:

> Three state of the art VLMs - Claude-3, Gemini-1.5, and GPT-4o

Literally none of those are state of the art. Academia is completely unprepared to deal with the speed Ai develops. This is extremely common in research papers.

That's literally in the abstract. If I can see a completely wrong sentence 5 seconds into reading the paper, why should I read the rest?

silveraxe93··on Firing programmers for AI is a mistake
oftwaresay engineeryay
silveraxe93··on 26% of students ages 13-17 are using ChatGPT to help with homework, study finds
If they are not listening to you, then I guess you did a great job then ;)
silveraxe93··on Ask HN: Am I the only one here who can't stand HN's AI obsession?
btw I'm not trying to defend LLMs here, I'd make the same comment if you flipped your question.
silveraxe93··on Ask HN: Am I the only one here who can't stand HN's AI obsession?
Look, I won't try to convince you're wrong. But if you come to a website with thousands of active users and ask: Do you agree with me? Don't reply if not.

What do you expect? Might as well ask an AI to generate that text, same level of information you'll be getting.

silveraxe93··on Be Aware of the Makefile Effect
This is the 40% that OP mentioned. But there's a proportion on people/engineers that are just clueless and are incapable of understanding code. I don't know the proportion so can't comment on the 50% number, but hey definitely exist.

If you never worked with them, you should count yourself lucky.

silveraxe93··on AI and Startup Moats
Larger models are more expensive to run (ceteris paribus). But we're seeing we can squeeze more performance from smaller models.

You need to compare like-for-like. You can't say that the cost of building a 5-story apartment is increasing by pointing at the burj khalifa.

silveraxe93··on AI and Startup Moats
ChatGPT is one of the fastest growing apps ever. Saying that's there's no products is willful blindness by this point.

This is hackernews. I'd expect users to have a basic understanding of VC investment. The expected value of next-gen models times the probability to create them is higher than the billions than they are throwing at it.

silveraxe93··on AI and Startup Moats
See situational-awareness[1], see the "algorithmic efficiencies" section. He shows many examples of how models are getting cheaper. With many citations.

Costs are not just down on a specific service. Even though I don't see the problem in that, as long as you get the promised level of performance, without being subsidised. See the deepseek model I linked above. It's an open model and you can run it yourself.

> At best, older models are getting cheaper to run.

What's your definition of old here? If you compare the literal bleeding edge model (o3) to 2 years ago best model (GPT-4)? Not only is this a ridiculously misleading comparison, it's not even valid!

o3 is a reasoning model. It can spend money at test time to improve results. Previous models don't even have this capability. You can't look at one example of where they just threw a lot of money and say this is the cost. The cost is unbounded! If they want, they can just not let the model think for ages and have basically "0-thinking" outputs. This is what you use to compare models.

If you compare _todays_ cost for training and inference of a model as good as GPT-4 when it was released, this cost has massively gone down on both counts.

[1] - https://situational-awareness.ai/from-gpt-4-to-agi/#The_tren...

silveraxe93··on AI and Startup Moats
But the cost is _definitely_ falling. For a recent example, see DeepSeek V3[1]. It's a model that's competitive with GPT-4, Claude Sonnet. But cost ~$6 Million to train.

This is ridiculously cheaper than what we had before. Inference is basically getting an 10x cheaper per year!

We're spending more because bigger models are worth the investment. But the "price per unit of [intelligence/quality]" is getting lower and _fast_.

Saying that models are getting more expensive is confusing the absolute value spent with the value for money.

- [1] https://github.com/deepseek-ai/DeepSeek-V3/tree/main

silveraxe93··on Python 3.13.0 Is Released
Install `uv` then run `uv run --python 3.13 my_script.py`
silveraxe93··on SQL Tips and Tricks
I'll try to give some constructive criticism instead of a drive by pot shot. I'm sorry, it's just that the leading commas make my eyes bleed and I really hope the industry moves away from it.

On point 3: What I do is use CTEs to create intermediate columns (with good names) and then a final one creating the final column. It's way more readable.

```sql

with intermediate as (

select

  DATEDIFF(DAY, timeslot_date, CURRENT_DATE()) > 7 as days_7_difference,

  DATEDIFF(DAY, timeslot_date, CURRENT_DATE()) >= 29 as days_29_difference,

  LAG(overnight_fta_share, 1) OVER (PARTITION BY timeslot_date, timeslot_channel ORDER BY timeslot_activity) as overnight_fta_share_1_lag,

  LAG(overnight_fta_share, 2) OVER (PARTITION BY timeslot_date, timeslot_channel ORDER BY timeslot_activity)as overnight_fta_share_2_lag
from timeslot_data)

select

  iff(days_7_difference, overnight_fta_share_1_lag, null) as C7_fta_share,

  iff(days_29_difference, overnight_fta_share_2_lag, null) as C28_fta_share
from intermediate ```
silveraxe93··on SQL Tips and Tricks
Yeah, unfortunately you're right that they are real conventions. Quite common too.

I also _understand_ why they exist. It's simple: It makes code marginally easier to write.

But writing confusing, unintuitive and honestly plain ugly code. Just so you can save a second after clicking run and the compiler tells you the mistake is a bad reason.

silveraxe93··on SQL Tips and Tricks
> I want to scratch my eyes out every time I see someone formatting with comma starting the lines

Right!? I _physically_ recoil every time I see that. I think that's the clearest example of normalisation of deviance [1] I know. Seems like anyone that enters the industry straight from data instead of moving from a more (software) engineering background gets used to this.

And the arguments in favour are always so weak! - It's easier to comment out lines - Easier to not miss a comma

Those are picked up in seconds by the compiler. And are a tiny help in writing code vs violating a core writing convention from basically every other language.

[1]- https://danluu.com/wat/

silveraxe93··on SQL Tips and Tricks
The "readability" section has 3 examples. The first 2 are literally sacrificing readability so it's easier to write, and the last has an unreadable abomination that indenting is really not doing much.
silveraxe93··on Girls in Tech closes its doors after 17 years
You're on a boat with a hole in the bottom. The water is rushing in. You grab a bucket and keep scooping water out, but not as fast as it rushes in.

Throwing water out of boats do not make it more buoyant.

silveraxe93··on UK General Election called for July 4th
The optimal time to call an election is when public support is highest. If you expect it to go up, you wait. If you think it will go down, call it today.
silveraxe93··on Oracle dumps Terraform for OpenTofu
I think you mean hypocritical, instead of hubris.
silveraxe93··on GPUs Go Brrr
NVIDIA is so damn good at its job that it took over the market. There's no regulatory or similar barriers to entry. It's literally that they do a damn good job and the competition can't be as good.

You look at that and want to take a sledgehammer to a golden goose? I don't get these people

← PreviousPage 2 of 6Next →