HNHacker News
TopNewBestAskShowJobs

aliljet

1,528 karma · joined February 22, 2016

contact me here: pav.gup@gmail.com
submissionscomments
aliljet··on Gemini 4 Argon (High): Intelligence, Performance and Price Analysis
It's hard to not see this as a gut punch for OpenAI. They're lead was largely captured by scoring on value (by way of reset after reset) and now they're getting eaten up on price and being bestes and equalled on performance. I'll still pay a premium for Opus 5.5 right now because it's nearly unlimited use, but Google is the quiet sleeping king Everyone is happy to watch everyone else, but I'd wager google burns more tokens through their search product than basically anyone else and now they're just quietly pacing the frontier...
aliljet··on The largest electric aircraft just flew [video]
To be clear, this is a hybrid aircraft, but it's still an awesome step forward!
aliljet··on GPT-6 Astra
The ARC-AGI-3 score is ridiculously high. Is this benchmaxxing or something way different? It's really hard to discern how we're approaching breakthroughs...
aliljet··on GPT-6 Astra
The ARCC-AGI-3 performance is absolutely incredible. The magnitude of change here is so high that I'm almost incredulous. Is this real? Did the benchmark get gamed?
aliljet··on ChatGPT Is Throwing 404
This is probably Astra getting ready for release.
aliljet··on At-home test for infected ticks could improve Lyme Disease diagnosis
Does diagnosis offer a better prognosis for those that are infected?
aliljet··on GLM-5.3: Frontier coding with emergent cyber capabilities
This is absolutely still shy of Sol and Fable, but only just by a hair. Ridiculous results. There's still not a compelling economic reason to drop OpenAI courtesy of the ludicrous reset addiction that's taken place, but it feels like we're on the precipice.

How are you all toying with running this kind of thing in a mega quantized way locally? Two weeks out from released weights, but this is still just GLM 5.2 with post-training magic.

aliljet··on Mistral OCR 4.1
Accuracy is truly what people die for in the OCR game. Price isn't the primary function here.. it's an equation of price, accuracy, speed, and in mayn cases regulation.
aliljet··on Grok 4.6 scores 61 on the Artificial Analysis Intelligence Index
Can you explain what you mean? These days courtesy of an addictive reset game OpenAI is playing, I can't find anything with frontier intelligence that's more cost efficient...
aliljet··on Pixel 11 Pro Fold
What genuinely disappointing result. Long time pixel user here and I've been routinely buying these phones with the argument that you're getting the most value of any modern smart phone. Now? I'm just waiting for the pixel 10 family to drop in price. Happy to just wait.
aliljet··on Qwen3.8 Max now ranked as the best overall model by agentic index
Is there a path to distill this model to do very specific things? Like a RAG strategy for a small (or even large) corpus?
aliljet··on Beating GPT-5.6 Sol on retrieval with 100x cheaper open models
There is a more serious question in here that's not being answered. How effective is the retrieval in finding buried needles in larger and larger haystacks. And there's a correlary question, how effective could you be in finding paired needles in that haystack where you need to hold a needle to unlock finding another needle.
aliljet··on Qwen3.8-Max: A New Bar for Coding and Cowork
I'm trying and failing to find value running a potential Qwen 3.8 27b dense model on a 16 core, 128 GB of ram, 2080ti box. Yes, the GPU yells for help, but the problem is that no math works to upgrade this machine even when pouring $200 in rent every month into the large model providers...

How are you all justifying economical use of these local models right now? What's the cost efficient way to do this and do better (even with models evolving over time and losing now vs later) than the big labs?

aliljet··on RTX 2080 Ti Memory Upgrade to 22 GB
Honestly, I have a 2080ti that I use to play and I can tell you the math isn't there to upgrade it. It's much easier to just find a 3090/4090/5090 and keep pace with the software and hardware simultaneously.
aliljet··on Show HN: Getting GLM 5.2 running on my slow computer
I'd be curious about an.option that would allow glm use with a low end GPU like a 2080 ti...
aliljet··on Qwen-AgentWorld: Language World Models for General Agents
The benchmarks here are confusing at best. Am I reading correctly that this model is essentially as good or better than all frontier models right now?
aliljet··on Mistral OCR 4
I was just using infinity parser 2 (flash, to be fair) for pennies self-hosted to run through thousands of pages of documents with remarkable confidence. I decided to use https://huggingface.co/datasets/allenai/olmOCR-bench to determine what was the best OCR tool, yesterday, but I've got no idea what the best is now. What is the dominant OCR eval right now? Between Baidu and Mistral this morning, I wonder if there's a new tool to switch to..
aliljet··on Unlimited OCR: One-Shot Long-Horizon Parsing
I'm curious about this. What models/tools have you been using?
aliljet··on Unlimited OCR: One-shot long-horizon parsing
How does this compare with infinty parser 2 which seemed to be running the table on every other OCR tool (https://huggingface.co/datasets/allenai/olmOCR-bench). To be fair, there's no single winning OCR benchmark and this isn't showing up anywhere yet..
aliljet··on Qwen-Robot Suite: A Foundation Model Suite for Physical World Intelligence
This sounds incredible. Have these models effectively solved the problem of trying to use a fast-processing network to predict the world's state ahead? For example, to catch a ball?
aliljet··on Running local models is good now
The problem here is always the cost-benefit. For $200/mo, you're receiving subsidized best of breed access. There's no model competing for that price anywhere. If a 27B param model is what you choose, show me your hardware! I would love to be wrong...
aliljet··on Expanding Project Glasswing
Is this just one giant marketing plot?
aliljet··on Qwen3.7-Max: The Agent Frontier
Where can a user reasonably host this in an affordable way to access the local LLM revolution?
aliljet··on Gemini 3.5 Flash: frontier intelligence with action
I'm really running into this deep at the edges of content creation. Take, for example, a need to general some kind of legal work. The cost of painstakingly checking and rechecking each case cited is reducing the value of these frontier models immensely.

Coding, however, is solved like magic. Easier to add tests, to be fair.

aliljet··on Gemini 3.5 Flash
Is there a good benchmark tracking hallucinations? The models are all incredibly good now, even the open ones, and my hope is that the rate of hallucinations is something that's falling off in concert with larger and larger context lengths.
aliljet··on 1966 Ford Mustang Converted into a Tesla with Working 'Full Self-Driving'
This is so cool. I would love to revitalize a generation of great, but perhaps boring older cars with FSD. Just so much work...
aliljet··on Spirit Airlines Is Winding Down All Operations
Why did Spirit die? Was there any last of this that had to do with their abysmal customer service?
aliljet··on I'm Peter Roberts, immigration attorney who does work for YC and startups. AMA
What systems are you actively using? And what systems have you tried? It seems like law, generally, may be hitting a tipping point on LLM use...
aliljet··on Claude.ai and API unavailable [fixed]
This is a tough moment. Claude is simultaneously becoming substantially more expensive, substantially less reliable (single 9 of reliability), and substantially less performant. It's really hard to justify the cost of a subscription over there right now.
aliljet··on HERMES.md in commit messages causes requests to route to extra usage billing
I wonder how this kind of response from Anthropic is actually being read by the community at large. If you consider the rough sentiment of the r/ClaudeCode subreddit against the r/Codex subreddit, you can see that there is a definite loudness among the folks departing ClaudeCode for Codex. Something big is shifting on the ground, I think.
Page 1 of 9Next →