HNHacker News
TopNewBestAskShowJobs

KaoruAoiShiho

3,254 karma · joined October 1, 2011

Looking for investors & partners: kaoruaoishiho@gmail.com
submissionscomments
KaoruAoiShiho··on Ember-1
https://www.reddit.com/r/LocalLLaMA/comments/1wj3s31/thank_y...
KaoruAoiShiho··on ARC-AGI Leaderboard
Appears to be benchmaxxing

https://x.com/quietnning/status/2080786711861407883

KaoruAoiShiho··on Nvidia, CoreWeave, and Nebius: Inside the Circular Financing of the GPU Boom
You're probably just responding to the headline but this person is an AI bull and isn't claiming it's a big deal, she's going into it and explaining it.
KaoruAoiShiho··on GPT‑Live
Click through to the link, the answer is no it uses the latest gpt models now.
KaoruAoiShiho··on Most arguments are about ego, not ideas
I think the parent's point is that if you are genuinely open to losing, the arguments can be productive because you can learn something instead... So stopping arguments is just another way of closing yourself off.
KaoruAoiShiho··on GLM 5.2 beats Claude in our benchmarks
And before you know it, you invented some openrouter provider from first principles...
KaoruAoiShiho··on GLM 5.2 beats Claude in our benchmarks
Are you sure fireworks is unquant? It's not listing precision on openrouter like everyone else.
KaoruAoiShiho··on GLM-5.2: The Most Powerful Open Model yet and the Brutal Reality of Running It
Terrible zero value article, I am extremely surprised it is upvoted.

That being said Artificial Analysis just came out with a brand new benchmark where it scored between opus 4.8 and gpt-5.5 and well behind fable-5 so it's definitely frontier-ish https://x.com/ArtificialAnlys/status/2067744637155226101

KaoruAoiShiho··on GLM-5.2 is the new leading open weights model on Artificial Analysis
This is really held back by one bench (omniscience accuracy) where it's really very far behind otherwise i think it's got at least a couple of points higher.
KaoruAoiShiho··on Running local models is good now
Fable largely fixed the annoying chatterness so sucks that it's gone now.
KaoruAoiShiho··on Show HN: Number Gacha, a gacha game distilled to its essence
Then that's just a video game might as well as play a video game why limit yourself to still confined to the rules of chess.
KaoruAoiShiho··on Cursor Introduces Composer 2.5
Kimi 2.5 has the best long context. For raw coding benchmark scores you can just post train on top of it with more specialized data. 2.5 is kinda old, 2.6 is the current release which is exactly just that and catches up to the frontier in most aspects.
KaoruAoiShiho··on Show HN: Number Gacha, a gacha game distilled to its essence
Do people like "gacha"? I thought people played games for the game experience, story, etc, and the gacha is just the monetization mechanic. It's like making a big deal out of paying $20 bucks a month, or buying loads of DLCs for example.
KaoruAoiShiho··on China blocks Meta's acquisition of AI startup Manus
No I think the best agent with hundreds of millions in ARR should be worth more than the 15th best model company with tiny revenue. ur the joke.
KaoruAoiShiho··on China blocks Meta's acquisition of AI startup Manus
Manus is saved, 2 billion is such an undervaluation considering much worse companies like minimax is valued at 30 billion.
KaoruAoiShiho··on DeepSeek v4
SOTA MRCR (or would've been a few hours earlier... beaten by 5.5), I've long thought of this as the most important non-agentic benchmark, so this is especially impressive. Beats Opus 4.7 here
KaoruAoiShiho··on Kimi K2.6: Advancing open-source coding
Huh, that's not a thing?
KaoruAoiShiho··on Claude Opus 4.7
Might be sticking with 4.6 it's only been 20 minutes of using 4.7 and there are annoyances I didn't face with 4.6 what the heck. Huge downgrade on MRCR too....

256K:

- Opus 4.6: 91.9% - Opus 4.7: 59.2%

1M:

- Opus 4.6: 78.3% - Opus 4.7: 32.2%

KaoruAoiShiho··on Pro Max 5x quota exhausted in 1.5 hours despite moderate usage
Talking nonsense.
KaoruAoiShiho··on Pro Max 5x quota exhausted in 1.5 hours despite moderate usage
Well they're not public yet so you'll have to put up with rumors. But the numbers are available for companies like DeepSeek say they have an 80% profit margin, so it stands to reason OAI etc would do similar numbers considering they charge much more.
KaoruAoiShiho··on Pro Max 5x quota exhausted in 1.5 hours despite moderate usage
After googling https://www.reddit.com/r/singularity/comments/1psesym/openai...
KaoruAoiShiho··on Why do we tell ourselves scary stories about AI?
TLDR: Writer hasn't heard of agents.
KaoruAoiShiho··on Netflix Prices Went Up Again – I Bought a DVD Player Instead
I feel like netflix is definitely very cheap, with OpenClaw or whatever your favorite agent is, it's trivial to subscribe to watch one show and then have it cancel immediately.
KaoruAoiShiho··on GLM-5.1: Towards Long-Horizon Tasks
Blog post is new but the model is about 2 weeks in public.
KaoruAoiShiho··on GLM-5.1: Towards Long-Horizon Tasks
The non-awesome context window is the sad part, but I think a better harness can deal with this.
KaoruAoiShiho··on Issue: Claude Code is unusable for complex engineering tasks with Feb updates
Can you paste the relevant section in your soul please?
KaoruAoiShiho··on What Claude Code chooses
Buy more GPUs.
KaoruAoiShiho··on I guess I kinda get why people hate AI
Sam Altman gave millions to Andrew Yang for pushign UBI, so they are trying to forewarn and experiment with finding the right solution. Most of the world prefers to shove their heads in the sand though and call them grifters, so of course we'll do nothing until it's catastrophic.
KaoruAoiShiho··on Qwen3-TTS family is now open sourced: Voice design, clone, and generation
Something like this:

Character Name: Marcus Cole Voice Profile: A bright, agile male voice with a natural upward lift, delivering lines at a brisk, energetic pace. Pitch leans high with spark, volume projects clearly—near-shouting at peaks—to convey urgency and excitement. Speech flows seamlessly, fluently, each word sharply defined, riding a current of dynamic rhythm. Background: Longtime broadcast booth announcer for national television, specializing in live interstitials and public engagement spots. His voice bridges segments, rallies action, and keeps momentum alive—from voter drives to entertainment news. Presence: Late 50s, neatly groomed, dressed in a crisp shirt under studio lights. Moves with practiced ease, eyes locked on the script, energy coiled and ready. Personality: Energetic, precise, inherently engaging. He doesn’t just read—he propels. Behind the speed is intent: to inform fast, to move people to act. Whether it’s “text VOTE to 5703” or a star-studded tease, he makes it feel immediate, vital.

KaoruAoiShiho··on Qwen3-TTS family is now open sourced: Voice design, clone, and generation
Have you tried specifying the emotion? There's an option to do so and if it's left empty it wouldn't surprise me if it defaulted to rng instead of bland.
Page 1 of 34Next →