HNHacker News
TopNewBestAskShowJobs

Imanari

144 karma · joined April 25, 2018

submissionscomments
Imanari··on GPT-6 Sol and Luna
There are multiple models competing with 5.6sol on AA but none of them have the same feel (intuition,taste,judgement) - actually they are very far behind. I would say open source models are farther behind the the big labs than the benchmarks make you believe.
Imanari··on MiMo v2.6
For simple edits it produces huge reasoning traces. Kind of disappointing. Super long repetitive reasoning always sits wrong with me. It feels like a way for the labs to brute-force higher benchmark scores but not actually like a smarter model. Disappointing, Mimo2.5 was such a nice model.
Imanari··on MiMo v2.6
ish… at least we can be sure they don’t benchmaxx the pelicans lol
Imanari··on I am often wrong
Steps 1–5 of his framework are increasingly formal ways of saying “figure out what’s going on before doing something,” followed by step 6: “then do it fast”
Imanari··on Introducing System One Models and Jev
> AI Map Reduce over Big Data

> Search for relevant information over giant corpuses

Do you mean as an alternative to embeddings?

Imanari··on Introducing System One Models and Jev
Seems like LLM can do everything Jev can do (just structured outputs?) but Jev is highly optimized and purpose built for it and thus way faster and cheaper. Is that a fair description?
Imanari··on Gemini 3.8 Flash and 3.8 Flash Cyber
Is that possible? What subscription would that be? 'Google AI Plus'?
Imanari··on DeepSeek-V4-Pro outperforms Fable 5 after fixing runtime inference control
I find it very intriguing. I reminds me of the insight in the early days that prompting the models to "think step by step" gave big performance boosts. Here we basically tell the models "You have an internal monologue/J-space and you can use it!". Quasi an opaque 'think step by step' prompt. The models were always able to think step by step but needed to be prompted to do so. Maybe the same thing is possible with the J-Space? The huge claims are yet to be replicated/proven, though.
Imanari··on Gemini 3.7 Flash
Testing it in Pi.dev and liking the speed a lot! Huge improvements from past gemini models regarding tool calls and agentic capabilities
Imanari··on How Compaction Works in Pi
Great thread, I was just thinking about compaction. My current line of thought is that compaction/pruning/ctx management in general should be something ongoing and maybe recursive. For example:

User:'How is auth implemented?' -> [thinking] [codebase exploration with [thinking] in between, 10 file reads, 3 of which were "wrong"] [thinking] -> agent_response

This little exchange contains a WHAT (how auth actually is implemented) and a HOW (where that info is and how to retrieve it). Maybe this question was part of a larger task. I think that whole exchange could be summarised before it enters context, kind of like what happens with subagents. The main thread would then consist mostly of [summaries]. Eventually the context will fill up anyway and we would summarise those summaries again. Alternatively one could maintain a [master_summary], kind of like an internal state. So new [summaries] get integrated directly and the [master_summary] gets updated.

Imanari··on How Compaction Works in Pi
I am surprised at 'removes actual tool calls/results, tool output'. Your approach with /prune seems to be 'keep the WHAT, remove the HOW (we got here)'. I would have thought that the HOW contains some useful signal.
Imanari··on Flint: A Visualization Language for the AI Era
What is wrong with "please make plot XYZ in plotly?"
Imanari··on Advancing the price-performance frontier with GPT‑5.6
How do you run 'deep research'?
Imanari··on Kimi K3-256k
Here is what most people miss…
Imanari··on Writing by hand is good for your brain
There is a technique where you keep your hand almost still and move the whole arm instead.

https://youtu.be/-F8SA_QySkc?is=z0P7I8loxuHVwKXn

Imanari··on You only need the frontier model for one single edit
So its basically (a plan + a bunch of file reads + a first edit) injected into the context of a cheap model so the cheap model does not feel like it has to re-read the files and just continues executing and editing?
Imanari··on How version control will evolve for the agent boom
I think a Q&A-based approach is could also be a suitable way to capture the reasoning behind a project. Instead of writing documentation afterward, an AI could interview the developer throughout the process and preserve the questions and answers as project context.

Matt Pocock’s “grill-me” skill is a nice example of this idea: https://github.com/mattpocock/skills/tree/main/skills/produc...

Imanari··on Pruning RAG context down to what the answer actually needs
Did you experiment with the Pruner completely replacing the Reranker?
Imanari··on The worthlessness of Vitamin D is mildly exaggerated
I'll take it ;)
Imanari··on The worthlessness of Vitamin D is mildly exaggerated
I started to take VitD on a whim and my mood and energy sure improved. Also I feel like I get sick less often. 3 friends whom I recommended VitD told me the same. Living in northern germany. YMMV.
Imanari··on GLM 5.2 vs. Opus
This mirrors my experience. I have been using it in Pi. It is smart and output is good but it is not efficient in getting there.
Imanari··on MiMo Code is now released and open-source
Only tangentially related: MiMo-2.5Pro is fast, cheap and very capable, although not quite gpt5.5 level iontelligence (I dont use the claudes). It works flawlessly in Pi and is an excellent workhorse. I expect big things from their next model.
Imanari··on DeepSeek V4 Pro beats GPT-5.5 Pro on precision
I always feel GPT5.5 is better at ‘getting the bigger picture‘ when I am describing something vaguely vs Chinese models. What’s your experience with that?
Imanari··on LLM Paper Trading
What do the model have as inputs? What’s their harness like? Just price data or are they free to pull reports etc. from the web?
Imanari··on Show HN: Forge – Guardrails take an 8B model from 53% to 99% on agentic tasks
Very cool work! Regarding your finding "the tool ran successfully and returned data" and "the tool ran successfully but found nothing." Couldn’t this be solved by designing better tool responses instead of adding another layer in between? Just curious and probing my understanding.
Imanari··on Show HN: Semble – Code search for agents that uses 98% fewer tokens than grep
Good old aider ahead of its time
Imanari··on Agents need control flow, not more prompts
That’s exactly the approach of smolagemts. The only “tool“ available is writing python code
Imanari··on Agents need control flow, not more prompts
As with so many things aider.chat was ahead of its time with its ability to create deterministic scripts.
Imanari··on DeepSeek v4
Like a special parser? Would you mind elaborating?
Imanari··on DeepSeek v4
How can they fix it after the release? They would have to retrain/finetune it further, no?
Page 1 of 5Next →