HNHacker News
TopNewBestAskShowJobs

gillesjacobs

1,965 karma · joined January 29, 2018

[ my public key: https://keybase.io/gillesjacobs; my proof: https://keybase.io/gillesjacobs/sigs/naKVxhthsAZCj70Vsuk8ooWuF4gl1Ugl39KGf77zvrU ]
submissionscomments
gillesjacobs··on Unreal Agent
OP is a coding agent and harness tool named Unreal, not the game engine by Epic.
gillesjacobs··on RTK reports token savings, but our cost benchmarks disagree
Main takeaway:

  Average cost per attempt, without → with RTK:
  
  Claude/Fable: $1.72 → $1.64 (~5% cheaper)
  DeepSeek: $0.115 → $0.121 (~5% more expensive)

  Almost all Claude savings came from a single task.
  Excluding it, savings were under 1%.
It took me a few rereads to parse out the top-line. This article really buries the lede.
gillesjacobs··on Jason Arday, ex-Cambridge professor at centre of plagiarism row, found dead
He was an immoral man taking full advantage of an immoral system.
gillesjacobs··on Pixel Watch 5
The custom complications on WearOs watch faces do it for me. I can see my investment portfolio value change right from my wrist.
gillesjacobs··on 2026 Eclipse Webcams
Cool, there is no check on the direction of the camera. Various camera's I checked do not seem to point in the direction of the sun during totality. It does filter candidate cameras down though.
gillesjacobs··on AI-Generated Images Discourage Me from Reading Your Blog
And I'd rather see the Mona Lisa.
gillesjacobs··on Agent Skill to Force Docs in ASD-STE100 Simplified Technical English
https://youtu.be/uJblcC4lKYw

This video benchmarks slop-style indicators with different skill/prompt solutions including the STE skill vs. George Orwell's six rules of writing prompt: Orwell came out on top overall.

Additional bonus: it doesn't add much more tokens to input context. I have compared prose prompts with these rules and without and I am liking the results.

  1. Never use a metaphor, simile, or other figure of speech which you are used to seeing in print.
  2. Never use a long word where a short one will do.
  3. If it is possible to cut a word out, always cut it out.
  4. Never use the passive where you can use the active.
  5. Never use a foreign phrase, a scientific word, or a jargon word if you can think of an everyday English equivalent.
  6. Break any of these rules sooner than say anything outright barbarous.
gillesjacobs··on Future euro banknote design proposals
They asked to put EU buildings on there to grab power and cultural capital. The EU is centralizing and consolidating by any means necessary.
gillesjacobs··on Claude Opus 5
Because their fear-based marketing gave the US gov justification in blocking them for a while. They wisely didn't do that for Opus 5.
gillesjacobs··on Flux 3
Same person that was mocking the hands in image generation in 2023, is the same person that was saying 'hands are fixed but it can't generate "the red dog jumps over the jump rope held by the blue pelican while juggling 5 balls"' in 2024, is the same person that posted this.
gillesjacobs··on You only need the frontier model for one single edit
In ML, you want to test general capability of a model (generalizability), because you want it to perform well on unseen tasks. In that benchmark, the literal reference is leaking through web search, the agent can see the matching real codebase online and the commits so that's test set leakage. I know no programmer that was ever paid to rewind an existing codebase to a previous commit and implement a feature/fix a bug that exists in the next commits.
gillesjacobs··on Chat Control passed first round in EU Parliament
Except the only need Chat control is fulfilling is the governments and bureaucrats need to control and check all civilian communications.
gillesjacobs··on Munich 1991: The Roots of the Current AI Boom
> this is wildly incorrect in an academic context.

Worked for me in my academic career ¯\_(ツ)_/¯ Of course if you suspect a work is poignant, you should read it, then cite it.

gillesjacobs··on Munich 1991: The Roots of the Current AI Boom
Which work has more value: the abstract description of a catalogue of potential model architectures or their validated application trained on real data?

In the Schmidhuber case their is 20 years and a chain of countless other works in between the two.

gillesjacobs··on Munich 1991: The Roots of the Current AI Boom
Of course, but if you haven't read them you also shouldn't cite them.

And that's where Schmidhuber goes off the rails: publicly shaming published papers into citing you isn't good academic practice. It's bullying.

gillesjacobs··on Nvidia RTX Spark
With some caveats, you wouldn't be able to connect two 4k monitors to a dock without TB5.
gillesjacobs··on The disturbing white paper Red Hat is trying to erase from the internet
https://web.archive.org/web/20260402155236/https://www.redha...

Archive URL to original paper

gillesjacobs··on Cursor Composer 2 is just Kimi K2.5 with RL
I stand corrected, that is pretty scummy.

I bet Moonshot is going to make them open their wallets to avoid legal trouble.

gillesjacobs··on Cursor Composer 2 is just Kimi K2.5 with RL
They probably licensed it. Still a bit deceptive not to mention it on the model card/blog post, but companies whitelabel all the time without mentioning.

It goes against the ML community ethos to obscure it, but is common branding practice.

gillesjacobs··on Cursor Composer 2 is just Kimi K2.5 with RL
Cursor is mostly an IDE / coding-agent harness company. So it probably makes sense for them not to train their own base model, but instead license something like Kimi and fine-tune it for their own harness and workflows.

Their moat looks pretty thin. A VSCode fork with an open-source LLM fork on top. In the fast-moving coding-agent market, it’s not obvious they keep their massive valuation forever.

gillesjacobs··on Scott Adams has died
I liked his cartoons and he did no wrong.
gillesjacobs··on Show HN: Replacing my OS process scheduler with an LLM
You're underselling this as a process manager, it could also be a productivity tool with some prompt changes; Determine procrastination apps: games, non-professional chat, video streaming and kill it.
gillesjacobs··on A Developer Accidentally Found CSAM in AI Data. Google Banned Him for It
https://archive.ph/awvmJ
gillesjacobs··on How to sequence your DNA for <$2k
They save money by cheap labour and batching large quantities for analysis. For the consumer this means long wait times and potentially expired DNA samples.

I tried two samples with Nebula, waited 11 months total. Both samples failed. Got a refund on the service but spent 50usd in postage for the sample kit.

gillesjacobs··on A PM's Guide to AI Agent Architecture
Nice framing for PMs, but technically it is way too rosy. MCP is real but still full of low utility services and security issues, so “skills as plug-ins” is not production ready. A2A protocols were only just announced this year (Google, etc.) and actual inter-agent interoperability is still research grade, with debugging across agents being a nightmare. Orchestration layers (skills, workflows, multi-agent) look clean in diagrams but turn into brittle state machines under load. LLM “confidence scores” are basically uncalibrated logits dressed up as probabilities.

In short: nice industry roadmap, but we are nowhere near robust, trustworthy multi-agent systems yet.

gillesjacobs··on Show HN: PageIndex – Vectorless RAG
I am always on the lookout for new document extraction tools, but can't seem to find any benchmarks for PageIndex-OCR. There are several like OmniDocBench and readoc. So... Got benchmark?
gillesjacobs··on Show HN: PageIndex – Vectorless RAG
Extracting structure and elements from HTML should be trivial and probably has multiple libraries in your programming language of choice. Be happy you have machine-readable semantic documents, that's best-case scenario in NLP. I used to convert the chunks to Markdown as it was more token-efficient and LLMs are often heavily preference trained on Markdown, but not sure with current input pricing and LLM performance gains that matters anymore.

If you have scanned documents, last I checked Gemini Flash was very good cost/performance wise for document extraction. Mistral OCR claims better performance in their benchmarks but people I know used it and other benchmarks beg to differ. Personally I use Azure Document Intelligence a lot for the bounding boxes feature, but Gemini Flash apparently has this covered too.

https://getomni.ai/blog/ocr-benchmark

Sidenote: What you want for RAG is not OCR as-in extracting text. The task for RAG preprocessing is typically called Document Layout Analysis or End-to-End Document Parsing/Extraction.

Good RAG is multimodal and semantic document structure and layout-aware so your pipeline needs to extract and recognize text sections, footers/headers, images, and tables. When working with PDFs you want accurate bounding boxes in your metadata for referring your users to retrieved sources etc.

gillesjacobs··on Show HN: PageIndex – Vectorless RAG
A suspicious lack of any performance metrics on the many standard RAG/QA benchmarks out there, except for their highly fine-tuned and dataset-specific MAFIN2.5 system. I would love the see this approach vs. a similarly well-tuned structured hybrid retriever (vector similarity + text matching) which is the common way of building domain-specific RAG. The FinanceBench GPT4o+Search system never mentions what the retrieval approach is [1,2], so I will have to assume it is the dumbest retriever possible to oversell the improvement.

PageIndex does not state to what degree the semantic structuring is rule-based (document structure) or also inferred by an ML model, in any case structuring chunks using semantic document structure is nothing new and pretty common, as is adding generated titles and summaries to the chunk nodes. But I find it dubious that prompt-based retrieval on structured chunk metadata works robustly, and if it does perform well it is because of the extra work in prompt-engineering done on chunk metadata generation and retrieval. This introduces two LLM-based components that can lead to highly variable output versus a traditional vector chunker and retriever. There are many more knobs to tune in a text prompt and an LLM-based chunker than in a sentence/paragraph chunker and a vector+text similarity hybrid retriever.

You will have to test retrieval and generation performance for your application regardless, but with so many LLM-based components this will lead to increased iteration time and cost vs. embeddings. Advantage of PageIndex is you can make it really domain-specific probably. Claims of improved retrieval time are dubious, vector databases (even with hybrid search) are highly efficient, definitely more efficient that prompting an LLM to select relevant nodes.

1. https://pageindex.ai/blog/Mafin2.5 2. https://github.com/VectifyAI/Mafin2.5-FinanceBench

gillesjacobs··on Belgian CVD is deeply broken
Had many a friend in the Belgian hacker scene who were threatened with legal action after responsible disclosure. To my knowledge, these threats always remained empty: if there is one thing more expensive than engineering a fix, it is starting a lawsuit in Belgium.

It is a sad state-of-affairs that the culture is like this. Ultimately it results in a less secure society, where vulns are anonymously disclosed and shared.

gillesjacobs··on MCP: An (Accidentally) Universal Plugin System
It doesn't do it magically. The "tools" an LLM agent calls to create responses are typically REST APIs for these services.

Previously, many companies gated these APIs but with the MCP AI hype they are incentivized to expose what you can achieve with APIs through an agent service.

Incentives align here: user wants automations on data and actions on a service they are already using, company wants AI marketing, USP in automation features and still gets to control the output of the agent.

Page 1 of 14Next →