HNHacker News
TopNewBestAskShowJobs

lisa_coicadan

1 karma · joined July 9, 2025

submissionscomments
lisa_coicadan··on Ask HN: Who is doing the best Word/PDF RAG tool with deep research?
Great thread, we’ve seen the exact same pain points around working with large volumes of complex PDFs/Word docs.

At Retab.com, we focus on the “hard pre-RAG” layer: turning raw documents : including scanned reports, OCR messes, financial statements, or regulatory filings... into clean, structured, model-ready data.

Instead of relying on embeddings over noisy text chunks, we use schema-driven generation, multi-LLM consensus, and an evaluation UI to ensure output is accurate, complete, and explainable. No manual parsing, no hallucinations, just structured JSON (or any format you want), ready for retrieval, agents, or analytics.

We work with teams doing RAG on contracts, audits, earnings reports, etc.. anywhere that “close enough” isn’t good enough. Happy to run your hardest docs through Retab if you want to benchmark against WFGY or LlamaParse

lisa_coicadan··on Show HN: SwellDB – Query AI-generated tables with SQL
Really interesting project, love the idea of skipping traditional ETL by generating structured views on demand.

We’re building something in a similar space at Retab.com, but with a different philosophy: instead of querying live across unstructured sources, we focus on reliably turning raw inputs (PDFs, scanned docs, images, etc.) into clean, structured outputs, using schema-guided LLM generation, multi-model consensus, and an evaluation dashboard. So it’s less about on-the-fly queries, and more about building robust pipelines where you can trust the output and audit how it was produced. Curious if you’ve thought about integrating evaluation or schema validation layers downstream, or if SwellDB is mainly about exploration? Excited to follow the project either way!

lisa_coicadan··on Coding with LLMs in the summer of 2025 – an update
I’ve seen this exact workflow (PDF → extract data → update structured files) come up a lot, and it’s impressive that Claude handled it end-to-end like that. We’ve been building Retab.com to handle those kinds of tasks more reliably, especially when you want structured output (like JSON) from messy documents like PDFs, scans, or even images. Instead of writing ad-hoc scripts or chaining LLM calls, you just upload the file, define what you want (via schema), and it gives you clean structured data. It’s AI-native but deterministic, no need to install PyPDF2 or debug model behavior. Just wanted to share in case others are solving similar problems repeatedly.