748 karma · joined October 17, 2018
Feel free to reach out to me at felix.faust@everfind.ai :)
zoom out, then play around with cols +/-1 and observe the pattern change. I observe the pattern from -7 to +5; same on #1-200-420
TRANSCRIPTION_PROMPT = """Task: Transcribe the page from the provided book image.
- Reproduce the text exactly as it appears, without adding or omitting anything. - Use Markdown syntax to preserve the original formatting (e.g., headings, bold, italics, lists). - Do not include triple backticks (```) or any other code block markers in your response, unless the page contains code. - Do not include any headers or footers (for example, page numbers). - If the page contains an image, or a diagram, describe it in detail. Enclose the description in an <image> tag. For example:
<image> This is an image of a cat. </image>
"""
Feeding an empty prompt to a model can be quite revealing on what data it was trained on
https://github.com/BlueFalconHD/apple_generative_model_safet...
idk whats the hype about gemini, it's really not that good imho
At retrieval time, our approach involves a broad "prefetching" step: we quickly identify the most relevant schemas, perform targeted vector searches within these schemas, and then rerank the top results using the LLM before agentic reasoning and execution. The LLM is provided with carefully pre-selected tools and fields, empowering it to dive deeper into prefetched results or explore alternate queries dynamically. This method significantly boosts RAG pipeline performance, ensuring both speed and relevance.
Additionally, by limiting visibility of the "agentic execution context" to just the current operation span and collapsing it in subsequent interactions, we keep context sizes manageable, further enhancing responsiveness and scalability.
Bottom line: Googles "least-privilege" rhetoric sounds noble, but in practice it gives Big Tech first-party apps privileged access while forcing independent vendors to ship half-working products - or get kicked out of the Play Store. The result is users lose features and choices, and small devs burn countless hours arguing with a copy-paste policy bot.
edit: I think I read HN comments more than HN articles. Interesting
At refind.ai (https://refind.ai/), we’ve been working on a similar challenge of navigating large volumes of documents, but with a more AI-driven and user centric approach. Beyond just search, we focus on transforming unstructured data into structured insights. This includes features like automatic metadata extraction, natural language search, and integrations with external systems like email and CRMs.
To those who struggle with file management at scale, especially in environments where tools like Spotlight or Windows Search fall short, I think there’s a lot of potential for tools like Buzee to evolve further. If anyone wants to collaborate or learn more about our approach, feel free to reach out!
Best of luck with Buzee’s journey—I’d love to see how the community builds on it!