HNHacker News
TopNewBestAskShowJobs

piterrro

451 karma · joined November 20, 2012

https://logdy.dev Real-time logs viewer https://leadjobs.dev/ Engineering Leadership job board

posinsk () gmail.com

submissionscomments
piterrro··on Is sandboxing sufficient to contain rogue agents?
I’m thinking about implementing a Jev like model into an agentic harness I’m building. Still it woildnt be enough since Jev like model woild only judge single actions, the case is that agent can build a rogue strategy step by step where each one in isolation is totally safe but as a whole they make up danger behaviour.

We come down to the question - who observes the agent and how its implemented

piterrro··on GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price
Oh buddy you have no idea what a good plan and agent harness can do with deepseek…
piterrro··on Malleable software: Restoring user agency in a world of locked-down apps (2025)
This whole article is true and to the point but... reality is: most people still don't build their own setups: work, home - most of them follow a sh*tty software or decorate with Ikea. I think this applies only to a small group of users how will be willing to go extra mile and prompt their ideal interface for whatever they need to do.
piterrro··on DeepSeek Elastic Compute (DSec)
12 sandboxes per code is insane, I wonder how many of these sandboxes are idle at a time. Depending on the tasks assigned the resource requirements are different. Compare an agent doing pdf conversion and one responding to a simple question. One is cpu bound the other is mostly network wait.

This is an interesting problem from infra perspective since you cannot predict the workload. On a bigger scale you may get away with forecasts.

Im waiting for tech that elastically allocates cpu/mem without restarting a container.

piterrro··on I don't want the details
Post-mortem, 5-why, retrospection - all this to understand what and how it happened so the org may introduce change so that it doesnt happen again. Then another edge case happens and the cycle repeats. Humans do this cycle all the time - big orgs invented the above structure for this process to be visible and collaborative and so that the rest of the org can learn too. Then you put execs on top, they will come, ask some questions so they can sleep better at night. Corporate playbook, some people hate it, the rest understands it… and plays the game.
piterrro··on Pion, an agent designed to run any company autonomously
Interesting to see these experiments. This is early but imagine in few years there will be companies mostly run by agents with a light overview from a human operator. What then happens to scaling of the bussinesses? I would assume, just like today anyone can vibe code an app, there will be vibecoded bussinesses. Maybe its time to start building infrastructure for these bussinesses instead, its a non existent market yet, but give it a few years.
piterrro··on Mercury 2.5
They can run it on „commodity” hardware as gpus, that gives them ability to change direction fast without burning money on ASIC? The space moves fast so model on ASIC can be outdated in couple of months?
piterrro··on Mercury 2.5
I’m using this model to „rerank” results from vector store. The model is provided a set of results and asked to produce a string of 1s and 0s where the offset reflects the position in the result set. The prompt goes along the line „do this set of result match the provided query X”. Works like a charm, normally I would use a small non reasoning model, but given how Mercury produces the output its blazingly fast - which is what I was optimizing for - not to increase the search latency. It helped improving our search in a way that reranker could get close to.
piterrro··on OUI-1: world's first model for Generative UI
The problem with that is that it only works with simple, least interactive UIs. Each new UI a human will be presented needs to be learned to be ised effectively otherwise a user will be lost.

Having said that, imo, Gen UI only makes sens as a presentation layer - not controls. Unless LLM will be using a set of very well defined and homogenic components like table, forms, small widgets.

piterrro··on OUI-1: world's first model for Generative UI
I think there is a big misunderstanding in the space around what Gen UI is and what its used for. Lots of folks refer to it as a framework for building web apps - its not. Gen UI is a DSL for LLM to build UIs on the fly in a multi turn converstation - those are - throw away, one off interfaces or visualization. The reason for the DSL is pragmatism - standardisation and token savings.

The html/css/js or a react app built by an LLM is not Gen UI.

piterrro··on AI models ran real businesses: They sent $12,431 in fake invoices, lost $3,200
The thing Im missing the most is the goal of this experiment. Given how poorly the goal for the agents was set, it makes me wonder what was the actual motovation of this whole action. Lets get the „make as much money as possible” goal broken down.

Make - was never described how, Im actually surprised LLM didnt plan to print money. As much money - what does it mean? How much is much? As possible - there is no flavour of time, effort, cost and profit for the LLM. Could be even infinite, the result would be the same.

Given that the above goal is closest to „use cheating or unethical actions to create a profit” - I think the authors of it actually expected LLM to go wild.

Also

> Going forward, we plan to recreate this experiment with longer time horizons but using simulated environments instead.

Watch out, they will try that again.

piterrro··on Show HN: TERMy – A fast terminal assistant that does not use LLMs
You can get determinostic output (mostly) by setting the temperature to zero. Using couple of other tricks you can get close to 100% of determinism with LLMs.
piterrro··on Invisible Companies
If you own or work for an invisible company, upvote this comment
piterrro··on P99 0 ms* autocomplete for 240M domain names
I just typed a random sequence of the characters, long enough to be certain such domain doesnt exist. No only, the browser send an autocomplete request for every keystroke but for each request it returned a set of proposed domain names (which I'm 100% certain doesnt exist). At this point, how do I understand which results are legit and which are fake? Also, it would be nice to highlight the typed part in the result set so I can visually see what matches exactly.
piterrro··on SQLite as a Document Database (2020)
I do that in psql and it works really well. But your post got me thinking since I need to find a solution to store content of multiple documents, have a way to do FTS as well as vector similarity. I dont need that for all of the documents at once - I need to do it either for one document or at most couple of documents.

Now I'm thinking I could have a separate database file per "batch", store it in object storage and then download on demand and query it as I want. This way I'll not bloat my primary storage size as well as I dont need a special vector DB since sqlite vector search will be enough for up to 50k vectors.

piterrro··on Show HN: A lightweight, stateless database for agent memory
I get the closed source open binary approach - I would test it if I woild be in a need!
piterrro··on RAG Is Simpler Than You Think
RAG only makes sense if you have an LLM review the results, pick the most relevant ones and iterate further if there's a need running another query and repeating the process. Raw dump of vector search (even with reranking) is asking for troubles (or rather weird user questions like 'why this crap popped up in the results?')
piterrro··on Show HN: Huzzah – a novel approach to coding with AI
You could write that pseudocode as a prompt for the agent and get the same result. Use plan mode to understand what agent wants to do.

Am i missing anything?

piterrro··on PostgreSQL for Everything
true to that - currently using psql (in a single monolithic codebase) as: sql db, json db, vector store, logs store, full-text search, queue, message bus.

multiple processes connected to it.

piterrro··on Geolocating a random island using geometry and CUDA programming
really impressive, could that be the way to locate yourself without GPS? assuming we know more/less where we are
piterrro··on Cloudflare's AI Psychosis
Companies specialized in building DCs and AI labs specialized in niche models?
piterrro··on Cloudflare's AI Psychosis
The problem is, beside their size, they have not moat to compete in AI space. Their data centers are spread across the world in other companys DCs (btw do they own any DC actually from top to bottom?).
piterrro··on Don't classify, hallucinate
I would propose the following, query vector store for 10 closest categories based on a query, feed it to an LLM, in the prompt ask it to produce a single digit 0-9 representing the number of the most appropriate choice. Use plain text prompt, dont inflate token count with JSON. There you go, you just drastically reduced the output pricing.

Additionally you could experiment with a reranker instead of an LLM or after reranking take top-3 results and then feed to LLM as input in order to reduce input token costs.

piterrro··on Choose Boring Technology (2015)
https://grugbrain.dev/ Similar on this topic
piterrro··on Mistral OCR 4.1
You can even run it on your desk if you want, a single gtx 4090 is enough. It can be fully air gapped.
piterrro··on Mistral OCR 4.1
Accuracy can have different dimensions, depends on what you can tolerate and whether you can detect it to apply more powerful methods.

Imagine you have a cheap and 99% accurate ocr. The other 1% you can detect and apply more powerful (more accurate but slower and more expensive) ocr method. What would you use? At scale these things add up.

piterrro··on Mistral OCR 4.1
Most use cases dont need that kind of accuracy, just doesnt justify the 3-4usd range. I build for that exact case (tender documents, we’re processing north of 100k pages per day), it doesnt need to recognize scanned written text from 1930s, its usually pdf/docs/scanned printed pages. The accuracy is great, bounding boxes are must have for proper grounding for building answers by LLMs. Tesseract was too slow and not enough in some cases (for example tables or images which we also recognize and describe)
piterrro··on Mistral OCR 4.1
For anyone interested, I have an ocr pipeline running on rented GPUs, doing around 1000pages for 0.05-01 usd with around 0.8 seconds per page with full bounding boxes support for grounding.

If you’re interested you can find contact to me via this profile.

3.5 usd/1000 pages is just too expensive…

piterrro··on Don't be a meat proxy
In these situations I always use AI to respond
piterrro··on Why we write our own C and C++ inference engines
Could this vllm port be faster to install? Im starting gpu machine multiple times a day and it takes 5 minutes to set vllm up. If Inise this port that time is minimized?
Page 1 of 8Next →