HNHacker News
TopNewBestAskShowJobs

qudent

105 karma · joined January 3, 2024

submissionscomments
qudent··on Google API keys weren't secrets, but then Gemini changed the rules
I think the fact that it is not possible to put hard spending caps on API keys might be ruled illegal by some EU court soon enough, at least when they sell to consumers (given the explosion of vibecoding end-users making some apps). When I use OpenAI, Openrouter etc., I can put 10 $ on my API key, and when the key leaks, someone can use these 10 $ and that's it. With Google, there is no way to do that - there are extremely complicated "billing alerts" https://firebase.google.com/docs/projects/billing/advanced-b... , but these are time-delayed e-mails and there is no out of the box way to do the straightforward thing, which is to actually turn off the tap automatically once a budget is spent. The only native way to set a limit enforced immediately is by rate limiting - but I didn't see params which made it safe while usable in my case.

(a legal angle might be the Unfair Contract Terms Directive in the EU, though plenty of individual countries have their own laws that may apply to my understanding. A quite equivalent situation were the "bill shock" situations for mobile phone users, where people went on vacation and arrived home to an outrageously high roaming bill that they didn't understand they incurred. This is also limited today in the EU; by law, the service must be stopped after a certain charge is incurred)

qudent··on Google API keys weren't secrets, but then Gemini changed the rules
In Google AI Studio, Google documentation encourages to deploy vibecoded apps with an open proxy that allow equivalent AI billing abuse - giving the impression that the API key were secure because it is behind a proxy. Even an app with 0 AI features exposes dollars-per-query video models unless the key is manually scoped. Vulnerable apps (all apps deployed from AI Studio) are easily found by searching Google, Twitter or Hacker News. https://github.com/qudent/qudent.github.io/blob/master/_post...
qudent··on Tell HN: Google AI Studio docs encourage Google-discoverable open wallets
Google AI Studio documentation encourages developers to deploy vibecoded apps, claiming the API key is secure because it is protected by a proxy - however, there are no checks on the open proxy the deployed app exposes, which allows anyone to use the developer's wallet for arbitrary queries. Vulnerable live endpoints are discoverable by a single google search for us-west1.run.app . The proxy processes Gemini requests even if the deployed website has no AI features itself. Not even a documentation update 2.5 months after reporting.
qudent··on Two different tricks for fast LLM inference
Because inference is autoregressive (token n is an input for predicting token n+1), the forward pass for token n+1 cannot start until token n is complete. For a single stream, throughput is the inverse of latency (T = 1/L). Consequently, any increase in latency for the next token directly reduces the tokens/sec for the individual user.
qudent··on Two different tricks for fast LLM inference
It does affect the throughput for an individual user because you need all output tokens up to n to generate output token n+1
qudent··on Google AI Studio's API key protection is as exposed as the key itself
Google AI Studio's Build Mode hides API keys behind a proxy during deployment, which the docs imply is secure. But the proxy forwards arbitrary requests to any Gemini model with zero auth, quota or validation, using the real API schema, even for apps with no AI features. Deployment URLs are discoverable by searching the URL scheme. This was reported to Google in late November and classified as a documentation issue.
qudent··on Show HN: Claude skill+design pattern for managing worktrees for parallel agents
The "correct" way to have agents working on repositories in parallel - which you want to do to speed up AI development rather than just waiting around - is to use git worktrees, which are separate directories checking out different branches of a repo - so agents can work in parallel on different directories without interference. This also allows comparing different prompts, frameworks, agents etc.

But manually, I always found it hard to manage the complexity of creating a new branch, merging, propagating updates in time for the agents without creating a mess. So I created these super-simple wrapper functions to create a new "child" branch+worktree, runs pnpm install (which uses hardlinks to speed things up), merges from the parent worktree branch, merge to the parent worktree branch etc. This script also imposes a predictable directory structure on the worktrees, which allows understanding easily which task is the "parent" of another task.

The whole thing is a Claude skill. Check out the README.md and the skill description itself at https://github.com/qudent/parallel-working-made-simple/blob/... !