HNHacker News
TopNewBestAskShowJobs

eclipsetheworld

362 karma · joined November 18, 2016

submissionscomments
eclipsetheworld··on A €0.01 bank transfer could compromise a banking AI agent
I have been working on this issue for a bit, and the most interesting approach I have seen so far comes from the research domain of information-flow control, specifically Microsoft’s FIDES work.

The idea is not to distinguish instructions from data. It is closer to having different privilege levels. Not all code has to run in kernel space, some code runs in unprivileged user space. So what is the equivalent for LLM agents?

In FIDES-style systems, every piece of information that enters the agent context is labeled along two dimensions: integrity and confidentiality. Integrity captures whether the data is trusted or untrusted (i.e. could it contain a prompt injection attack). Confidentiality captures who is allowed to see or receive it [0].

The privileged agent, sometimes called the planning agent, should not directly see untrusted data because it would be susceptible to prompt injection attacks. In the article’s example, a bank transaction’s sender-supplied reference would be untrusted. Instead, the planning agent receives a variable token. It can then either delegate processing of that variable to an unprivileged / quarantined agent with no or limited tool access, or pass the token as a reference to a tool.

Tools then have policies attached to their arguments and outputs. These policies specify which integrity and confidentiality levels are allowed, and whether the tool call may proceed. The policy also determines how the result should be labeled.

For example:

1. High-confidentiality data should not be allowed to flow into a `send_email` tool call addressed to an external recipient.

2. A tool call whose result depends on untrusted input should generally produce untrusted output.

3. A sensitive side-effecting tool should be able to reject calls that are influenced by untrusted context.

So the answer to “how do you separate data from instructions?” may be: you do not rely on the model to do that separation. You track provenance and privilege outside the model, and then enforce the security policy at the tool boundary.

[0] In the simplest implementation, confidentiality is assessed with a binary low/high value, however, in a more advanced implementation, confidentiality can be represented as the set of users or principals allowed to learn that information.

eclipsetheworld··on OpenAI Is Preparing to File for an IPO Soon
Collectively, Alphabet (Google), Amazon, Microsoft, and Nvidia already own approximately 25 - 35% of OpenAI and Anthropic respectively. They already are a part of your portfolio.
eclipsetheworld··on Software factories and the agentic moment
I have been working on my own "Digital Twins Universe" because 3rd-party SaaS tools often block the tight feedback loops required for long-horizon agentic coding. Unlike Stripe, which offers a full-featured environment usable in both development and staging, most B2B SaaS companies lack adequate fidelity (e.g., missing webhooks in local dev) or even a basic staging environment.

Taking the time to point a coding agent towards the public (or even private) API of a B2B SaaS app to generate a working (partial) clone is effectively "unblocking" the agent. I wouldn't be surprised if a "DTU-hub" eventually gains traction for publishing and sharing these digital twins.

I would love to hear more about your learnings from building these digital twins. How do you handle API drift? Also, how do you handle statefulness within the twins? Do you test for divergence? For example, do you compare responses from the live third-party service against the Digital Twin to check for parity?

eclipsetheworld··on Article by article, how Big Tech shaped the EU's roll-back of digital rights
I think you’re conflating security with compliance.

If the goal is to stop breaches, we should mandate MFA and ban default-public cloud buckets. Those are technical solutions. GDPR, instead, mandates a massive administrative layer. No data breach has ever been stopped by a well-drafted Privacy Impact Assessment or a 50-page DPA. Those are legal shields, not security measures.

> then don't automate them: just add it to your DPO's job description.

The DPO isn't an engineer. To let them fulfill a request, I still have to build the internal tooling to query, redact, and export data from distributed production databases. Also, "I'll have my DPO do it manually" never sounds good when going through an audit.

> they may simply be being kind.

The simpler explanation is that the average person has no clue what these rights are because they’ve never had a reason to care. In healthcare, patients care that their data is secure and the service works. They aren't losing sleep over "data portability."

Ultimately, this "level playing field" only benefits incumbents. Unethical players ignore the rules until they’re caught, while legitimate startups are hit with a compliance tax that makes it nearly impossible to compete with US-based firms that can focus 100% of their energy on the product.

eclipsetheworld··on Article by article, how Big Tech shaped the EU's roll-back of digital rights
Thanks for the comment. It actually perfectly illustrates my point. Most people equate GDPR with a "Delete My Account" button, but that’s just the tip of the iceberg.

We didn't spend thousands of hours on a deletion feature (or just development time). We spent them in total to be compliant in a healthcare environment. That time goes into:

Documenting the entire lifecycle (how, why, and where) of every single data point we process. Conducting and documenting formal risk assessments for every major processing activity (Privacy Impact Assessments (DPIA)). Drafting and negotiating data processing agreements (DPAs) with every single partner and vendor we use. Building strict role-based access and logging systems to track exactly who views and edits data and why. Implementing pseudonymization and logical data separation to ensure we meet "privacy by design" standards. Constantly coordinating between the product and dev team and the DPO to update policies and communicate changes to users.

The point I’m making is that the EU has built an incredibly expensive regulatory environment to support rights that, in practice, the vast majority of users don't seem to care about. We’re over-engineering for a "loss of control" that the average user hasn't shown much interest in reclaiming.

eclipsetheworld··on Article by article, how Big Tech shaped the EU's roll-back of digital rights
As a European founder building startups since 2015, I’ve spent a massive chunk of my career navigating the "alphabet soup" of EU regulation: GDPR, DSA, DMA, AI Act, CSRD, SFDR, CBAM... the list is exhausting.

While the goals are usually noble, I’m increasingly convinced we’re regulating ourselves into irrelevance. I’m not a Big Tech company yet my interests align with theirs. We desperately need an EU that prioritizes actual growth over well-intentioned paperwork. To me, the AI Act and the GDPR are the worst offenders here, representing the largest possible gap between "good intentions" and the actual effect they have on the ground.

Consider frontier LLM labs. We have the talent, the Nordic data centers, and access to the GPUs. But why would any investor drop $100B on a frontier LLM lab here when the legislative environment is fundamentally more hostile than the US? It feels like we’ve already watched Mistral and Aleph Alpha get left in the dust.

To give you an idea of the "compliance vs. reality" GDPR gap: I worked on a project processing healthcare data for millions of people. We had a clear, easy-to-find privacy policy and a responsive DPO. Total GDPR requests for info or deletion? Exactly 53. Out of millions. We spent thousands of hours building systems for rights that only 0.001% of our users cared to use.

If you look at the courts, the "damage" being prevented is equally vague. Since EU courts don't really do punitive damages, most awards are tiny unless there’s actual identity theft. Most of what GDPR protects is "mental distress" or "loss of control"-concepts so ambiguous that courts rarely award anything for them unless something else went wrong.

The result of all this "protection"? No FAANG-equivalent, no frontier AI leader, and no homegrown ad-tech. It turns out the most perfectly regulated company is the one that never exists in the first place.

eclipsetheworld··on Agent design is still hard
Interestingly, sticking to the "Agent = REPL" mental model is actually what helped me solve those specific scaling problems (sub-agents and shared data) without the SDK bloat.

1. Sub-agents are just stack frames. When the main loop encounters a complex task, it "pushes" a new scope (a sub-agent with a fresh, empty context). That sub-agent runs its own REPL loop, returns only the clean result with out any context pollution and is then "popped".

2. Shared Data is the heap. Instead of stuffing "shared data" into the context window (which is expensive and confusing), I pass a shared state object by reference. Agents read/write to the heap via tools, but they only pass "pointers" in the conversation history. In the beginning this was just a Python dictionary and the "pointers" were keys.

My issue with the heavy SDKs isn't that they try to solve these problems, but that they often abstract away the state management. I’ve found that explicitly managing the "stack" (context) and "heap" (artifacts) makes the system much easier to debug.

eclipsetheworld··on Agent design is still hard
We're repeating the same overengineering cycle we saw with early LangChain/RAG stacks. Just a couple of months ago the term agent was hard to define, but I've realized the best mental model is just a standard REPL:

Read: Gather context (user input + tool outputs). Eval: LLM inference (decides: do I need a tool, or am I done?). Print: Execute the tool (the side effect) or return the answer. Loop: Feed the result back into the context window.

Rolling a lightweight implementation around this concept has been significantly more robust for me than fighting with the abstractions in the heavy-weight SDKs.

eclipsetheworld··on One thing Tesla and Comma.ai overlooked in self-driving
It depends on how you define “safer.” Cities have a higher frequency of accidents, but with lower severity. Highways have a lower frequency of accidents, but with higher severity.

So in this case, you probably want to opt for accidents of lower severity. Metal undents more easily than flesh.

eclipsetheworld··on One thing Tesla and Comma.ai overlooked in self-driving
Yeah, they should have called them "human fleet response agents" instead. [0]

[0] https://waymo.com/blog/2024/05/fleet-response/

eclipsetheworld··on DeepL Voice: Real-time voice translations for global collaboration
I agree, LLM translations are not only more convenient but also much more capable. I often find myself giving instructions on how to translate text, such as asking the LLM to use formal language in the target language or to apply specific gender-neutral wording. Additionally, it can translate text while preserving the structure (e.g. values in a JSON object) or even adapt to a new target structure. It's just so much more convenient.
eclipsetheworld··on B2C billing is harder than B2B billing
Lemonsqueezy, Gumroad, and Paddle also act as the merchants of record, meaning they assume liability for every transaction. Their role extends beyond simply handling invoicing.
eclipsetheworld··on Perplexity AI's new tool for researching the stock market
While I agree with your statement and recognize that, for now, Perplexity has only introduced a financial information platform comparable to Google Finance or Yahoo Finance, the true value of any forward-looking financial model is rooted in the depth of the qualitative research supporting it.

Building a useful forward looking financial model mostly involves qualitative analysis. This means thoroughly examining the company's and competitors' 10-Ks and 10-Qs, digesting industry reports, understanding the company’s business model, breaking down the underlying mechanics of the income statement, balance sheet, and cash flow statement, identifying the core processes driving value creation, forming solid hypotheses on how the business will evolve, etc.

I believe Perplexity, as an advanced answering engine, provides a strong foundation for supporting this kind of in-depth research and hope to see the platform evolve into this direction.

eclipsetheworld··on Lufthansa brings the inflight experience to life with mixed reality
The international business class is simply inferior to most of its competitors. Other airlines offer larger seats (4 seats per row vs. 6 seats per row). Their onboard entertainment systems use more modern hardware, provide a better selection of content, and sometimes even integrate better with personal devices. The culinary options are less refined. Additionally, Lufthansa still charges $27 for in-flight Wi-Fi. Their lounges are also subpar — smaller, more crowded, and with fewer culinary options. This is based on my personal experience flying with Lufthansa, Austrian, Swiss (all part of the same group), British Airways, Singapore Airlines, and Turkish Airlines. However, other business travelers echo these sentiments. It's noticeable that they enjoy a lack of competition in their home country of Germany, since some healthy competition would force them to actually improve their business class offering.
eclipsetheworld··on Lufthansa brings the inflight experience to life with mixed reality
Lufthansa will do anything for its business class except make it any good.
eclipsetheworld··on Meta does everything OpenAI should be
> [...] Meta (or Facebook) democratises AI/ML much more than OpenAI, which was originally founded and primarily funded for this purpose. [...]

I believe this statement is accurate. Your comment does not alter this fact and merely imposes an arbitrary requirement instead of giving credit where credit is due.

If another company were to openly share alternatives to Meta's core assets, I would welcome that as well.

eclipsetheworld··on Show HN: I made a multiple runtime version manager that can be used on Windows
This looks really cool! Working on multiple projects with different versions before tools like pyenv or nvm existed was a real challenge. As someone working with different programming languages as well, this tool looks like the next logical step.
eclipsetheworld··on How Hertz’s bet on Teslas went sideways
The comments in this thread appear to overlook a crucial point. The business model of rental companies fundamentally revolves around earning more from vehicles than the depreciation costs incurred before resale. However, the recent price reductions by Tesla, coupled with the accelerated depreciation of electric vehicles (EVs) due to rapidly evolving technology and increasing competition, seem to clash with this business model.
eclipsetheworld··on Paris preserves its mixed society by pouring billions into public housing
Allocating public housing could be done through an auction system. Bidders would submit offers for annual rent. The surplus, after deducting costs, could then be allocated to buying or building additional housing for this program. Alternatively, it could be used directly to subsidize rent for low-income individuals.

In the end, this would solve the allocation problem while maximizing the available public housing. It would take a couple of years or decades to reach an equilibrium state I guess.

I'm probably missing something obvious here. Can somebody point out my mistake?

eclipsetheworld··on Apple unveils App Store Award winners, the best apps and games of 2023
https://developer.apple.com/app-store/app-store-awards-2023/
eclipsetheworld··on The Tailscale Universal Docker Mod
I just love that this blog post includes an AI-generated image with the caption of course being the name of the model and the given prompt.
eclipsetheworld··on Low-Background Steel
HUMAN: Write a short piece for an HN post about how low-background steel might be a good analog for a word to describe text and image content before the advent of large generative AI models.

AI:

As the world of artificial intelligence (AI) continues to advance, the impact of large generative AI models like GPT-4 cannot be overstated. However, before the advent of these powerful models, there was a time when text and image content were created with a different approach. Just as low-background steel is used in sensitive scientific experiments due to its minimal radiation contamination, the content generated before the rise of GPT can be likened to "low-background steel" - free from the influence of large-scale AI models.

Low-background steel refers to steel that has been carefully selected and processed to have low levels of radioactive contamination, making it ideal for experiments that require precise measurements. Similarly, content created before the era of GPT was shaped by human creativity, expertise, and personal touch, without the reliance on massive AI models.

Just like low-background steel retains its inherent properties without interference from external factors, pre-GPT content carries the unique perspectives, creativity, and individuality of human creators. It reflects the diverse range of voices and ideas, and embodies the authenticity and originality of human expression.

As we continue to harness the power of large generative AI models, it's important to appreciate the value of "low-background" content created before their advent. It serves as a reminder of the human touch in content creation and the unique qualities that make it distinct from AI-generated content. Just as low-background steel is treasured for its purity in scientific experiments, pre-GPT content holds a special place in the history of creative expression, and it's worth acknowledging its significance in our rapidly evolving AI landscape.

In conclusion, low-background steel serves as an apt analogy for pre-GPT text and image content. It represents the authentic, human-driven approach to content creation before the rise of large generative AI models. As we navigate the future of AI, let's remember the value and richness of content created by human ingenuity and creativity, and the unique insights it brings to the digital world.

eclipsetheworld··on Samsung considers moving to Bing as default search engine
This article seems a little hyperbolic since I assume the default search engine will simply be the highest bidder...
eclipsetheworld··on The Twitter API is now effectively unmaintained
Performance-focused advertisers. Growing this segment was already set as a strategic direction for Twitter in 2021, however, Musk's takeover and the subsequent departure of brand-focused advertisers seem to have accelerated these plans [0].

As for brand-focused advertisers, I can see parallels to 2017's "YouTube Adpocalypse" in which advertisers paused their spending due to their ads being displayed before videos containing offensive content. They recovered from it by demonitizing offensive content. Twitter will have to address "brand safety" concerns, however,

[0] https://www.wsj.com/articles/have-the-ads-in-your-twitter-fe...

eclipsetheworld··on The Twitter API is now effectively unmaintained
I'll take the contrarian stance here. During the Morgan Stanley Conference interview, Elon Musk forecasted that Twitter had the potential to achieve a positive cash flow status by the second quarter of 2023. Advertisers are returning to the platform and they just announced their new API monetization plans [0]. The company's strategic direction notwithstanding, Twitter may be a profitable company soon.

[0] https://twitter.com/TwitterDev/status/1641222782594990080

eclipsetheworld··on FDIC Takes over Silicon Valley Bank
I'm wondering the same thing. The current rise in interest rates must have been a scenario that the bank considered. How can we still have a system that allows situations like this to happen? Other than negligence or malpractice, I cannot fathom a reasonable explanation as to how this happened / was allowed to happen (again).
eclipsetheworld··on Toolformer: Language Models Can Teach Themselves to Use Tools
I was waiting for a proof of concept like this! IMHO the next-wave of GPT-3 productization will involve mapping problem and solution domains to a text-based intermediary format so that GPT-3's generalization abilities can be applied to these problems.
eclipsetheworld··on DetectGPT: Zero-Shot Machine-Generated Text Detection
This approach seems to require knowledge of which LLM was used to generate the given text. I wonder if e.g. model fine-tuning - as already provided by Open AI [0] - could evade this detection approach.

[0] https://beta.openai.com/docs/guides/fine-tuning

eclipsetheworld··on U.S. moves to bar noncompete agreements in labor contracts
In Germany we have non-competes, however, the employer has to continue paying the ex-employee (a part of) their salary for the non-compete to have any effect.
eclipsetheworld··on Princeton researcher apologizes for GDPR/CCPA email study
As soon as the client's IP address touches your server you are processing personal information. E.g. I have seen many webserver which save these in their access logs.

Again, this is the reality of GDPR. It is not okay to operate a website serving EU visitors without considering GDPR implications. This is how GDPR is intended. Don't operate a website serving EU visitors if you don't have a plan on how to respond to these emails. I'm not trying to be harsh or dissuade these small websites from operating. It is just the reality of GDPR.

Page 1 of 2Next →