HNHacker News
TopNewBestAskShowJobs

TrueDuality

2,337 karma · joined July 21, 2016

submissionscomments
TrueDuality··on Defcon stiffs badge HW vendor, drags FW author offstage during talk
It's a bit of inference on my part, I'll give you that but the premise doesn't make sense if he wasn't paid.

If he was working as an unpaid volunteer with the "compensation" being a part of a talk on stage... What was he protesting before the event even started when he injected the screen asking for money? That's a pretty garbage thing to do, but he did it before the consequence he claimed he was protesting.

TrueDuality··on Defcon stiffs badge HW vendor, drags FW author offstage during talk
Nah there is a difference between fun and light thing to find, and asking for money and disrupting an event.
TrueDuality··on CrowdStrike debacle provides road map of American vulnerabilities to adversaries
Offense is also easy in that there is a ton of software out there, and you just need to find one vulnerability. There is a "win" condition" Defense is impossible as there is a ton of software and you need to protect all of it every time, there is only a "lose" condition.
TrueDuality··on CrowdStrike debacle provides road map of American vulnerabilities to adversaries
> Was there ever such a time? If so then tell me when it was.

It was a goal for a long time, and I'd say we use to be more resilient pre-cloud SaaS auto-update everything. When every software solution installation is on private networks, with fundamentally different architectures (both machine and topology), along with a wide selection of even very poor quality software, was a lot more resilient than what we have today.

Today a single outage in a single service (say AWS) can grind a large number of companies to a halt. A bad update like this one immediately impacts everyone all at once and has a domino effect. That didn't use to happen.

We've been concentrating our collective architecture into a few best practice tools but that all become single points of failure for not only digital attacks, but misconfigurations, mismanagement, company failures, exhausted underpaid engineers, optimizations, etc.

> Hardening systems against vulnerabilities means making them less convenient/easy to use and people instantly balk against that.

This isn't necessarily true, and I'd argue quite the opposite direction has been happening in the security industry over the past decade or so. People realized that hard security would only cause users to find simple predictable bypasses that would overall _weaken_ the security posture. You just have to look at the evolution of NIST recommendations around passwords to see this happening.

Must change a password every 90 days that can't be the same as your last 10 passwords and complex password requirements? Well users are going to use the minimum size in predictable patterns and just increment a number at the end. Those old password hashes you have to keep around to check if the user is reusing the password? Those are a liability that, when broken, tell the attacker which pattern each user is using. Not the case anymore and there is a lot more usable security rolled that is entirely transparent to end users or almost entirely transparent.

Think about how prevalent and bad captchas used to be on the website and how easy they were to circumvent. Cloudflare's and Google's captcha solution are pretty transparent and has much greater efficacy than the old ones.

Did Microsoft's general and on-going laxness contribute to bad security practices? Absolutely, but that is one ecosystem that had weird other by the nature of how inherently unstable that environment was and is not and hasn't except for maybe a brief peak ever been a core foundation of the internet infrastructure, just enterprise infrastructure unfortunately. They definitely never got the memo about usable or transparent security. I hope they're at least trying behind the scenes now.

TrueDuality··on Ask HN: What's Wrong with IRC?
This assumes that there is actually a friendly frontend meeting the equivalent feature and UI/UX experiences of things like Slack and ignoring the cost of maintaining that service by IT.
TrueDuality··on My VM is lighter (and safer) than your container (2017)
There is a lot of discussion on here about the different isolation levels available, but these micro-VMs aren't playing in the same field and can't be compared apples-to-apples.

If you go read the paper this requires a specialized Xen kernel, which in turn requires processor virtualization extensions directly available where you're running these containers. Those extensions aren't generally available if you're already running inside of a VM.

This is a solution that only works on bare metal which I would bet money the vast majority of people using containers, outside of development environments at least, are not running their containers in bare metal but in an existing VM such as on AWS or GCP where this solution is simply a non-starter.

Neat, niche, and doesn't operate in the same world as containers.

TrueDuality··on Falcon 2
It unfortunately isn't worded that it only affects new usage. If you needed to check once before your initial use that would be shady but unquestionably legal as the terms of the contract are clear when you are entering into the agreement.

This clause allows them to arbitrarily change the contract with you at will, with no notice. That _shouldn't_ be enforceable but AFAIK that kind of contract has never been tested. It is _likely_ unenforceable though.

TrueDuality··on GPT-4o
Weird visiting the page crashed my graphics driver using Firefox.
TrueDuality··on Telegram has launched a pretty intense campaign to malign Signal as insecure
Google Play, for all its failings, IS the safest way to get an APK right now.

F-Droid for as much as I love the open platform, does not provide any security guarantees about what you're downloading. It is a volunteer run project and does not have the extensive security policies and practices that Google has. From https://f-droid.org/en/about/

> Although every effort is made to ensure that everything in the repository is safe to install, you use it AT YOUR OWN RISK.

Likewise downloading and side-loading it from their website, requires you to disable some security guarantees by doing things like enabling developer mode.

TrueDuality··on Popover API
Very much a anti-pattern. The entire concept of these is malicious compliance against the GDPR and not something that the GDPR requires.
TrueDuality··on Popover API
I suspect this isn't going to get major usage. Making it a dedicated API makes these things easier to target by extensions and thus easier to block. There are legitimate usages but almost all of the ones I encounter are marketing call to actions and invasive support chat boxes which are both almost universally hated and would the target of those blocking action.
TrueDuality··on Popover API
Which people still hate because its mostly marketing trash that gets in the way of whatever they were trying to do on the underlying page.
TrueDuality··on Ask HN: The Problem with "AI Startups"?
Domain expertise and narrow models can definitely be an advantage but you need to be somewhat successful before you can get sufficient data to train or refine a model with sufficient domain expertise to be a differentiator.

You definitely do not need to train your own base model to be successful, but if your entire pipeline is a system prompt or three in a small agent graph... You didn't build a product, you built a hobby tool over the weekend no matter how much UX/UI polish you put on it.

I do think this question, and many start-ups are thinking about how they want that sweet AI money and are starting their business from there rather than from a problem. If you see a problem that is labor intensive and could be done by mechanical turk... Well you're probably on to something and an AI language model can probably solve that problem.

The companies and AI products I think that are going to last either aren't starting with AI they're focusing on a problem and have reasoned their way to AI OR they're doing some deep stealth research into the problem over a long period of time to be able to develop domain-specific data, techniques, and plans so they can fine-tune their own model before ever showing it to a potential customer. You better be sure you have a good plan if you're going to try and be the latter.

TrueDuality··on Fast, simple, hard real time allocator for Rust
Even with interrupts there is almost always a single execution unit, and thus single-threaded. An interrupt is just a jump instruction to a predefined handler. The code that was running is entirely stopped until the system decides to resume from that point.
TrueDuality··on FridgeLock: Preventing Data Theft on Suspended Linux with Memory Encryption (2020)
Intel TME and AMD SME (both on boot discardable unique memory encryption technologies running in silicon) are both pretty common in consumer grade hardware and has great Linux kernel support.

Both Android and iPhone's use their secure enclave's for storing their encryption keys limiting the effective targets of these attacks (and would be quite difficult to physically extract from).

I suppose this is still useful for older hardware and ultra budget phones... But this is a protection against state actors and high end espionage which wouldn't use those classes of devices...

Soooooo who is this for? What threat model is this meaningful for? In what world am I trusting a random unaudited security module that taints my kernel for _any_ security sensitive application?

TrueDuality··on Memary: Open-Source Longterm Memory for Autonomous Agents
"Semantic Knowledge Graph" is even worse! The term is intended for the design of semantic networks with edges restricted to a limited set of relations. A knowledge graph is already about semantics!

Gotta say txtai seems like a useful tool to throw in my toolbox!

TrueDuality··on Memary: Open-Source Longterm Memory for Autonomous Agents
@dbish nailed it, but I can give you a bit more concrete example. Continuing off the light example I started off with. An ontology for knowing what city a person currently is in. We have two classes of entities, a person, and a city. There is a single relationship type "LocatedAt" you can add and remove edges to indicate where a person is and you can construct some verification rules such as "a person can only be in one city at time".

To have an LLM construct a knowledge graph of where someone is (and I know this example is incredibly privacy invasive but its a simple concrete example not representative). Imagine giving an LLM access to all of your text messages. You can imagine giving it a prompt along the lines of "identify who is being discussed in this series of messages, if someone indicates where they are physically located report that as well" (you'd want to try harder than that, keeping it simple).

You could get an output that says something like `{"John Adams": "Philadelphia, PA, US"}`. If either the left or right side are missing create them. Then remove any LocatedAt edges for the left side and add one between these two entities. You have a simple knowledge graph.

Seems easy enough, but try to ask slightly harder questions... When did we know John Adams was in Philadelpha? Have they been there before? Where were they before Philadelphia? The ontology I just developed isn't capable of representing that data. You can of course solve these problems and there are common ontological patterns for representing it.

The point is, you kind of need to know the kind of questions you want to ask about your data when you're building your ontology and you're always going to miss something. Usually you find out the unknown questions you want to ask of the data only after you've already built your system and started asking it questions. It's the follow-ups that kill you.

There has been a lot of work on totally unstructured ontologies as well, but you're moving the hard problem elsewhere not solving it. Instead of having high quality data you can't answer every question with, you have arbitrary associations that may mean the same thing and thus any query you make is likely _missing_ relevant data and thus inaccurate.

Huge headache to go down, but honestly I think it is a worthwhile one. Previously if you changed your ontology to answer a new question, a human would have to go through and manually and painstakingly update your data to the new system. This is boring, tedious, easy-to-get-wrong-due-to-inattention kind of work. It's not complex, its not hard, its very easy to double check but it does require an understanding of language. LLMs are VERY capable of doing this kind of work, and likely more accurately.

TrueDuality··on Memary: Open-Source Longterm Memory for Autonomous Agents
Yeah precisely. Knowledge graphs are simple to think about but as soon as you look into them you realize all the complexity is in the creation of a meaningful ontology and loading data into that ontology. I actually think LLMs can be massively useful for building up the ontology but probably not in the creation of the ontology itself (far too ambiguous and large/conceptual task for them right now).
TrueDuality··on Memary: Open-Source Longterm Memory for Autonomous Agents
This seems like its overloading the term knowledge graph from its origins. Rather than having information and facts encoded into the graph, this appears to be a sort of similarity search over complete responses. It's blog style "related content" links to documents rather than encoded facts.

Searching through their sources, it looks like the problem came from Neo4j's blog post misclassifying "knowledge augmentation" from a Microsoft research paper with "knowledge graph" (because of course they had to add "graph" to the title).

This approach is fine, and probably useful but its not a knowledge graph in the sense that its structure isn't encoding anything about why or how different entities are actually related. A concrete example in a knowledge graph you might have an entity "Joe" and a separate entity "Paris". Joe is currently located in Paris so would have a typed edge between the two entities of something like "LocatedAt".

I didn't dive into the code but what I inferred from the description and referenced literature, it is instead storing complete responses as "entities" and simply doing RAG style similarity searches to other nodes. It's a graph structured search index for sure but not a knowledge graph by the standard definitions.

TrueDuality··on Show HN: Podlite - a lightweight markup language for organizing knowledge
Ouch. IMHO the reason markdown became so popular is because how simple it is, how well the syntax generally stays out of the way of the content, and its almost entire lack of mixing formatting with document structure.

This is VERY heavy, and VERY invasive. The entire lack of opinion in "use whatever markup you like as long as its wrapped in our markup" seems like an extra awful complication. I'd be curious to hear about what concrete problems the authors of this spec were trying to solve and why they went so far.

TrueDuality··on Snowflake Arctic Instruct (128x3B MoE), largest open source model
Most of those are fine tuned variants of open base models and shouldn't be included in the "every tech company" thing you're trying to communicate. Most of those are researcher or engineers learning how to work with these models, or are training them on specific data sets to improve their effectiveness in a particular task.

These fine tunes are not a huge amount of compute, most of them are doing these trainings on a single personal machine over a day or so of effort, NOT the six+ months across a massive cluster it takes to make a good base model.

That isn't wasted effort either. We need to know how to use these tools effectively, they're not going away. It's a very reductionist and inaccurate view of the world you're peddling in that comment.

TrueDuality··on Meta Horizon OS
You're not going to get HDMI (or at least not directly), but you can now get DisplayPort with many of the XR headsets. They're primarily using the alt-mode of USB-C as the cable is re-used for power as well.
TrueDuality··on What makes a great software engineer [pdf]
This reads like the result of a team bonding exercise and doesn't contain much material content.
TrueDuality··on BeeTrove – OpenAI GPTs Open-Source Dataset
The data is directly linked to in the repository if you'd like it:

Data: https://drive.google.com/drive/folders/1hUGnQ_AWeL2wi5UhUTt0...

Repo Source: https://github.com/beetrove/openai-gpts-data/tree/main/OpenA...

TrueDuality··on Identifying Stable Diffusion XL 1.0 images from VAE artifacts (2023)
Very interesting! Well broken down and explained!
TrueDuality··on Top Israeli spy chief exposes his true identity in online security lapse
It's always a small detail that gets you... You only have to fail once for your OPSEC to crumble into nothing.
TrueDuality··on Big Tech's underground race to buy AI training data
LinkedIn built a whole platform inside their platform for doing exactly this. I think you get a badge or something on your profile claiming your an expert on something if you write a couple paragraphs on a topic using the provided prompt.

They're very clear its going into an AI generated article on the topic but you better believe that is also now core training data.

TrueDuality··on Top Israeli spy chief exposes his true identity in online security lapse
From that line and a few others in the article Amazon doesn't seem to be involved at all. It seems like there was a Gmail account created to receive feedback, corrections, reviews, and media outreach that was also supposed to be relatively anonymous but was registered with his actual name which would be visible to anyone that he responded to at the very least, but may also show up through other account associations.
TrueDuality··on White House directs NASA to create time standard for the moon
> ... that accounts for its differing gravitational forces ...

I salute maintainers of timezone libraries. This is going to be a fun one. You are now responsible for space and time.

TrueDuality··on First practical SHA-256 collision for 31 steps. fse2024
Defacing git repositories doesn't even make sense. You won't mess with people's checkouts, its trivial to detect and identify the responsible party. It's the security equivalent of a child throwing a tantrum in their own room. You want to replace it? Everyone that comes after you and tries to push a change will immediately notice like opening the door to the proverbial child's room. You're busted and you've accomplished nothing.

You want to inject malicious code yourself? When it gets caught, or the file is inspected or reviewed you're busted.

This attack has the opportunity to get malicious code injected into a repository that will never show up in a PR, code review, or any existing checkout (so the senior developers that would notice the change most likely will never receive it). This is re-using an existing trusted and known good commit in your history, even the signature on it, to say "yeah this has always been here, this is perfectly safe and hasn't been modified since the author wrote it".

This is far more subtle, sneaky, and extremely valuable as an attack vector (and it gets more juicy, stay tuned) to get targeted vulnerabilities and backdoors into specific software. This isn't a novel attack method, as I mentioned the Linux kernel goes through a very rigorous process just to avoid this kind of attack.

Aside: You keep trying to use gibberish for your bad examples. You don't need to use garbage, that's the point I keep trying to hammer home to you. The added details can be from any generator and isn't constrained to living exclusively in comments. Garbage is what people use as examples for these attacks because its the easiest, and if you can demonstrate it for garbage then it works for any generator. With garbage you've made the point.

Back at the security issue. So now you have a poisoned repo that contains malicious code and is effectively undetectable through normal use. Meanwhile your production artifacts include the unaltered malicious code from the repository. It will remain unchanged and referenced until someone else creates _any_ change to the file you targeted (as once again git doesn't actually store diffs but whole files in a particular commit). That change might be something like a developer adding some print statements to try and diagnose why the CI system is failing.

When another change happens for that file the evidence mostly vanishes or at least is extremely obscured. There will be _some_ object in your repository that has the SHA-1 object, whether its the original or the malicious one depends entirely on when your checkout occurred.

On the receiving end your best case scenario is that the changed code doesn't work and causes weird bugs in your CI system that can't be reproduced in local checkouts and goes away magically as soon as anyone tries to diagnose it. This capability is worth STUPID amounts of money and I would be shocked if this isn't a technique used selectively in the wild by nation states.

SO how do you solve this problem?

* One of the inherent problems is that signatures don't actually cover the content of the commit. This is another regular complaint of git's behavior and would allow you to side-step this issue using the existing signing infrastructure. This is a bandaid but it's what most people argue for as it is significantly less of a lift than changing the hash function. If you're worried about the attack you just have to sign your commits and tags. If you sign your commits NOW without a change to git, you're still 100% vulnerable to this attack. and because the signatures will still be valid is likely to either make someone innocent look guilty of injecting a vulnerability, or will have audits look less closely at the code because it came from a trusted source causing more harm than good.

* Change out SHA-1 to something that isn't as vulnerable to collision attacks. The problem is collision attacks. Let me say that again, the core issue is with collision attacks. If you can create a chosen plaintext or chosen prefix attack the security guarantees of the git ledger goes away. You can't trust it. It needs to be replaced.

* If neither of those are options for you, your third option of protection is to adopt the Linux kernel policies. Releases are done directly from engineer's machines from a trusted known good repo that has patches added by hand by the most senior engineer.

← PreviousPage 4 of 17Next →