HNHacker News
TopNewBestAskShowJobs

TrueDuality

2,337 karma · joined July 21, 2016

submissionscomments
TrueDuality··on First practical SHA-256 collision for 31 steps. fse2024
I don't need to prove that I can do a thing to prove that a thing is possible and the burden of proof is on you claiming that this isn't an active security problem because that's basically well known and well understood. The only outstanding questions is how-detectable, impactful, and available those attacks are.

Specifically the things you need to counter is at least one of the thing in the following list:

* Hash security: SHA1 collisions are feasible to generate and companies are actively moving away from them with good reason and have been doing so for at least seven years (https://security.googleblog.com/2017/02/announcing-first-sha..., https://www.howtogeek.com/238705/what-is-sha-1-and-why-will-...)

* Content generation: As I've already discussed, the contents of what you use to make that collision can be anything you want and meet any requirements you have the ability to produce a generator for. To meet this you're going to have to prove to me that no engineer can make a seeded random number that uses a language's grammar to produce plausible and valid to compile token, or to just use a language model to produce plausible code and comments (also requiring a seed). This is a _trivial_ thing to do.

* The attack: Git relies on a chain-of-hashes based on SHA1, those hashes are over the complete files included in the repository if you can generate a collision for a file in git's history you can replace the files in that commit and all subsequent commits will remain valid. This is the attack everyone is worried about related to git. The only thing that protects against this right now is the security of SHA1. Additionally signatures on commits and tags DO NOT protect against this, they're over the hash, commit message, and list of objects not the objects themselves. The attacked files will still look like they came from a valid signed commit.

The extra scary part of that attack is the malicious/changed file will not be visible to any existing checkouts, those clients will believe they have the correct object and will continue to show that correct object. But anything that does regular fresh checkouts, like say a CI system that deploys to prod, will get the poisoned object. Even if its checking the signatures on every commit, it won't see this coming.

So the security of all our git repos, our production environments, new devs are foundationally rooted in the security of either write access to the repository OR the foundational security of SHA1.

I would say that is a practical and useful attack. A faster hashing algorithm will EXACERBATE this problem as you're almost always trading collision resistance for speed. Any hashing algorithm that allows you to calculate its hashes faster is MORE vulnerable to collision attacks not less.

"Computationally astronomical" isn't a very good argument. 20 years ago SHA1 was insane in its security. These thing get weaker over time and need to be periodically replaced, not because they're failing, but because increased resource capacity has fundamentally changed the original assumptions the algorithm was designed for.

Even with the computationally astronomical argument that is a matter of cost and resources, not practicality. It absolutely is practical to do if the result is worth the outcome. What is the most famous git based project? Maybe the thing it was originally designed to manage... Think maybe _any_ nation state would be happy to pay less than ~$100k USD (https://sha-mbles.github.io/) to get some malicious code running in production builds of the Linux kernel? The kernel project specifically has extra manual checks and multiple "known good repos" with commits literally being added by hand to protect against this attack. It's practical, it's a problem. It needs to be fixed.

If you still insist on a working example pay me $125k and I'll produce one for you.

TrueDuality··on First practical SHA-256 collision for 31 steps. fse2024
Even without comments your additional requirements aren't relevant, but not in the way I think you're assuming.

When you're searching for a practical collision you only need a way to generate systematic output that semantically will be interpreted with your intent. The easiest way to do this is to include semantically irrelevant data to something that was manually produced that is semantically relevant.

In the programming domain, source code specifically, comments are the easiest way to include semantically irrelevant information but you could also include unused functions, variable names etc. You are literally limited by the constraints of your imagination and your ability to dodge CI failure checks.

Aha! You might say, but any human that saw that change or PR would immediately notice the garbage produced and catch the collision attempt! (this is your argument) Unfortunately no, that assumes your search space that I talked about is over semantic garbage. It's a bit more work, but your search space for a collision could be "Shakespearean sonnet's that would make a literary buff cry" as long as you had a generator that could produce it and produced different outputs from different seeds.

We now have access to a generator that can take an incrementing seed number, and produce both semantically meaningful and meaningful semantically irrelevant content. The language models. Interestingly this moves the compute cost to the generator (usually the compute restriction is on the hash being attacked).

It's definitely not practical with our current compute capabilities to attack a search space of 2^256 through brute force for a simple hash much less including waiting for a language model to produce an output using a different input seed for each check but that's not what this article is about either...

What these collision attacks (such as the linked article) do is _decrease the search space_. Without any algorithmic tricks the search space for sha2-256 is 2^256. These tricks are eating away at that exponent. This work results in a reduction of a collision to 2^49.8. That is a massive drop in the search space. Is it still feasible to attack today? Absolutely not. But a few more of these tricks and I can see those "garbage comments" collision happening, but wait a tiny little fraction of a time beyond that and include language models for your search space?

Hell your changes could be _productive_ and produced incrementally through a series of commits if you really wanted to limit your search space and get creative about it.

With SHA-1 collisions attacks using semantic garbage are already considered practical. We're still probably computationally constrained in using language models to produce semantically viable collisions but we're not that far off either. Those comments won't be garbage. You will not be able to distinguish it from any other AI generated code being committed which is rapidly improving in quality and efficiency to generate.

Even without language models you could use something like a language's EBNF grammar as a token generator for source code which would probably pass any glance checks, but definitely not dedicated inspection like a code review. That is probably something that IS PRACTICAL TODAY for SHA1.

TrueDuality··on Algorithms for Modern Hardware
Compiler optimization is an extremely hard problem. Unsolvable by our current standards and usually the compiler has less knowledge than the author of the code and has to make assumptions about the hardware it is going to run on. Language design feeds heavily into the types of optimizations that are even possible to reason about.

So yes there has been an enormous amount of research into both the programming language design side of this equation, the compiler design, and the area in the middle. Probably hundreds of thousands of research papers since the 50s or 60s.

Compilers and automated systems, in the general case at least, still can't out compete a domain expert that is familiar with the specific hardware being targeted, a good profiler, and the time to fiddle with the raw assembly.

TrueDuality··on :syntax off (2016)
I used more or less a bare vim (not minimal, not neovim) config with no plugins and < 100 lines in my vimrc for about 5 years. Minimal syntax highlighting, zero popups or autocomplete, no jump to definition, etc.

It wasn't intentionally by choice, I remember getting frustrated with some interaction between a couple of plugins and resolved to start from a clean slate and carefully configure and curate my config. The clean slate happened but the subsequent configuration never did. I got used to it pretty quickly and it forced me to keep more of the code base in my mental RAM at any point in time.

That still resulted in quite a bit of hunting around, slower refactors but I was thinking about the code quite a bit more. I finally started using a much more feature rich environment and while I'm glad I did the long stretch without it, I even feel I benefited from it, I don't think I could go back.

For autocompletes specifically I do find the immediate suggestions annoying even when I can type right through them. I have no idea how people work when they have to pause their coding to actively dismiss suggestions either via the mouse or through a keyboard shortcut.

For the time being my autocomplete suggestion window is a key combination to even display for a single instance and no additional actions need to be taken unless I want to accept the suggestion (I can still type right through it). I'm sure everyone's mileage will vary but autocomplete feels significantly more intrusive to me than syntax highlighting.

TrueDuality··on Brits left baffled by Brexit's 'not for EU' food labels
Probably not if I had to guess, but trade deals between the EU and Britain still need to be worked out, and that trade needs to be facilitated by customs enforcement which there hasn't been infrastructure for before between the EU and Britain.

Complicating the issue is Northern Ireland which effectively falls under the jurisdiction of both. If British products need to be sold there, they need to be distinguishable somehow from the versions that are allowed to leave Northern Ireland into the rest of the EU.

That combination of shared jurisdiction and lack of trade infrastructure is how you get these labels. Its the laziest solution to a self-imposed problem.

TrueDuality··on Monte-Carlo graph search from first principles
I would argue that this saves both compute and memory for any significant search. In a traditional tree-search without treating states as identical and revistable still requires you to store every node that you visit in memory to track the N, Q, and child relations.

Applying the algorithm adds a hash associated with each node, and a counter on each edge. On the other side you're de-duplicating identical states in your search space, saving the memory of duplicate nodes but more importantly saving compute by not requiring independent searches of the same state space. Whether this trade off actually saves memory will be determined based on how likely duplicate states are in your search space.

Otherwise your summary is absolutely spot on.

TrueDuality··on The Pile is a 825 GiB diverse, open-source language modelling data set (2020)
Most of the distribution for this is via torrents/magnet links and in person hard drive exchanges. I'd go look at some public trackers if you want a copy and don't know someone that already has it.

Do be aware that it does include copyrighted content so distribution is piracy.

TrueDuality··on OpenIPC is an alternative open firmware for your IP camera
They have a list of specific chips the firmware supports which is very useful but I couldn't find any concrete camera models this was for or the features the firmware supports.

Maybe I missed it but it'd be nice to raise the visibility of that a bit more.

TrueDuality··on LongRoPE: Extending LLM Context Window Beyond 2M Tokens
Legitimately curious, what is lacking from current playthrough and movie descriptions available now? For games the only things I feel like might not be represented well in text is descriptions of the observer/player reactions and maybe details about their specific choices through a game... Kind of like turning a choose your own adventure book into a normal book by tracing one path through it?

I can see Q&A on a movie being useful but can't think of descriptions of the overall movie itself lacking... I'd definitely be worried about spoilers.

TrueDuality··on StreamingLLM: tiny tweak to KV LRU improves long conversations
There was a really interesting post a while ago about adjusting the softmax function to allow attention heads to not make a choice (https://www.evanmiller.org/attention-is-off-by-one.html). It seems like that might remove the need for these attention sinks entirely. I keep meaning to go in and perform tests on this but boy time gets away from you...
TrueDuality··on Ubisoft Says Out Loud: We Want People to Get Used to Not Owning What They Bought
Your analogy falls down because the act of taking a physical good deprives the current owner of that good without recompense. That is the definition of stealing.

Software and digital content piracy does not deprive any other owner of anything else, this is isn't stealing in the legal sense. It may be financial harm or copyright abuse but those are both separate crimes from theft.

Piracy isn't legal, but it isn't theft either.

TrueDuality··on Changes we're making to Google Assistant
I was really hoping something had forced Google to allow the full disabling of the assistant. I really don't find it useful and accidentally trigger it more often than I care to admit.

Make it truly useful or get out of my way.

TrueDuality··on Airbus Shatters Record for Jet Orders as Demand Soars
Probably less. A lot of US military groups have dedicated personal for maintaining operational hardware including planes that are cross-trained at these large companies.
TrueDuality··on Ask HN: How much code do you write in your job?
My team is pretty small right now so yes I do end up doing a fair amount of technical decisions but at my last spot, I was more of a head engineer over groups of engineers and I really need to learn to let the teams below me make their own technical decisions only stepping in when there was a conflict, they were explicitly asking for guidance, or in a few rare cases where I could see they were clearly going down a wrong a road.

I think I would need a bit more information about your current seniority and skill levels before I could give you more specific direction. There is always the default answer if you want to move up higher that communication skills are important, and that's true but what I don't see as often is developing the skill to let decisions go and to not only find ways to live with them but actually accept those decisions and move forward with them.

Absolutely speak up if you disagree, and communicate clearly the prior experiences that is helping guide your decision. Talk about your assumptions and where you think future pitfalls might be. If your team or manager decides to go down a road you don't agree with or don't think is optimal... Let it go. Work with it. Try and accept that as a path forward and embrace it. You might not have the full picture, it might be a legitimately better strategy, or you might be right. In any of those scenarios though fighting against the current, trying to go rogue to prove your path is correct. Those will only lead to ill-will. If it does turn out to be the wrong path, don't play the "I told you so card", take the high road and just move forward.

A lot of engineers, myself included, look back at prior decisions more than looking forward. Reflection is important, and lessons can be learned but we're there to build things, not be right all the time. The bigger your team the bigger and more intense things you can build, but only if you're all working with the current of development.

The more you work with your team, the more they trust you and as you gain in seniority you will have a peak of silent "I told you so moments", that's usually when I've found it was time for me to go up to the next level of seniority as its probably time to start mentoring the people around you.

None of this is about getting more coding time, but I don't think I've ever really optimized for that. Maybe think about what you're trying to get out of that coding time. Maybe what you're actually looking for is unbroken focus time, that I have optimized for but also involves working with your team. Communicate with your team that you would like more focus time, propose things like protected afternoons (meetings only in the morning) or meeting free days (I've found Tues/Thurs usually work best for this).

If you provide me with a bit more detail I might be able to give less general advice.

TrueDuality··on 802.11ah Wi-Fi HaLOW: The 1 Kilometer WiFi Standard
But dropping from 2.4Ghz and 5Ghz down to 900Mhz does significantly improve the penetration from varying forms of obstructions and makes tropospheric ducting more likely to occur.
TrueDuality··on Ask HN: How much code do you write in your job?
I'm doing the CTO thing for a <10 person start-up that is actively building our initial product. I am rapidly and randomly switching from anywhere from 0-80% coding during any given week depending on what my team and the company needs. Probably closer to 20-30% on average.

I come from a history of being an individual contributor (15 years of that) and I would say overall as I've gotten more senior in my positions I've done less coding as a % of my time, but the coding that I have been doing is individually more impactful or sensitive.

The non-coding part of my time has gone more into mentoring, reviewing, design, business discussions & goal setting, policy writing and implementation, team organization & management, and interactions with operations.

I think that about covers it, but feel free to follow up if you have anymore questions.

TrueDuality··on "Anna's Archive" blocked following publishers' complaint
Frequently these ISP level blocks also redirect the normal port 53 traffic to their own servers so these settings are meaningless. What does work is DNS-over-HTTPS which can be enabled quickly and easily in both Chrome and Firefox and will side-step this block even against adversarial ISPs.
TrueDuality··on Gentoo goes Binary
I agree on customizing the package flags, and features. When using Gentoo in production it became an important part of our security posture to omit the features and integrations with unused software.

That being said we've always had a build host dedicated to producing binaries, but the actual support for binaries in Gentoo hasn't been great. Unsigned serving over HTTP or NFS of compiled artifacts is about all you get. I'm really pumped to see that the new package format adds in cryptographic verification that really should have been there all along even for internal only serving.

TrueDuality··on Ferret: A Multimodal Large Language Model
I'm just inferring myself, but I believe it's referring to discussing things in the foreground / background or in a specific location in the provided image (such as top right, behind the tree, etc) in user queries.
TrueDuality··on Ferret: A Multimodal Large Language Model
Modalities and models are not the same thing.
TrueDuality··on AI companies required to disclose copyrighted training data under new bill
This is interesting with the recent work showing that models using different architectures, but the same dataset converge to similar capabilities. It makes the case that the data set itself and the compute over it is the secret sauce for these models.

It'll be especially interesting with the ruling indicating that the output of these models isn't copyrightable as that and internally generated (from scratch) data will be the only bits omitted from this reporting. I'll be curious if the gap gets bridged on including the disclosure of data generated by a model containing copyrighted data.

TrueDuality··on The changing face of post-pandemic New York City
If you click into the article, they have it broken down by boroughs.
TrueDuality··on Show HN: Beeper Mini – iMessage client for Android
The biggest question I have is what does the exit story look like? I know it's a common problem for people switching to Android from an iPhone (keeping their phone number) and not receiving any messages from iPhones as their contacts are still trying to send via iMessage.

What happens at the end of the trial when I choose it's not for me. Will I suddenly loose my ability to text with every Apple user? What about in the event of a service failure of your infrastructure (it says it talks directly to Apple's servers but does it also rely on you?)... What if Apple does go after you or makes a breaking change to the protocol?

In all those scenarios, what happens to my messages? Can I export and keep them?

I'm very interested to try but the potential to loose access even for a day or two is a really hard ask.

TrueDuality··on Loongson releases next-generation CPU
That sounds like it would be an absolute blast.
TrueDuality··on IKEA sensors for doors and windows, motion, water leaks
Doesn't Matter run over Thread? I thought it was just an application protocol not a physical connectivity protocol.
TrueDuality··on The IDP Knowledge Base System
Sorry missed this comment. There are a surprising number of them, basically any formal solver is a deductive reasoning engine. Prolog is itself a deductive reasoning engine (and if you're going to implement a reasoning system, its probably worth going through the WAM Book over a weekend, search Google for wambook.pdf, its a publicly released book with the title "Warren's Abstract Machine: A Tutorial Reconstruction". The harder ones are inductive and abductive reasoners, which I'm not aware of any opensource ones only old papers that all do little tiny bits. A lot of deductive reasoners boil down to just depth-first-search with backtracking.

How I'm storing them is maybe interesting... Graph databases that support tagged or typed relations get you most of the way to a basic reasoning system. I'm currently storing the graph content of the knowledge base in a database called SurrealDB with the RocksDB backend (this is new to me). Learning SurrealDB has been giving me a bit of Deja Vu around building simple reasoners, there is a lot of overlap between graph theory and formal reasoning.

Rules are encoded in the program. I find it beneficial to think about this as building up lots of small tools that are aware of what they're capable of and can indicate whether they think they can make single step progress towards the goal. You might have a small tool that if you have a location and you're looking for a different location (its prerequisites for "knowing" it may be useful) you could have that tool walk your graph finding the smallest shared domain that both are a part of that can become part of the context. Building up context and learning to disambiguate things using that context is kind of the "hard part" if you want to optimize your solvers or ask complex questions.

For the Rules to be meaningful you need to have an Ontology for your database. Your Rules will be operating on top of fixed concepts, and known relation types between them. Ontologies are a BIG subject that has also faded into the mist a bit, there are some samples on http://ontologydesignpatterns.org but there are a lot of dead links on there as well. The more complex your Ontology the harder to write and/or less useful the individual tools become as they will probably have fewer opportunities to run.

Hope this helps!

TrueDuality··on Linux being secure is a common misconception
iOS being secure is also a common misconception. A bit of an extreme example but that is a meaningful problem as opposed to Linux on the Desktop cases.

The problems pointed out in this post are almost all around usability as Linux on Desktop. The author admits the tools exist but are hard to use in most of the cases. Where features are missing its a misunderstanding of the Linux world.

A lot of these protections either live under different names outside of the Windows world and the ones that aren't don't exist because Linux protects things in a different less vendor-lock-in way. This is very apparent in the virtual machine section which are about protecting host kernel primitives from the VM... Linux doesn't expose _any_ host kernel primitives to the VM. The closest is narrow drivers, properly sandboxed and isolated in userspace, that are minimal and have their own security guarantees.

eBPF is far from a dangerous feature, C is only a dangerous language when that danger isn't managed, and the Linux kernel is top-of-class for managing those features, there is the same root boundary issue in Windows and is a deep source of security problems (root on Linux CAN be restricted through both seccomp and SELinux policies unlike SYSTEM).

Things could be better for sure in the Linux world, but pretty much everything besides secure defaults here requires a level of effort or access to attack that requires full compromise of the machine already. There is much lower hanging fruit that we need to clean up and the funny thing about Linux is that it trends hard toward security and quality over time.

You just have to look at time to fix patches for security vulnerabilities, not just in the kernel but in any packages maintained by security distributions. The author calls out not getting patches back-ported, without looking into the patches that don't get back-ported. RHEL won't backport fixes for features that aren't compiled in for example.

There was a post earlier this week about the CVSS scores being different between NVD and RHEL's bug trackers... It was because the networking functionality of that package wasn't compiled in so there was no possibility of remote execution .

TrueDuality··on The IDP Knowledge Base System
While I agree with your statement I would add "yet" to the end of it. Language models are powerful tools on their own, but we're already augmenting their knowledge with external data stores. Personally I think the next stage is going to blend in knowledge graphs.

Populating these Ontologies is very manual and time-consuming right now. LLMs without any additional training (I'm currently using a mix of the various size Lllama2 models, along with GPT 3.5 and 4 for this) are capable of few-shot generation of ontological classification. Extending this classification using fine-tuning is doing REALLY well.

I'm also seeing a lot of value in using LLMs to query and interpret proofs from deductive reasoners against these knowledge graphs. I have been limiting the scope of my research around this to two domains that have a kind of eccentric mix of formal practices, explicit "correct" knowledge, and common sense rules of thumb to be successful. Queries can be quite onerous to build which a fine-tuned model can help with, and LLMs can both assist in interpreting those logic chains and doing knowledge maintenance to add in the missing common sense rules or remove bad or outdated rules. Even selecting among possible solutions produced by the reasoners is really solid when you include the task, desires, and constraints of what you're trying to accomplish in the prompt performing the selection.

The formal process and knowledge is handled very well by knowledge graphs along with a deductive reasoning engine, but they produce very long winded logic chains to reach a positive or negative conclusion where a simpler chain might have sufficed (usually due to missing rules, or a lack of common sense rules) and are generally incapable of "leaps" in deduction.

LLMs on their own are capable of (currently largely low-level) common sense reasoning, and some formal reasoning but are still very prone to hallucinations. A 20% failure rate when building rules that human lives may depend on is a non-starter. This will improve but I don't think a probabilistic approach will ever fully remove them. We can use all of our tools together in various blends to augment and verify knowledge, fully automatically, to make more capable systems.

TrueDuality··on The IDP Knowledge Base System
I really want to see an open knowledge base project tackle exactly what you're describing. Cyc has a pretty solid foundation for accomplishing this using what they call microtheories (personally a weird term to me). It's a great step in the right direction but its a closed ecosystem with a huge amount of cognitive overhead to do anything against. A lot of other projects seem to do better with the ingestion and updating of information (albeit with less automated tooling).

I've been refreshing myself on ontology design and knowledge representation lately. I'm kind of surprised that is hasn't moved very much in the last two decades. I'm guessing we're about to hit an inflection point there as it becomes a cool topic with the current AI trends.

TrueDuality··on RustGPT: ChatGPT UI Built with Rust, Htmx, SQLite
This comes up a lot but I never get tired of answering it. There might not be anything about Rust that is interesting to you personally but generally the benefits are different depending on what language you're coming from:

If you're coming from C/C++/C# you get very good memory guarantees, a standard library that is very well fleshed out, a well supported package manager and a fairly robust (though a bit immature still) ecosystem.

If you're coming from Java or Golang, you have the benefit that now your code can be consumed by other languages. You also get the memory guarantees, better error handling ergonomics, and in the case of Java the ecosystem is much higher quality and safer (at least in my extensive experience).

From Java specifically (and any of its derived languages), you loose all the operational and memory problems that come with just starting the JVM.

If you're coming from Python, Ruby, JavaScript the biggest benefit is probably performance but you get all the benefits that Java and Golang do as these are also dead end languages.

There is a lot of quality of life, safety, security, and performance benefits but what matters the most to YOU is going to largely depend on where you're coming from, what you want to build in it, and what your priorities are. It legitimately isn't for everyone, but its a VERY good choice as a default language.

← PreviousPage 5 of 17Next →