HNHacker News
TopNewBestAskShowJobs

segfaultbuserr

13,013 karma · joined December 17, 2018

submissionscomments
segfaultbuserr··on Data-only attacks are easier than you think (2024)
correction: s/static analysis/data-flow analysis/
segfaultbuserr··on Data-only attacks are easier than you think (2024)
Computing is a huge subject, and there are many branches. For context, one branch, which I call "memory corruption studies", ever since the 1990s, is almost an entire research discipline with its own standing in infosec. As a result, entire OS concepts were invented to mitigate them (there even exists entire operating systems nearly dedicated to this this, such as OpenBSD and HardenedBSD), CPU hardware was literally modified to help mitigating them, new complier code generators were invented to mitigate them. Thousands of PhD degrees were awarded based studies on them. The research arm of every major tech firm has teams dedicated to them. There are infosec companies dedicated to this single subject.

In this context, practical data-only attacks as reported by this paper are the holy grail in this field.

As an outsider critic, you may say it's a side-effect of C/C++. If you use dynamic programming languages, this issue doesn't exist in your universe because there's no fundamental difference between code and data. But in another parallel universe of systems programming, it's a huge subject. Both the Morris worm and the publication of the article Smashing the Stack for Fun and Profit in Phrack were regarded by many as the milestones of hacker culture and canonical models of hacking, all the exploitation and mitigation studies that followed it were partially motivated by new hackers who wanted to "advance the field of hacking", so any advancement would be considered significant by a hacker. This is Hacker News, and I thought most people would understand this background. But apparently many developers are application-focused nowadays and hack in different universes, and it's not the case.

segfaultbuserr··on Data-only attacks are easier than you think (2024)
It depends on what you overwrite. If you overwrite a CPU instruction or a function pointer that followed the buffer, it's a code-execution attack. If you overwrite a data variable that followed the buffer, it's a data-only attack. I said nearly all conventional fuzzing found code-execution attacks, not data-only attacks. Isn't that clear? The former method is considered common, well-studied, with defenses, the latter method is considered rare, niche, and defenseless.
segfaultbuserr··on Data-only attacks are easier than you think (2024)
In microcontroller programming, redundant data, checksumming and token-passing are sometimes used to mitigate CPU malfunctions due to electromagnetic interference (microcontrollers are often used as "programmable logic", so there's no hard layering between hardware and software, layering violation is made on purpose). If anything looks wrong, you trigger an assertion failure and reset the chip via the watchdog timer. For example, when you pass LS_CMD_STR, you would also pass the name of the caller and the CRC32 checksum of the string as arguments, and the function on the receiving side should validate them.

So I think adding assertion is definitely a way to discover data-only attacks in fuzzing, or even as a partial mitigation of these attacks. It's just stack canary for variables and strings (but as the paper authors said, complete mitigation can be impractical).

segfaultbuserr··on Data-only attacks are easier than you think (2024)
First, there is no "AI", the author showed even a simple static analysis as proposed by them was sufficient to find data-only attacks, in contrary to the common belief that you need heavily customized exploits per application for this kind of attacks. This is the whole point of the research.

Next, I believe they found 944 available "data-only gadgets" usable by a pre-existing memory corruption bug. You still need to find a memory corruption bug first to use them, in the same sense that you need to hunt for ROP gadgets to get arbitrary code execution on a W^X system.

segfaultbuserr··on Data-only attacks are easier than you think (2024)
Corrupting program memory via malicious input data is known as a code-execution attack, not a data-only attack. The fuzzed program usually crashes because its executable code or the control flow got overwritten directly by the input, or indirectly by the program code itself when it tries to process bad data. An exploit involves injecting external code, or overwriting memory addresses (like a virtual table or a stack return address) to override the original logic flow to do something else.

A data-only attack would be an attack that reuses the original logic by only corrupting data inputs (such as a flag or a file path), without overwriting code or overriding the logic. W^X, stack canary, or CFI won't work in these cases since no code is tampered by the attacker. In almost ever talk about compiler mitigations, you always hear a passing-by mention of data-only attacks - before the speaker immediately dismisses them as an academic curiosity when the software industry is still facing a flood of stack smashing and ROP attacks.

segfaultbuserr··on WordPress: Unauthenticated path traversal leading to conditional RCE
WordPress began as a FOSS replacement of Movable Type after Moving Type 3 made the controversial decision to change its licensing fee for some features. So it was like a Unix/Linux situation.
segfaultbuserr··on ReBarUEFI: Resizable BAR for almost any UEFI system
ReBar's commercial name is AMD Smart Access Memory, it allows a PCIe device such as a GPU to map more VRAM to the system at once, which improves performance by reducing access overhead. This generated much fanfare in the early 2020s after AMD officially supported it in the newly released AMD Zen 3 CPUs with RX6800 series GPUs. It was marketed as a new technology to boost GPU/gaming performance. What AMD did was just rebranding an obscure feature in the PCIe specification [1]. As shown by this project, it was actually supported by the PCIe controller since Sandy Bridge, just disabled in the firmware. For a decade nobody bothered to use it. Presumably, AMD saw an opportunity and enabled it, presumably after validating the hardware and fixing any driver compatibility problems.

[1] This is nothing new in the tech industry. Intel rebrands DVFS as SpeedStep, IOMMU as VT-d, AMD rebrands the NX bit as Enhanced Virus Protection, etc.

segfaultbuserr··on SSH3: Faster and rich secure shell using HTTP/3
TCP port knocking.
segfaultbuserr··on Scientists uncover how the brain washes itself during sleep
The brain truly is a system with terrible service availability. On average, after running for just 16 hours, it must be offlined for 8 hours to run maintenance tasks such as "scrub", "garbage collect", "trim", and "fsck".
segfaultbuserr··on The Gambler Who Cracked the Horse-Racing Code (2018)
Is a binary search involved in this "gambling system"?
segfaultbuserr··on Unix Programmer's Manual Third Edition [pdf] (1973)
+1. Gorden Bell said a new computer category would enter the market every decade [1]. PDP-8 and later the PDP-11 were the quintessential minicomputer category makers. They were basically the microcomputer-equivalent in the 1970s. Both brought great cost reduction in their respective eras.

[1] https://en.wikipedia.org/wiki/Bell%27s_law_of_computer_class...

segfaultbuserr··on Bill Atkinson doxxed Douglas Adams in 1987
In 1987, doxxing yourself was the norm. ~80% of the Usenet messages from the 1980s (to early the 1990s) had names, institutions, office addresses, and phone numbers attached to them too. Most were universities, governments, or corporate R&D addresses, but there were many small businesses and home users as well. Some phone numbers are probably still valid today. In fact, an Usenet archive (UTZoo) has already been taken down from the Internet Archive due to an alleged legal threat made by an individual (despite that this archive was indispensable if anyone wants to find any historical information from this era, and that it had been available online for the last ~20 years before it was taken down, with multiple copies still online). I suspect the legal status of these kinds of early online community archives will be increasingly problematic over time.
segfaultbuserr··on AWS data center latencies, visualized
Sorry, it was a typo. I meant 2/3 (including common cables and fiber optics), not 1/3.
segfaultbuserr··on AWS data center latencies, visualized
What an embarrassing typo! I was thinking of 0.66, and somehow I thought 0.66 = 1/3 (must've been distracted by the "2" in 1/2). I should've written 0.66 or 2/3.
segfaultbuserr··on AWS data center latencies, visualized
1/2 c in circuit boards (FR-4), 1/3 c in cables, two useful numbers to remember.
segfaultbuserr··on An intuitive guide to Maxwell's equations (2020)
I remember reading a great answer [1] from Stack Exchange, that claims:

> the 1873 treatise used a pre-Heavisde form of vector calculus cannnibalized from Hamilton's quaternions ... only sparingly, to present the equations in capsule summary form.

Thanks for the reply. From your link, I now understand what does "vector calculus cannnibalized from Hamilton's quaternions ... only sparingly" means.

[1] https://hsm.stackexchange.com/a/15618

segfaultbuserr··on The Discovery of Superconductivity (2010)
> You read some Wikipedia pages and Feynman lectures of physics. I'm a physicist who has done well over a decade of research in magnetic materials.

In the same way that a geodesist navigates using a reference ellipsoid defined by WGS-84, while a city commuter uses Cartesian coordinates on a flat map. The commuter's navigational tool will never work in geophysics research, and it doesn't need to be.

> To the parent and its sibling comments: There is no atomic or subatomic current that can explain ferromagnetism in any approximation. [...] Any such explanation attempt fails spectacularly if you actually try to do the math (which gives an electron surface that is moving faster than speed of light, as Uhlenbeck/Goudsmit who proposed this incorrect idea quickly found out), so it doesn't even work as an approximation of any kind.

I consider "circulating currents create ferromagnetism" to be as true as "an atom's structure is similar to a solar system." Both concepts break down when it's examined in details, so its use by research physicists is obviously unacceptable, but I consider it's nevertheless as an useful mental image in introductory discussions among non-physicists.

Would you consider Rutherford's original atom model to be a first approximation? Can it be considered a very oversimplified but useful heuristic, at least when people who know anything about atoms are first introduced to this concept? Alternatively, would you consider Rutherford's atom to be "an explanation attempt that fails spectacularly if you actually try to do the math (which gives an electron that collapses into the nucleus in picoseconds, as Rutherford's colleagues quickly found out)?

If you believe the latter case, everyone can stop this conversation right now. Because it means the entire disagreement is entirely down to what kinds of "metal images" are acceptable, rather than any factual, like "whether a full quantum treatment of ferromagnetism is necessary to completely explain ferromagnetism (of course it is)." The rest of us who don't solve research problems believe a toy model is still interesting, but don't deny (nor mention) better models. You, as a professional physicist, believe many "what if?" metal models from history are just not legitimate physics, and should not be mentioned at any circumstances to avoid poisoning the minds of youths - an approach known as Whig history, in which scientific progress marches from one victory to another, and all losers be damned - a perfectly valid approach for teaching physics to students who only care about pure physics science, instead of "who said what."

As a side note, I know some engineers who really hate the idea that electric circuits works due to an electron flow. The most extreme one I've seen of wanted to ban this concept in introductory textbooks, calling it a big lie (an explanation attempt that fails spectacularly if you actually try to do the math, which gives the speed of an electron 30 billion times slower than the speed of light in free space). As we all know, the steady-state electron flow was only a result of the transient propagation and reflection of electromagnetic waves in free space or dielectric materials. Thus, they believe the wave model should be the only interpretation in a science textbook, since "they're high-school teachers, I'm a design engineer who work with high-speed digital systems with 20 years of experience, and I know for sure that high-speed circuits and computers can't even be made functional if you ignore fields and transmission line effects." Meanwhile, I believe the electron flow model still works as an introductory mental image (although the field view perhaps needs to be mentioned earlier).

> Who developed this theory in quantum mechanics, where and when? Pauli, who first introduced it into quantum mechanics and the namesake of spin 1/2 matrices, insisted that it is purely quantum mechanical with no classical analogue.

The earlier "electron as a rotating ball" idea was considered by Ralph Kronig and Uhlenbeck-Goudsmit in 1925. Pauli personally never accepted it due to its unphysical flaws. Only in 1927 did Pauli publish a rigorous QM treatment. Thus, "electron spin using classical rotation as analogue" was still an intermediate step before establishing this concept in QM. It was a footnote in history since Pauli was a great physicist and already considered the problem himself earlier and found the solution before everyone else. Otherwise this intermediate step may last longer than 2 years.

> Furthermore, such magnetic moments (called magnetic impurities in that context) ruin the superconducting order by breaking the time-reversal symmetry, so trying to make a connection to ferromagnetism in the context of superconductivity is even worse.

This, in comparison, is a more interesting criticism.

segfaultbuserr··on The Discovery of Superconductivity (2010)
You said,

> Ferromagnetism has nothing to do with currents

This is why I said ferromagnetism is circulating current in the sense of "to a first approximation" and "heuristically". Wiktionary defines "heuristic" to be:

> a practical method [...] not following or derived from any theory, or based on an advisedly oversimplified one.

I think that if you ask Feynman, he would probably agree or sympathize with the naive idea of "atomic currents" as a heuristic argument in the introduction of this topic... which is nothing new anyway, and has been a heuristic argument used in electromagnetism for a long time, at least before QM.

In Feynman's own words,

> These days, however, we know that the magnetization of materials comes from circulating currents within the atoms—either from the spinning electrons or from the motion of the electrons in the atom. It is therefore nicer from a physical point of view to describe things realistically in terms of the atomic currents [...] sometimes called “Ampèrian” currents, because Ampère first suggested that the magnetism of matter came from circulating atomic currents.

You said,

> Spin is a type of intrinsic angular momentum that is not associated with any spatial motion.

Yet the concept of spin in quantum mechanics was originally developed using macroscopic rotations as an analogy, although today we know that spin is an intrinsic property of subatomic particles (thus the joke, "Imagine a ball that is spinning, except it is not a ball and it is not spinning.") In the same sense that Ampère's concept of "atomic currents" was developed using circulating electric current as an analogy.

> The Feynman lecture you linked to is an explanation why currents fails to explain ferromagnetism. You need to read the next chapter.

Of course, "The actual microscopic current density in magnetized matter is, of course, very complicated." This is surely explained in the next chapter. I could've mentioned "atomic currents" without citing any link, but I included it to allow anyone who's interested to read the whole thing in context.

segfaultbuserr··on The Discovery of Superconductivity (2010)
Knowing superconductivity makes magnets less mysterious. Once you accept that physics absolutely allows the creation of a static magnetic field from a circulating current that flows forever in a zero-resistance inductor coil, then the existence of ferromagnetism is no stranger than that - to a first approximation, it also comes from circulating currents, "just" on a subatomic scale. [1] It's kind of surprising that the Atomic Current Hypothesis of ferromagnetism was already proposed by Ampere back then. Following the same heuristics, the fact also becomes clear that the energy in an inductor coil can't really be "spent" to do useful work forever without de-energizing it, and the same is true for permanent magnets. [2]

[1] https://www.feynmanlectures.caltech.edu/II_36.html

[2] This intuition debunks many types of incorrect "infinite energy of magnets" ideas that lead to perpetual motion. Although it can't debunk the "perpetual motion solely from an uneven static (electromagnetic or gravitational) field" idea, which is even older.

segfaultbuserr··on Simple tasks showing reasoning breakdown in state-of-the-art LLMs
> I must confess, when I tried to answer the question I got it wrong...! (I feel silly).

In programming there are two difficult problems - naming things, cache invalidation, and off-by-one error.

segfaultbuserr··on Why do electronic components have such odd values? (2021)
> There’s also part of, good designs don’t depend on high precision components. I think TAoE emphasized that.

If I call correctly, TAoE said engineering calculations should never keep too many significant digits, since no real-world components are that accurate, and all good designs should keep component tolerance in mind - they should not have an unrealistic expectation of precision. It also mentioned that designing a circuit for absolute worst-case tolerance is often a waste of time.

But I don't think TAoE told you to "avoid precision components in your design, use trimmers instead" (Do you have a page number?) when the application calls for it. For example, 0.1% feedback resistors in precision voltage references are often reasonable.

> For high precision one can use trim potentiometers

From what I've read (from other sources), mechanical trimmer used to be extremely popular, but they went out of favor in recent decades because tuning could not be automated and that increased assembly cost. Using a 0.1% resistor is favorable if it allows trim-free production.

> or maybe even digital potentiometer with an ADC at the other side to measure and get as close as possible

Yes, digital trimming and calibrations is today's go-to solution.

segfaultbuserr··on An intuitive guide to Maxwell's equations (2020)
Also, in page 452 [5], Heaviside wrote:

    div B = 0
Finally in page 475 [6]:

    div D = ρ
So yes, essentially all 4 Maxwell's equations were here.

[5] https://archive.org/details/electricalpapers01heavuoft/page/...

[6] https://archive.org/details/electricalpapers01heavuoft/page/...

segfaultbuserr··on An intuitive guide to Maxwell's equations (2020)
> Maxwell was also not really taking advantage of them much and always split them up into scalar and vector part.

Did Maxwell actually use quaternions? If I recall correctly, at least in A Treatise on Electricity and Magnetism, quaternions were not actually used. Instead, he did most things in Cartesian coordinates, and all equations were applied to a vector's x, y, z components tediously. But many sources claimed Maxwell used quaternions, including quotes from Lord Kelvin. My reading on this part of history is limited, so my guess is that he did use them in personal research or in later papers. On the other hand, some other physicists of the same era used quaternions extensively, including applying them to Maxwell's electromagnetism, that is a sure fact...

Coincidentally, A Treatise on Electricity and Magnetism was written as an overview all electromagnetic phenomena as a whole, so it paid very little special attention to the generation and transmission of electromagnetic waves. Combining that with its difficult math, the book would puzzle physicists for another decade before they see the light from the book, and made it a rather curious period of history in electromagnetism.

> I've been trying to find these 4 equations in Heaviside's writing but so far have not been successful.

In 1885, Heaviside published Electromagnetic Induction and Its Propagation in The Electrician, and formulated what he called the "Duplex Form" of Maxwell's equation. This was a long series of papers published in several months, and later republished in Electric Papers, Volume I. Basically, following his physical intuition, he felt that electric and magnetic fields should be symmetric and generate each other, and that should be directly highlighted in equations.

The logic of the paper went like the following.

First, he started with a definition of electric current [1]:

    C = kE
    D = cE / 4π
    Γ = C + D
in which, E denotes electric force, C denotes conduction current, k denotes specific conductivity constant, D denotes displacement current, and c denotes dielectric constant. Finally, Γ denotes true electric current, which is the sum of the conduction and displacement terms.

Next, a definition of magnetic current [2]:

    B  = µH
    G  = Ḃ / 4π = µḢ / 4π
    G' = gH + µḢ / 4π
H denotes magnetic force, B denotes magnetic induction, µ denotes permeability, G denotes magnetic current, Ḃ and Ḣ are derivatives of B and H (Newton's notation). Hypothetically, suppose that magnetic monopoles exist (Heaviside did so), G' would denote the "true magnetic current", with an extra conduction term gH, where g is a constant similar to k.

Then, he introduced the concepts of divergence and curl, and their physical significance [3]. After more discussion and derivation, he finally wrote [4]:

    curl (H - h)  = 4πΓ = 4πkE + cĖ
    -curl (e - E) = 4πG = 4πgH + µḢ
in which, e and H denote impressed electric and magnetic forces to take static fields into account. Finally, since magnetic monopoles don't exist, he made g = 0, but kept this term in the equations for symmetry and elegance. [0]

This is the core of Heaviside's Duplex Form of Maxwell's equations. one can clearly see the co-evolution of electric and magnetic fields, and is the precursor of the modern Maxwell's equations as we know today in its vector calculus formulation. As far as I know, his treatment of "physical" vectors as first-class objects is his original invention (independently invented by Gibbs as well), although the concepts themselves came from quaternions.

This is not a complete summary, as he continued his analysis in a series of publications.

A good book on this part of history is Oliver Heaviside: the life, work, and times of an electrical genius of the Victorian age, by Paul J. Nahin.

[0] So the claim "Maxwell's equations need modifications if magnetic monopole has been discovered" is historically inaccurate, it should rather be, "be restored to Heaviside's original form."

[1] Electric Papers, Volume I, Page 429, https://archive.org/details/electricalpapers01heavuoft/page/...

[2] Page 441: https://archive.org/details/electricalpapers01heavuoft/page/...

[3] Page 443: https://archive.org/details/electricalpapers01heavuoft/page/...

[4] Page 449: https://archive.org/details/electricalpapers01heavuoft/page/...

segfaultbuserr··on Herbie: Optimize Floating-Point Expressions
I tried the quadratic equation, which is infamous for its loss of precision in numerical computations. The best output was merely a rewrite of the original equation using FMA (I'm not sure if I was using the tool correctly, perhaps it could find more creative solutions given the right inputs). But at least it also found the standard approximation -c/b & -b/a, applicable when the real roots are far from each other (which is the case for many practical problems) [0].

[0] https://web.mit.edu/6.969/www/readings/quadratic-formula.pdf

segfaultbuserr··on The CompCert C Compiler
> and the syntax isn't that great, nobody even knows they exist.

Linux kernel maintainers certainly do - they even invented a non-standard extension of C99's VLA notation (their own invention, not implemented in GCC)... In newer versions of the package "Linux man-pages", most libc functions are documented using a non-standard extension of that syntax. For example, the prototype of memcpy() now reads [0]:

    void *memcpy(void dest[restrict .n], const void src[restrict .n], size_t n);
It means this function accepts two arrays (dest[] and src[]), with n bytes of data type "void", src[] is read-only ("const"), and both src[] and dest[] are non-overlapping ("restrict"). Two non-standard notations are used here: a "void" array with elements of unknown type (not allowed in C99), also, ".n" means "a variable in the argument list, but defined after this variable" (which is not allowed in C99).

The equivalent C99 definition would be something similar to:

    void *memcpy(size_t n, char dest[restrict n], const char src[restrict n]);
As more people are exposed to these man pages, the C99 syntax hopefully will have more publicity. Finally, C23 interestingly states that:

> 15. Application Programming Interfaces (APIs) should be self-documenting when possible. In particular, the order of parameters in function declarations should be arranged such that the size of an array appears before the array. The purpose is to allow Variable-Length Array (VLA) notation to be used. This not only makes the code's purpose clearer to human readers, but also makes static analysis easier. Any new APIs added to the Standard should take this into consideration.

[0] https://git.kernel.org/pub/scm/docs/man-pages/man-pages.git/...

[1] https://www.open-std.org/jtc1/sc22/wg14/www/docs/n2611.htm

segfaultbuserr··on Psyche-C: automatic compilation of partially-available C programs
> What is the purpose of this?

It's a demo application of the research team's C compiler front-end [0], and demonstrates its automatic type inference feature. The compiler front-end itself is intended to be a general-purpose library to power high-level static analysis tools. While this demo itself is not very useful, being able to do static analysis on incomplete code snippet is certainly extremely useful.

[0] https://github.com/ltcmelo/psychec

segfaultbuserr··on What's the difference between a motor and an engine? (2013)
If you're not confused enough, here's another exercise: try distinguishing generator, dynamo and alternator.
segfaultbuserr··on The Mirror Fusion Test Facility (2023)
> I can see the need for maintaining what we have already [...]

The officially-stated goal of these labs (pulsed power, fusion, and hydrodynamic test facilities [0]) is indeed for maintaining existing nuclear weapons, not to design new ones (and also for doing basic research during free time). This was called the Science Based Stockpile Stewardship program [1] - ensure that existing nuclear weapons would remain functional in the foreseeable future. (Interestingly, the lesser-known hydrodynamic test facilities such as the Dual-Axis Radiographic Hydrodynamic Test Facility are more useful for weapon designs than fusion facilities).

The idea is to test materials under extreme lab conditions to help computer modeling, so that it would still be possible to do minor design changes to replace obsolete or end-of-life parts (the FOGBANK incident [1] came to mind). Understanding long-term aging is also a stated goal.

[0] http://www.wslfweb.org/docs/agex.htm

[1] https://en.wikipedia.org/wiki/Stockpile_stewardship

[2] https://en.wikipedia.org/wiki/Fogbank

segfaultbuserr··on A BSD person tries Alpine Linux
It was pretty clear to me that the comment was a description of the respective characteristics of glibc and musl in terms of security, while avoiding any conclusion: glibc has heap hardening, which is good for security, but a complex codebase, which is bad for security. Meanwhile, musl is small and understandable, which is good for security, but with a naive codebase that lacks hardening, which is bad for security. Which is better is intentionally left to the reader to avoid flamewars.
Page 1 of 34Next →