HNHacker News
TopNewBestAskShowJobs

devit

4,235 karma · joined September 10, 2015

submissionscomments
devit··on DeepGEMM: clean and efficient FP8 GEMM kernels with fine-grained scaling
Multiply FP8 matrices with FP32 scaling factors giving a bfloat16 matrix result on an nVidia Hopper or newer GPU.
devit··on The Shape of a Mars Mission
This seems to be a problem with rocket/lander technology resulting in a ~900kg weight limit on Curiosity.

According to Internet searches, Starship can bright 100 tons to Mars surface.

A common large Earth backhoe seem to weight 20 tons, so with Starship you can just ship one and it will be capable of driving at normal speeds (up to 100km/h), excavating for meters and not centimeters, etc.

(obviously it would need adaptations since diesel engines need air that isn't present on Mars and EV batteries might have problems with the cold, but it would be a similar weight magnitude)

devit··on RT64: N64 graphics renderer in emulators and native ports
Doesn't seem to be a noticeable improvement from N64-era graphics.

It probably needs generative AI based upscaling to high resolution meshes, textures and realistic materials to actually achieve a quality improvement.

devit··on Nvidia Security Team: “What if we just stopped using C?” (2022)
That would be a good start as the current syntax is absurd.
devit··on Intel's Battlemage Architecture
I wonder if a multiplexer would be feasible?

Hardware-wise instead of putting the chips on the PCB surface one would mount an 16-gonal arrangement of perpendicular daughterboards, each containing 2-16 GDDR chips where there would be normally one, with external liquid cooling, power delivery and PCIe control connection.

Then each of the daughterboards would feature a multiplexer with a dual-ported SRAM containing a table where for each memory page it would store the chip number to map it to and it would use it to route requests from the GPU, using the second port to change the mapping from the extra PCIe interface.

API-wise, for each resource you would have N overlays and would have a new operation allowing to switch the resource overlay (which would require a custom driver that properly invalidates caches).

This would depend on the GPU supporting the much higher latency of this setup and providing good enough support for cache flushing and invalidation, as well as deterministic mapping from physical addresses to chip addresses, and the ability to manufacture all this in a reasonably affordable fashion.

devit··on Asahi Linux lead developer Hector Martin resigns from Linux kernel
I think Linus should change this though, so that the maintainers that don't want to learn Rust have to step down, and there are no more conflicts.
devit··on The origins of 60-Hz as a power frequency (1997)
Actually, you want to use a DC-to-DC converter that properly delivers the constant desired reduced voltage, rather than making a ridiculous stroboscopic light (aka PWM).
devit··on The missing cross-platform OS API for timers
The article's conclusion is completely wrong because it misses two crucial points:

1. On multicore machines you want to process timers in parallel on multiple cores. With userspace timers you either set the same timeouts on all threads and have unnecessary wakeups or distribute timers to cores ahead of time which leads to increased latency if a thread is stalled for any reason. I think this is unfixable without a dedicated timer API.

2. Good timer APIs let you set a time _interval_ for when the timer expires, which is essential so that the system can group timers and reduce wakeups (i.e. you process all timers where the lower bound has been reached before going to sleep, but don't wake up until the upper bound arrives). Most or all "wait with timeout" APIs only have a single timeout, although this could be fixed.

devit··on Open Euro LLM: Open LLMs for Transparent AI in Europe
Huh? Most Firefox users presumably use it, and anyway it's obviously essential and extremely useful functionality.

And the really important languages for EU/US audiences are, in order, English, Chinese, Spanish, Japanese, Portuguese, German, Italian, which is, guess what, 7 languages...

devit··on Over 90% of U.S. airport towers are understaffed, data shows
It's probably not done due to political reasons.

Most of the complication is probably accidental, due to history or human traditions.

From first principles, the problem is obviously straightforward to formulate (assign nonoverlapping regions of spacetime to aircraft containing their position and destination, that they can maneuver in and such that the travel time is approximately optimal).

Applying simplifying constraints to the form of the regions (e.g. a discrete set of departure slots and fixed takeoff/landing envelopes, a route that follows the optimal trajectory in latitude/longitude plus a discrete lateral offset, discrete set of altitudes that change only at route crossings), it should be possible to reduce to a discrete optimization problem solvable in linear time.

devit··on Over 90% of U.S. airport towers are understaffed, data shows
Seems like it can be at least mostly automated.
devit··on Nvidia sheds almost $600B in market cap, biggest one-day loss in US history
That depends on how whether the demand increase multiplier due to the lower cost per result is lower or higher that the efficiency increase multiplier. It can be either in general.
devit··on Wild – A fast linker for Linux
It's feasible to write complex correct programs with optimal performance in Rust, unlike any other programming language (complex+correct is not feasible in C/C++/assembly/Zig/etc., optimal performance not possible in any other language).
devit··on Wild – A fast linker for Linux
I think the optimal approach for development would be to not produce a traditional linked executable at all, but instead just place the object files in memory, and then produce a loader executable that hooks page faults in those memory areas and on-demand mmaps the relevant object elsewhere, applies relocations to it, and then moves it in place with mremap.

Symbols would be resolved based on an index where only updated object files are reindexed. It could also eagerly relocate in the background, in order depending on previous usage data.

This would basically make a copyless lazy incremental linker.

devit··on OpenAI fails to deliver opt-out system for photographers
Well then the model weights would be a compilation of copies of the original works, which has the same effect as it being a derivative work unless the copyright holder chose to allow copies but not derivative works.
devit··on OpenAI fails to deliver opt-out system for photographers
Model weights, if they can reproduce something like the original, are just a form of lossy compression (or even lossless for text), where the LLM answering the prompt is a more powerful version of asking software to retrieve a specific file from a Zip archive (or a webserver answering an HTTP query) of such lossy compressed data.

So if model weights don't infringe, that would also imply that saving an image as a JPG or a video using AV-1 doesn't infringe, which would obviously effectively implies that copyright doesn't apply to images or videos on the web, which is not current law/policy, so I think that reasoning cannot possibly work.

devit··on OpenAI Fails to Deliver Opt-Out System for Photographers
Aren't lawsuits the proper way to address this?

Seems like there's an argument that model weights are a derivative work of the training data, at least if the model is capable of producing output that would be ruled to be such a derivative work given minimal prompting.

Although it may not work with photography since the model might just almost exclusively learn how the object of the photo looks in general and how photos work in general, rather than memorizing anything about specific photos.

devit··on Ropey – A UTF8 text rope for manipulating and editing large text
"Handling texts that are larger than available memory. Ropey is an in-memory data structure."

That seems to make it of dubious use, not really suitable for a well-engineered text editor.

The fact that it's UTF-8 only is also a serious problem since files can contain arbitrary byte sequences.

devit··on All clocks are 30 seconds late
That's only the case if you interpret them as HH:MM:00. The issue is fixed if you interpret them as HH:MM:30 instead.
devit··on ELKS: Linux for 16-bit Intel Processors
That makes no sense, since DOSEMU is based on virtual 8086 mode, which requires a 80386, while ELKS is for 8086-286 CPUs.

You can just run MS-DOS or FreeDOS directly on those machines, that's what the machines were made for.

devit··on Phase behavior of Cacio and Pepe sauce
50g of carbs per "meal" are only enough if you eat 10-15 meals per day.
devit··on One Dog vs. the Windows 3.1 Graphics Stack
Linux supports vgafb/vesafb so this is possible if the distribution is configured appropriately.

I think some/most distributions might not enable it out of the box because it would generally result in a low performance/quality experience and the user not realizing what the problem is, and nowadays almost all GPUs are supported natively, so nobody has invested in writing code to show a "Using unaccelerated VGA/VESA, you may want to fix this" popup.

devit··on Tesla Cybertruck sales are disastrous
The major issue seems to be that Musk somehow wants people to pay him to drive it, as opposed to him paying people to drive it.
devit··on Execution units are often pipelined
The amateurs usually run benchmarks (because they can't reason about it as they lack the relevant knowledge) and believe they got a useful result on some aspect when in the fact the benchmark usually depends on other arbitrary random factors (e.g. maybe they think they are measuring FMA throughput, but are in fact measuring whether the compiler autovectorizes or whether it fuses multiply and adds automatically).

A pro would generally only run benchmarks if it's the only way to find out (or if it's easy), but isn't going to trust it unless there's a good explanation for the effects, or unless they actually just want to compare two very specific configurations rather than coming up with a general finding.

devit··on AI companies cause most of traffic on forums
DDOS is different from crashing.

And I doubt Facebook implemented something that actually saturates the network, usually a scraper implements a limit on concurrent connections and often also a delay between connections (e.g. max 10 concurrent, 100ms delay).

Chances are the website operator implemented a webserver with terrible RAM efficiency that runs out of RAM and crashes after 10 concurrent requests, or that saturates the CPU from simple requests, or something like that.

devit··on Cable-cutting tanker seized by Finland 'was loaded with spying equipment'
Isn't the tanker much more expensive than repairing the cable damage? (and Russia has much lower GDP than the EU)

Seems like they can't play this game for long.

devit··on Carlsen disqualified from World Rapid and Blitz championship for wearing jeans
Probably just burned out from chess, especially training the whole day for competitions.

We are animals and eventually the brain will rebel against extreme repetitive mental effort that is perceived to be at least partially useless (and given he's already been the world champion, it's easy for part of him to think there's no point in training).

devit··on Debian's approach to Rust – Dependency handling (2022)
The correct solution to packaging Rust crates for a distribution seems to be to compile them as Rust dynamic libraries and use package names like "librust1.82-regex-1+aho-corasick" for the regex crate, semver major 1, compiled for the Rust 1.82 ABI, for the given set of features (which should be maximal, excluding conflicting features).

For unclear reasons it seems Debian only packages the sources and not the binaries and thus doesn't need the Rust version, and also apparently only packages the latest semver version (which might make it problematic to compile old versions but is not an issue at runtime).

devit··on GPT-5 is behind schedule
I think the problem of those is that they are special purpose, and probably too expensive and bulky for that single purpose.

A single general-purpose robot that can do everything would be much easier to sell.

devit··on Fermat's Last Theorem – how it's going
I don't think that's a good example, since it's obviously provably true and the proof is also obvious (either use well-known lemmas like log(poly(x)) in O(x^a) or show that lim_x->inf lhs/x = 0 via L'Hôpital's rule, differentiation algorithms and well-known limits at infinity).
← PreviousPage 2 of 34Next →