HNHacker News
TopNewBestAskShowJobs

eddyb

465 karma · joined February 28, 2014

submissionscomments
eddyb··on Zero-cost futures in Rust
Compile error if it's holding on to any borrowed data on the stack - however, borrowing is only the default, if you move from a capture your closure will contain that value instead of reference to the value, and if you have "move" in front of your closure, all the captures are contained in the closure, so by boxing or by using -> impl Fn(...) you can safely return it.

If you need to hold onto data accessed from multiple locations, you'd need to use Rc<T> instead of keeping that data on the stack, which is closer to GC'd languages (and pretty similar to Swift, IIUC, except Rust is more explicit).

eddyb··on Zero-cost futures in Rust
Not even in nightly yet - follow https://github.com/rust-lang/rust/pull/35091 for updates (we're trying to land it this week).
eddyb··on Mozilla Awards $585k to Nine Open Source Projects
It is a build system written in Rust: https://github.com/rust-lang/rust/tree/master/src/bootstrap

There is only one Python file: https://github.com/rust-lang/rust/blob/master/src/bootstrap/...

Its job is to download the previous Rust release (beta for master, stable for beta and previous stable for current stable), including Cargo, build the build system with Cargo and start it.

The build system then builds native dependencies (LLVM) and runs Cargo several times for different components.

eddyb··on Using ImGui with modern C++ and STL for creating game dev tools – Part 2
But how many times can you access the DRAM in that time?

One thing you have to keep in mind is that the more dynamic the scene graph, and the larger its memory footprint, the more frequent you'll miss the cache while traversing it.

It's possible to manually write optimized flex-box-like layout code, where the overall structure, if not exact position, is mostly static.

For example, if you always have one resizable panel to the side of two viewports, of equal width (as is the case in the editor I use atm), you can easily do `viewport_width = (total_width - panel_width) / 2` and be done.

Those operations, if done with integers, should be faster than a read from cache, or if done with floating-point numbers, (much) faster than going to DRAM in the even of a cache miss.

However, doing all of that by hand would be a pain, and it's hopeless in the face of user customization. SIMD is another resource that's completely unusable in the face of truly dynamic data, and a time sink to do manually.

What we're missing is automation of mostly-static UIs, generating specialized code and optimizing it into computations that run under a microsecond. For HTML, the hope would be JITs, but highly dynamic DOMs are still all over the place, so you're not going to have as many gains as a dedicated UI system, and it wouldn't be free for the user.

eddyb··on Iris: Fast back-end web framework for Go
With generics, yes, like C++ templates: each combination of parameters results in one instance.

However, you can opt into using trait objects instead, which do use virtual calls.

eddyb··on Afl.rs: Fuzzing Rust code with american-fuzzy-lop
You can do it with either -C debug-assertions (stable) or -Z force-overflow-checks=on (unstable). The difference is that the former will also enable debug_assert! in your own code.
eddyb··on Afl.rs: Fuzzing Rust code with american-fuzzy-lop
That's LLVM's abort intrinsic - no a lot of those left around, I believe OOM, panicking while panicking and stack overflow use it or used to use it.
eddyb··on Afl.rs: Fuzzing Rust code with american-fuzzy-lop
> These are mostly reported using return values of type Error

However, unlike more dynamic languages, the return types are usually Result<SuccessValue, SomeErrorType>, e.g. writing to a file or a socket will get you Result<usize, io::Error>, values of which can be Ok(usize) (the number of bytes written) or Err(io::Error) (an I/O error which you can further inspect).

Although if you want to, you can use Result<T, Box<Error>> where Error is a trait implemented by types representing some error information. You then lose the ability to inspect errors other than for reporting them wholesale, but the APIs are slightly more uniform, type-wise.

You can read about all of this and more in the dedicated chapter of the official Rust Book: https://doc.rust-lang.org/book/error-handling.html.

eddyb··on The Path to Rust
mio has nothing to do with callbacks in its design, unless you build such an abstraction yourself on top of it.

However, without an ergonomic way to create state machines (i.e. generators), it's hard to use at all in the intended fashion (small per-connection state instead of large boxed closures or coroutine/thread stacks).

eddyb··on Teaching C
I believe C++14 adds the first sane C-family array type:

template<typename T, size_t N> class array { T data[N]; };

eddyb··on Porting a Haskell graphics framework to Rust
My understanding was that the goal of that research project is to produce tools for proving code written in unsafe Rust as safe at the API boundary (with limitations around FFI and inline assembly, of course).
eddyb··on Introducing MIR
Well, actually, we store MIR in crate metadata. However, it's not a special serialization format, just the rustc_serialize infrastructure, i.e. you could also serialize MIR to JSON if you really wanted to.
eddyb··on Introducing MIR
You are correct that a Rust interpreter would work great on MIR, which is why Scott Olson (@tsion) has been working on one: https://github.com/tsion/miri (check out the slides and the report).
eddyb··on Introducing MIR
There is discussion of a WebAssembly backend which lowers MIR directly, see https://github.com/rust-lang/rust/issues/33205.
eddyb··on Segfaults are our friends and teachers
AFAIK this could've been fixed in LLVM a long time ago, which would also help C code compiled with clang, although I believe stack probes are opt-in there.

We can't fix it fully in rustc because we don't know the stack size, which can grow with aggressive inlining, for example.

I suppose we could summarily probe allocas we know are larger than the page size, which would solve this particular situation (one large variable), but it's not a panacea.

eddyb··on Concurrency in Rust
If you write `let data = data;` in the closure, then you can achieve the effect of `move` without the special syntax (unless `data` can be copied, in which case there is really no way to force it to be moved inside the closure without `move`).

However, in Rust, you get an error if you accidentally capture by reference in the closure passed to `thread::spawn` so it's not as hard to get right as it is in JS.

eddyb··on A previously unnoticed property of prime numbers
The hashing optimization isn't necessary as using String at all is wasteful - my code ended up being simpler, but in the end the largest gain came from replacing u64 with u32 - see https://news.ycombinator.com/item?id=11290955.
eddyb··on A previously unnoticed property of prime numbers
LuaJIT 2.0.4:

3.67user 0.01system 0:03.68elapsed 99%CPU

rustc 1.9.0-nightly (74b886ab1 2016-03-13) (-C opt-level=3):

5.18user 0.00system 0:05.20elapsed 99%CPU

Switching to BTreeMap gives me:

4.36user 0.00system 0:04.38elapsed 99%CPU

Using u8 as the key (last_digit0*10 + last_digit1) instead of a string:

4.18user 0.00system 0:04.18elapsed 99%CPU

I tried preallocating the vector of primes and it didn't help, strangely enough.

Replacing the floating-point sqrt with squaring in the comparison does bring it a bit lower:

4.04user 0.00system 0:04.05elapsed 99%CPU

I don't know how to bring that number lower without using a sieve, perf reports that most of the time is spent in:

86,31 │ div %rbx

I've also just noticed that the Lua and the Rust code don't give the same results, but I can't easily tell why.

Oh! The largest prime is 0x00ec4bab, so they can be stored as u32. Final Rust result:

2.33user 0.00system 0:02.33elapsed 99%CPU

Code: https://gist.github.com/eddyb/51a92fa2edf20d6e23fe

eddyb··on A previously unnoticed property of prime numbers
In a sibling answer, steveklabnik suggests the Rust version was compiled without optimizations - https://news.ycombinator.com/item?id=11285569

Could you try running both on the same machine? I'm curious if LuaJIT can still beat Rust if both have optimizations working.

I know it can beat native code sometimes, which is pretty impressive (it finds common cases and specializes to them AFAIK, almost like "sufficiently advanced optimizing compiler" fairy tales).

eddyb··on AlphaGo beats Lee Sedol 3-0 [video]
Accelerando[0] starts with this quote, which I really like:

"The question of whether a computer can think is no more interesting than the question of whether a submarine can swim."

– Edsger W. Dijkstra

[0] http://www.antipope.org/charlie/blog-static/fiction/accelera... (readable online at http://www.antipope.org/charlie/blog-static/fiction/accelera...)

eddyb··on No Compiler – On LLVM, and writing software without a compiler
That's what the well-typed check I mentioned before is, see https://github.com/rust-lang/rust/pull/31474
eddyb··on No Compiler – On LLVM, and writing software without a compiler
We don't really have our own optimizations atm, but we do want MIR optimization passes.

The benefits would be two-fold:

We could do transformations LLVM can't figure out itself, like NVRO: `fn bar() -> T {let mut x = ...; foo(&mut x); x} ... box bar()` can (theoretically) end up calling `foo` with a pointer to the heap instead of having to copy `x` from the stack later.

And we can lift the burden of LLVM trudging through monomorphized instances from MIR: since MIR still has type parameters, you can transform it and simplify the resulting LLVM IR for all instances, e.g.:

mem::replace(mut_ref, x) is {mem::swap(mut_ref, &mut x); x} which can be reduced once to {Return = * arg0; * arg0 = arg1;} in MIR and translated to a load, a store and a ret (for immediates) or two memcpy's (for indirect argument & return), depending on the type.

As most of the time rustc spends is in LLVM optimizations, this will result in compilation time speedups linear in the number of instances, sometimes exponential in the size of the original code.

eddyb··on No Compiler – On LLVM, and writing software without a compiler
Depends what you mean by "match" statements: if you're referring to complex patterns, yes, they get desugared.

If you're instead referring to matching over ADT (enum) variants, that is a MIR primitive, the "Switch" terminator (there is also a "SwitchInt" for integer-like values, as in C or LLVM IR).

As for the interpreter, it's quite likely that we'll use (something like) it for evaluating constants, but not entirely certain atm.

I personally hope we can get the CTFE semantics I described (see [1] above), which despite being pure, go far beyond C++17 constexpr capabilities (e.g. you could potentially run an entire compiler for a different language at compile-time).

eddyb··on No Compiler – On LLVM, and writing software without a compiler
Yes, there are no immediate plans for anything other than MIR at the imperative-semantic level (the compiler also uses HIR, which is mostly a desugared AST and not related to LLVM IR).

Currently, initial support for lowering MIR to LLVM IR is close to complete and it might be enabled in the next beta (or the one after that).

None of the analysis passes have started their transition to MIR, though (there is a well-typed check to catch bugs in MIR generation, but that's about it).

There's also talk of using MIR for compile-time (pure) function evaluation [1] and some out-of-tree experimentation with interpreting it [2].

[1] https://internals.rust-lang.org/t/mir-constant-evaluation/31...

[2] https://github.com/tsion/miri

eddyb··on What is “the stack”?
It has a predefined set of calling conventions (sadly, this doesn't list x86 and other architecture-specific ones): http://llvm.org/docs/LangRef.html#calling-conventions

GHC and HiPE appear to have gotten themselves LLVM calling conventions, I guess. But none of this can be customized further without modifying LLVM.

Besides, none of these choices seem to affect anything other than the (ordered) set of registers available for arguments, and the sets of caller/callee-save registers (i.e. what you expect from x86 calling conventions).

Once you get to the stack it's really all the same.

What I meant was that Rust could have picked any of those choices, which would make it incompatible with C (it already is for some argument types because of the simplistic handling around them), but it couldn't have used its own special one, not without adding it to all LLVM targets it wants to support.

eddyb··on What is “the stack”?
The x87 FPU "stack" is not truly a stack though: it's rather a ring buffer of 8 registers.

And you can model it pretty well even without a "head pointer", just with moves between all registers, e.g. push(x) is ST(7) = ST(6), ST(6) = ST(5), ..., ST(1) = ST(0), ST(0) = x.

I've implemented it as such in a lifter from x86 machine code and all the redundant moves just disappear, if the uses of the pseudo-stack are balanced, and you're left with traditional registers.

eddyb··on What is “the stack”?
At no point while using LLVM, could Rust have used a different calling convention.

It was managing its stack, but it does that even today: when creating a thread, it allocates a 8MB stack.

Within that stack, however, function calls work as they do in C, and they've been like that effectively forever (although the Ocaml bootstrap compiler might have had a non-LLVM x86 backend, not sure about that).

eddyb··on What is “the stack”?
"This isn't a theoretical question at all -- Rust actually has a different calling convention from C"

Technically correct, but not at the register/stack management level: LLVM still controls all of that, Rust only chooses the order of the arguments and whether they are passed "directly" (in registers or on the stack, when registers are exhausted) or "indirectly" (pointer to the value, whether it's on the stack or somewhere else).

So the following statement, "which means it sets up its stack differently" is incorrect: even if Rust wanted to, it couldn't do so with the LLVM backend.

eddyb··on 555 timer teardown: inside the world's most popular IC
It's been rewritten in JavaScript so you don't need Java at all.
eddyb··on Glibc getaddrinfo stack-based buffer overflow
What you can do for mmap is have several abstractions (or one using generics and phantom types), one for each different set of usecases, with different access modes.

Examples would be:

* read-only: &ROMemMap -> &[u8]

* read-write: &mut RWMemMap -> &mut [u8]

* write-only: fn set(self: &mut WOMemMap, i: usize, b: u8),

or more generally: &mut WOMemMap -> &mut [WriteOnly<u8>] where WriteOnly<T> has fn set(&mut self, T)

Now for the shared case, consider this: aliasing rules can be avoided with atomic operations, i.e. Arc<AtomicUsize> is shared and can be safely (atomically) read/written by multiple threads.

In the multi-process case, you could provide an atomic API, although we don't currently seem to expose byte-level atomics (likely not present on some platforms) so if you wanted to write a demo you'd need to use the unstable intrinsics atm.

FWIW restricted to single-threaded code, this results in the Cell get/set API which is not hardware-atomic but cannot overlap with other accesses to the same memory, as Cell doesn't implement Share so any threading abstraction will block you from doing any kind of sharing of Cells.

That is, you can share &[Cell<u8>] pointing to any bag of bytes that sits in read-write memory, and anything in the same thread can read or write to it, safely.

← PreviousPage 4 of 6Next →