HNHacker News
TopNewBestAskShowJobs

timhh

814 karma · joined December 24, 2017

submissionscomments
timhh··on How to improve the RISC-V specification
On 2 there is this now:

https://alasdair.github.io/#_modules

But it's very new and the RISC-V model doesn't use it yet. I think it's also just a first step rather than fully solving the problem.

timhh··on How to improve the RISC-V specification
That's not really what I meant. It's very easy to configure which instructions or CSRs exist, and excluding custom extensions there are really only two options for behaviour, so you just have a flag `haveSomeExtension()` to enable the instructions and CSRs. The Sail model has some of these flags already.

If you write to a WARL field in a CSR the chip can use more or less any legalisation logic it wants. Configuring that is very difficult (though there is a decent attempt in riscv-config).

timhh··on How to improve the RISC-V specification
Yeah I completely agree. Especially annoying if RISC-V is the first ISA you've learnt which is probably the case for a lot of people.

I don't think you meant D3-1 btw.

timhh··on How to improve the RISC-V specification
I've been doing a lot of work with Sail (not SAIL btw) and I'm not sure I agree with the points about it.

There's already a way to extract functions into asciidoc as the author noted. I've used it. It works well.

The liquid types do take some getting used to but they aren't actually used in most of the code; mostly for utility function definitions like `zero_extend`. If you look at the definition for simple instructions they can be very readable and practically pseudocode:

https://github.com/riscv/sail-riscv/blob/0aae5bc7f57df4ebedd...

A lot of instructions are more complex or course but that's what you get if you want to precisely define them.

Overall Sail is a really fantastic language and the liquid types really help avoid bugs.

The biggest actual problems are:

1. The RISC-V spec is chock full of undefined / implementation defined behaviour. How do you capture that in code, where basically everything is defined. The biggest example is probably WARL fields which can do basically anything. Another example is decomposing misaligned accesses. You can decompose them into any number of atomic memory operations and do them in any order. E.g. Spike decomposes them into single byte accesses. (This problem isn't really unique to Sail tbf).

2. The RISC-V Sail model doesn't do a good job of letting you configure it currently. E.g. you can't even set the spec version at the moment. This is just an engineering problem though. We're hoping to fix it one day using riscv-config which is a YAML file that's supposed to specify all the configurable behaviour about a RISC-V chip.

I definitely agree about the often wooly language in the spec though. It doesn't even use RFC-style MUST/SHOULD/MAY terms.

timhh··on Lezer: A parsing system for CodeMirror, inspired by Tree-sitter
> no, I don't need to accept your PR just because you put work in it, you should have checked with me first, before doing it

Which is a bit rude IMO, and not really in the spirit of open source. I never demanded that he accept it. A PR is like "here's some code, let me know if it's ok" not "you should merge this code without question".

Here is what I would have written:

"Hi, thanks for the code but I don't think I want to go in that direction because X Y Z." (he didn't give any clear reasons so you'll have to imagine those).

Anyway he's free to do his thing. I was just explaining why I moved away from that library.

timhh··on Lezer: A parsing system for CodeMirror, inspired by Tree-sitter
Yes it's unsolicited "advice" that shows an uncooperative attitude IMO. You can say "I don't want to do that" in a nice way without being patronising. If you look at the other closed MRs you'll see a similar attitude.

E.g. here (I'd forgotten about this actually): https://github.com/lezer-parser/generator/pull/6#issuecommen...

Here https://github.com/lezer-parser/lr/pull/64#issuecomment-1802...

It's nothing major but just emanates "difficult to work with" vibes so I didn't want to spend my time working with a project like that (looks like not many other people do either).

timhh··on Lezer: A parsing system for CodeMirror, inspired by Tree-sitter
I attempted to use this but was disheartened but the fact that it doesn't statically type node names. Tree Sitter doesn't either but it has much more of an excuse given that it targets C.

https://github.com/lezer-parser/lezer/issues/8

The dev seems mildly hostile to outside involvement too, so I moved on. These days I use Chumsky which is Rust rather than Typescript, but also way more awesome, if you can deal with the often incomprehensible compilation errors at least!

https://github.com/zesterer/chumsky

timhh··on Ludic: New framework for Python with seamless Htmx support
Haha someone didn't like that I'm pointing out flaws in Python and down-voted that question.
timhh··on Ludic: New framework for Python with seamless Htmx support
Here's a bug I ran into with Python async:

https://stackoverflow.com/q/78036302/265521

I've written about 200 lines of Python async code in total, so to run into a bug like that so soon was not encouraging.

I suspect it's just not a very popular feature and so it doesn't get a lot of use and debugging. And it's Python so it really needs a lot of ton of real world use to detect bugs.

Anyway I'm not going to waste my time debugging Python internals so I just switched to Deno.

timhh··on Parsing URLs in Python
IMO that URL crate is not especially high quality. I barely work with URLs and I quickly found an embarrassingly trivial bug:

https://github.com/servo/rust-url/issues/864#issuecomment-16...

timhh··on Show HN: Open-source, browser-local data exploration using DuckDB-WASM and PRQL
Neat. I did a similar thing for analysing the output of commands that can produce JSON (SLURM in this case). Worked pretty well but I immediately ran into PRQL's lack of support for DuckDB's struct types.
timhh··on Candy – a minimalistic functional programming language
That's completely feasible and there are languages that do this. It doesn't really eliminate the need to run your program unless the inputs to your program are also completely restricted types like One, Two, Three. In which case yeah, you don't need to run it and the type system can just tell you the answer.

I believe you can do that sort of thing in loads of type systems, e.g. Typescript, but there are languages that intentionally support it. I use a niche DSL that has fancy types like this called Sail. https://github.com/rems-project/sail

In my experience the downsides of these fancy "first class type systems" are

1. More incomprehensible error messages.

2. The type checker moves from a deterministic process that either succeeds or fails in an understandable way, to SMT solvers which can just say "yep it's ok" or "nope, couldn't prove it", semi-randomly, and there's little you can do about it.

Still my experience of Sail is that it's very comfortable to go a little bit further into SMT land, and my experience of Dafny is that it's very unpleasant to go full formal-verification at the moment.

I've done a fair bit of hardware formal verification too and that's a different story - very easy and very powerful. I'm hoping one day that software formal verification is like that.

timhh··on Pratt Parsers: Expression Parsing Made Easy (2011)
Ah yeah I should have said "a library for which the integration is easier than writing my own library from scratch"!
timhh··on Pratt Parsers: Expression Parsing Made Easy (2011)
I had to write an expression parser recently and found the number of names for essentially the same algorithm very confusing. There's Pratt parsing, precedence climbing, and shunting yard. But really they're all the same algorithm.

Pratt parsing and precedence climbing are the same except one merges left/right associativity precedence into a single precedence table (which makes way more sense). Shunting yard is the same as the others except it's iterative instead of recursive, and it seems like common implementations omit some syntax checking (but there's no reason you have to omit it).

Here's what I ended up with:

https://github.com/Timmmm/expr

I had to write that because somewhat surprisingly I couldn't find an expression parser library that supports 64-bit integers and bools. Almost all of them only do floats. Also I needed it in C++ - that library was a prototype that I later translated into C++ (easier to write in Rust and then translate).

timhh··on Show HN: I made TV Sort, a web-based game for ranking TV show episodes
Due to the regularisation yeah they all start at the same rating. But you don't need many votes to start getting good ratings.

I introduced this method to Dyson for objectively calculating very subjective measurements (e.g. "how frizzy does this hair look?"). We basically crowd sourced it to other engineers.

I did a load of studies on different methods by ranking something that's sort of hard to rank but you know the answer to - I used 10 grey squares that only differed by 2/255 and you had to pick the brighter one.

Some other things:

1. I don't remember the exact details but there's a slight extension of the method where you give each user a "how good are you" coefficient that you simultaneously solve for. This helps eliminate people that vote randomly, and also inverts the votes of people that deliberately pick the wrong answer (as long as they're consistently wrong).

2. You can put confidence limits on the values very easily too since it's a MAP estimate. Actually I showed curves for each item - basically how does the model probability vary as you sweep one rating up and down a bit. People didn't understand it at all though.

3. You can calculate the rankings incrementally very quickly (details in the answer) which means you can show users comparisons that give the most information. This usually means you end up showing users endless difficult choices which can frustrate them, especially if it's a forced choice.

4. I never found a principled way to incorporate a "they look the same" option. I tried some ad-hoc methods and IIRC a "much better, slightly better, can't tell, slightly worse, much worse" scale gave the fastest convergence but it was pretty unsatisfying that I just used some as hoc method to add the results.

It was all closed source and I haven't worked there for years so the code is lost to the wind unfortunately.

timhh··on Show HN: I made TV Sort, a web-based game for ranking TV show episodes
You can do much better than just sorting. The simplest is to use Bradley-Terry. It's a very simple algorithm and will let you combine results from multiple users and gives an actual rating rather than just a ranking.

It also handles the probabilistic nature of sorting better. Traditional sorting algorithms rely on comparisons being sensible (a>b and b>c implies a>b) but you probably won't get that if you use people.

I explained it here:

https://stats.stackexchange.com/a/131270/60526

Quite closely related to matchmaking in computer games.

I remember there was a website a while ago that used pairwise comparison to rank programming languages and I think whiskey. Does anyone remember this? I could never find it again.

timhh··on Cascade: CPU fuzzing via intricate program generation
Interesting. Difficult to tell how complex the bugs are though. Some of them seem to be just triggered by accessing non-existent CSRs which suggests those chips haven't been very well verified already?

Also:

> Cascade discovered 3 inaccurate performance counter bugs (Perfcnts) in Kronos, VexRiscv and BOOM (K4, V13, B2). They incur an offset in the retired instruction counters when written by software.

Funnily enough the Sail model had this bug too! https://github.com/riscv/sail-riscv/issues/256

timhh··on Benchmarking code LLMs using user preference
This is neat. I feel like it should be blind by default though.
timhh··on Show HN: Slint – A declarative UI toolkit for embedded and desktop
I have the same but I wouldn't recommend gRPC. It adds a ton of overhead and complexity that you don't want. Plus you have to deal with all the complexity of Protobuf and the fact that everything becomes `Option<>`.

Instead I wrote my own RPC system using Serde and Bincode. Communication is over stdio, that way you can support SSH access extremely easily (like VSCode remote).

Unfortunately it was for a company so the code isn't public, but there really wasn't much code to the RPC system at all since you don't need to worry about versioning, authentication, transport, etc. I write a very simple schema language, used Nom to parse it and generate Rust, Typescript and Dart code. The Rust code generation is trivial since you can just `#[derive(Serialize...]`. Typescript / Dart was a bit more complicated (you have to implement Bincode) but it's not very difficult really.

By far the most complex thing is trying to integrate with SSH. If you call `ssh` directly then you end up having to parse non-machine-readable prompts and errors and so on. But if you use a proper SSH library then you end up having to implement ssh-agent, read `~/.ssh/config` etc. yourself which isn't fun either.

timhh··on Where are my Git UI features from the future?
> Are you judging with that experience?

Yes. Again, merge conflicts are intuitive (there were two different edits to the same code; you have to resolve the conflict). But yet again Git makes things difficult - using words like "Ours" and "Theirs" (when they're often both "ours"). It even gets them backwards in some cases!

Rebases are also pretty intuitive. For once it's not too bad a name - you take all the changes in a branch and reapply them to a new base commit. Re-base.

> Internally, it's a bunch of hashed and interlinked files. Git only collects them to show us a snapshot.

That is the snapshot. It's deduplicated but it is still a snapshot. I'm not sure what your point is here.

> Thinking that these are snapshot operations will make the operation much more hard to imagine and predict.

It really doesn't. It just means you have to describe what things do properly. Rebase calculates the changes from one diff to another and tries to apply those changes again from a different starting point. Easy no?

timhh··on Where are my Git UI features from the future?
I disagree. The model is intuitive. It's just that Git dresses it up in confusing language, a terrible CLI and a gazillion half baked GUIs. It doesn't help that lots of people recommend not using a GUI which is terrible advice for learning.

Btw I would recommend GitExtensions for learning Git. It has the most intuitive interface I've found. For example it lets you browse the files at each commit which really shows how commits are snapshots, not diffs.

timhh··on Where are my Git UI features from the future?
> It should be possible to sync all of my branches (or some subset) via merge or rebase, in a single operation.

Not sure if this is exactly what the author is looking for but I made Autorebase exactly for this.

https://github.com/Timmmm/autorebase

Great article btw. A lot of those things are way harder than they should be.

timhh··on In Praise of Stacked PRs
That sounds great! I have partly solved this issue in my autorebase tool (https://github.com/Timmmm/autorebase) - it basically rebases every branch, and fixes the commit time so that stacked branches get preserved even after a rebase just because the hashes all match properly.

That obviously doesn't work if you modify or drop any of the commits, so this option is very welcome!

timhh··on Embed is in C23
People don't dislike it because they are unaware how helpful it can be. They dislike it because they are aware how hacky, fragile and error-prone it is. They want something more robust than text substitution.
timhh··on Embed is in C23
Ha I suggested this on the C++ proposals mailing list 7 years ago:

https://groups.google.com/a/isocpp.org/g/std-proposals/c/b6n...

Enjoy the naysayers if you like! I'm glad someone spent the time and effort to push past them. Bit too late for me - I have moved on to Rust which had support for this from version 1.0.0.

> There's also the standard *nix/BSD utility "xxd".

> Seems like the niche is filled. Or, at least, if you want to claim that

> (A) XPM

> (B) incbin

> (C) "xxd -i"

> (D) various ad-hoc scripts given in http://stackoverflow.com/questions/8707183/script-tool-to-co...

>...do NOT completely fill this evolutionary niche

> This ultimately would encourage a weird sort of resource management philosophy that I think might be damaging in the long run.

> Speaking from experience, it is a tremendously bad idea to bake any resource into a binary.

> I'll point out that this is a non-issue for Qt applications that can simply use Qt's resources for this sort of business.

(Though credit to Matthew Woehlke, he did point out a solution which is basically identical to #embed)

> I find this useless specially in embedded environments since there should be some processing of the binary data anyway, either before building the application

In fairness there was a decent amount of support. But given the insane amount of negativity around an obviously useful feature I gave up.

I wonder if there was a similar response to the proposal to include `string::starts_with()`...

timhh··on Enclave: An Unpickable Lock
I don't think those high end locks are made of plastic though!
timhh··on Enclave: An Unpickable Lock
Thanks!
timhh··on Enclave: An Unpickable Lock
Yeah. In fairness someone did link a method that I think could my lock (and this one). Basically you attach a laser to it so you can very accurately tell how far the key turned. That gives you a way to test the pins even though you can't directly manipulate them.

Very very tedious though and I never tried it.

timhh··on Enclave: An Unpickable Lock
The only one I've seen was the Bowley lock. But yeah I sent my unpickable lock to him and he never mentioned it. :(

https://youtu.be/7hUonUE1hEY

timhh··on Enclave: An Unpickable Lock
Yeah same idea behind my lock too:

https://youtu.be/7hUonUE1hEY

I'd guess many many people have had this same idea.

← PreviousPage 4 of 5Next →