Show HN: Mako – a full Bitcoin implementation in C
github.com
github.com
nb: looking it up on wikipedia, it seems that monomania was a 19th century psychiatric diagnosis, and is no longer considered a real condition.
It's the most productive you can ever be imho.
I imagine a similar principle is in play here.
There's value in rewriting what already exists and works, even if it will only ever be useful to yourself.
[1] https://github.com/chjj/mako
[2] https://github.com/bitcoin/bitcoin
[3] How I calculated: find <source-directory> -type f | sed 's/.*/"&"/' | xargs wc -l
EDIT: Also of interest is this "from-scratch tour of Bitcoin in Python" that was discussed here on HN a few months ago: https://news.ycombinator.com/item?id=27593772
wc -l *.c */*.c */*/*.c */*/*/*.c */*/*/*/*.c */*/*/*/*/*.c */*/*/*/*/*/*.c bash$ shopt -s globstar
Then you can use double star syntax like this: bash$ wc **/*.c
The **/ will match any number of directory components, including zero; I think it's like *.c */*.c */*/*.c and so on, like what OP wrote.Somewhat sadly, the Glibc glob function does not have an equivalent GLOB_ option for this; it's just in Bash.
-bash: shopt: globstar: invalid shell option nameIn the git repo, there is a 2009-dated commit 3185942a5234e26ab13fa02f9c51d340cec514f8 where the material appears, as a snapshot import.
Make sure you're using "shopt -s globstar" and not "shopt -o globstar".
wc -l $(find . -type f -name "*.c") wc $(git ls-files '*.c')
This avoids counting things like generated code that is not checked in (lex.yy.c, y.tab.c or what have you), and any local test code or other junk you have laying around in the tree.So, for example, `$ makod -datadir=foobar -chain=testnet` will create a data directory in `./foobar` and start syncing testnet.
`$ mako getblock 100 2 -chain=testnet` will return block #100 to you serialized as json.
In other words, it's something akin to this:
https://man.archlinux.org/man/bitcoind.1
https://man.archlinux.org/man/community/bitcoin-cli/bitcoin-...
Please let me know how things go on Apple. I have yet to test it on a Mac. It's possible Mac has some issues since the event loop backend is using poll(2). Apple has been known to break poll every now and then.
did you leverage anything from core or other impls at all, or entirely from scratch?
One thing that kinda bothers me is that most of the tests are just
#include <stdint.h> #include <stdlib.h> #include <string.h> #include "lib/tests.h"
int main(void) { return 0; }
The past 2 months of limited sleep were a massive scramble just trying to get everything implemented and trying to get the damn thing to sync properly. Things seem to work in practice, but it's very scary not having high test coverage, especially with a project like this. Now that everything has solidified and my sleep schedule has normalized a bit, it's my intention to take the time to write proper tests.
The GMP naming convention is something like:
- Pointer/Data - single letter followed by a "p"
- Size/Length - single letter followed by an "n"
So a function declaration might look like:
static void
process_bytes(uint8_t *zp, const uint8_t *xp, size_t xn);
The above function would do some processing on `xn` bytes at `xp` and store the result in `zp` (assuming there are `xn` bytes also allocated here).A function which accepts more inputs might have `yp` and `yn` also, so conceptually: `zp = func(xp, xn, yp, yn);` or to simplify: `z = func(x, y)`.
int btc_tx_import(btc_tx_t *z, const uint8_t *xp, size_t xn);
This function deserializes a raw transaction of `xn` bytes at `xp` and stores the result in the transaction `z`. Zero is returned on failure.What would be the alternative here? I suppose I could rename `xp` to `data`, `transaction_data`, `raw_tx_data`, or something like that? I don't think it adds any value and it just takes up extra space, making the code less readable.
int btc_tx_import(btc_tx_t *transaction, const uint8_t *raw_transaction, size_t raw_transaction_size);
Or since clearly `tx` is already a convention for "transaction", it could be `tx`, `raw_tx`, and `raw_tx_size`. And sure, I have no problem with the `p` and `n` stuff, so it could be `txp`, `raw_txp`, `raw_txn`.But from your description, the input is a "raw transaction" and the output is a "transaction". Using `x` to mean "raw transaction" and `z` to mean "transaction" is obtuse. You know that the input is a "raw transaction" and the output is a "transaction", but I as a fresh reader, don't, and your code does not help me understand.
I agree here. What they want is to be able to get a superficial understanding at a glance of what the code is doing--in other words they want to give the absolute minimum effort in terms of reading. But when you actually read/write the code and understand it, the shorter names are an advantage. IMO it comes down to who the names are really important for--the reader/reviewer who will likely move on to something else in the next hour, or the person who has actually given some attention to the meaning of the code?
I think the short names also have the advantage of making the logic of the function body understandable at a glance.
int alter_struct(some_struct_t *s, void *xp, size_t xn);
then it's pretty obvious that the data pointer xp with size xn is going to modify whatever struct object we have at pointer s.The variable names can be shortened because the function interface follows a common convention.
The really important thing here is a good function name, not a good name for the data pointer argument (a descriptive name could even incorrectly suggest it has a specific type rather than simply point to bytes).
Again, it is true in every language that there are some basic patterns to what interfaces look like. That doesn't tell you anything about what the application-specific logic implementing those interfaces is intended to do. For that, it is useful to name things such that they describe what they represent in the domain of the application or library.
If it is a generic function, there are good names for that as well (for instance, some kind of copy function would still have meaningful names like "from" and "to").
But the example given by the author here was not that, it specifically acted on transactions. In that case, the input bytes were meant to represent a "raw transaction" and the output struct is meant to represent a "transaction".
Instead, the variable names should be used to convey information that the types alone can't convey.
How do you differentiate the two input lengths? If I were to rename `xp` to `x` and `xn` to `n`, what should `yn` be renamed to? At the very least, there's going to need to be a `yn` somewhere.
It's very common for code to include the type when there are two inputs to a function (even when written more verbosely): e.g. `thing_len`, and `other_thing_len`.
The `p`-suffix convention can also save you in a situation like this:
int x = 1;
int *xp = &x;
int y = 1;
int *yp = &y;
It avoids naming collisions, and further down in the function, you'll be able to differentiate the pointer and the value. I find it very useful.If you write multi-precision integer code in C[1] without this convention, you will end up with an unreadable mess. I certainly wish Torbjörn Granlund were here to testify to this.
void f(int c) {
l.c = c;
...
}
or void setColor(int color) {
label.color = color;
...
}
Names are more important than matching data types and sizes because they convey intent and meaning better than types. grep -r '\<[a-z](' .
And there were no matches. There were some one-character macros but they were macros repeated dozens of times in a localized area.It seems like what's being proposed is descriptive function names with simple variable names. If the function bodies are short enough (e.g. fit on a single screen) then this seems like a good trade-off to me. IOW, the variable names are symbolic but the function names are descriptive.
The short variable names should be clear enough if you understand the purpose of the function.
I shouldn't have to read the function implementation to understand its purpose. Code is buggy! If the function has no name or comment explaining what it's supposed to do, I only have the (often buggy) implementation to go by.
There is no reason to use terse, non-descriptive names in 2021. It's an awful practice that guarantees easy-to-avoid bugs.
I also agree that if the function's name leaves something to be desired then it should be commented.
You are conflating function names--global and relatively non-contextual--with variable names--which have limited scope and rely on the function name for their meaning.
In the setColor example, I would use setColor for the function name and c for the parameter name (with the caveat that C language doesn't have method names, so my reasoning about context has limited applicability to non-C languages)
This is quickly becoming my goto standard for measuring how clean my code is, and in my case this means ultra-descriptive variable names. I usually code in two passes: first rough things out using single-character/short names, and then go back and use LSP features to rename the variables using the language-aware tools in any modern code editor.
Important point that I didn't really get until I read more on this topic coming out of Ethereum community
One question, in src/crypto/rand.c:25 you have a standalone RNG but comments say it is not used internally? what is it used for?
I originally wanted to vendor my libtorsion code and link to it, but it felt clunky since libtorsion pulls in a ton of crypto that bitcoin doesn't need. Also, since I was focusing on just a few algorithms, it gave me the opportunity to optimize a lot of them (in particular, the ECC backend was optimized for secp256k1 whereas in libtorsion it supports all kinds of curves).
Because of all of this, there's probably some leftover comments. That comment isn't true anymore. rand.c is definitely used internally for libmako, just not libtorsion.
edit: fixed link.
But in recent years I have used Go as a replacement for C whenever possible. I think of Go as an improved C, with slightly different performance characteristics.
In your opinion, what is the advantage of writing a program like Mako in C rather than Go?
What are you using as data storage? I understand btcd uses leveldb, are you using something similar?
Mako uses LMDB. Aside from Berkeley DB, it's pretty much the only key-value store in town if you want a pure C project.
In the end, it worked out well because I really like LMDB. It has a very intuitive API, and it's very small, which sort of matches the spirit of this project.
The zero-copy on reads feature is also amazing. When reading a UTXO from the database, we do no copying and no allocation whatsoever. The UTXO is simply parsed from the pointer LMDB returns. This is something LevelDB cannot do at all.
Also, consider opening a discord server to talk more about this project!
Bcoin was frequently used as a reference along with bitcoin core v0.8.0-v0.11.0 when I felt like double checking consensus functions (among other things).
As an aside, I personally think bitcoin core v0.8.0 is the best version of core if you want to learn bitcoin from it. It's a lot more straightforward than later versions. I personally don't enjoy reading any version beyond v0.11.0.
This is also the reason mako doesn't support taproot yet. That code is very new and isn't present in upstream bcoin. I could try to implement it from the BIPs alone, but I won't know what intricacies are present in the actual bitcoin core code until I actually read it.
I spent the past few days thinking of names. Mako came to mind because I had recently been playing through the original FF7 (maybe the first time in ~10 years). It's short, memorable, looks cool. It checked all the boxes. On top of that, pretty much every name involving the words "btc", "coin", etc. is already taken when it comes to bitcoin projects.
That said, apparently there is a wayland notifier also named "mako", so I may have to think up a different name here. =/
Edit: Hmmm.
"Mako (Japanese: 魔晄, Makō; literally meaning "magic light") is a liquid substance featured throughout Final Fantasy VII. It is the condensed form of lifestream and the primary source of energy used by human beings throughout the world. It is comparable to nuclear energy. Lifestream can condense into Mako via both natural and artificial processes. The terms "Mako" and "Lifestream" are often used interchangeably because one is a derivative of the other."
From here: https://finalfantasywiki.com/wiki/Mako
The thing that catches my eye there is that they say lifestream can condense into mako through natural processes as well as artificial ones.
You don't want your own "goto fail" do you? ;) [1]
I think you did a great job.
[1] https://nakedsecurity.sophos.com/2014/02/24/anatomy-of-a-got...
> You don't want your own "goto fail" do you? ;) [1]
Ah, that was my style 7+ years ago. Then I started regularly contributing to an open source project which did not add curly braces on one-liners. I wanted to match the style of the project despite it feeling unnatural to me. Within a few weeks, my style was changed forever (for better or for worse).
> I think you did a great job.
Thank you.
if ((err = SSLHashSHA1.update(&hashCtx, &signedParams)) != 0)
{
goto fail;
}
{
goto fail; /* MISTAKE! THIS BLOCK SHOULD NOT BE HERE */
}
Copy and paste bugs can play out in any number of ways.Newer GCC will catch the goto fail by noting that the indentation is wrong.
$ gcc-11 -Wall gotofail.c
gotofail.c: In function ‘main’:
gotofail.c:5:3: warning: this ‘if’ clause does not guard... [-Wmisleading-indentation]
5 | if (argc > 42)
| ^~
gotofail.c:7:5: note: ...this statement, but the latter is misleadingly indented as if it were guarded by the ‘if’
7 | goto fail;
| ^~~~
$ cat gotofail.c
#include <stdlib.h>
int main(int argc, char **argv)
{
if (argc > 42)
goto fail;
goto fail;
return 0;
fail:
return EXIT_FAILURE;
}
Of course, it could be that the indentation is not wrong. If somneone runs the code through some automatic formatting, the problem will not then be diagnosable that way.Sadly, if I fix the indentation then:
$ gcc-11 -Wall -Wunreachable-code -O3 gotofail.c
$ # silence
even though the return 0; is not reachable.Still, I think I'm going to stick with "brace the else part if the then part is braced, and vice versa*.
Yep, it shows.
I got curious over the GP's comment, and I spot checked json.c[1] and json.h[2].
They're cleanly written, your data structures reflect your use and I can get an idea of what does what through the code alone.
Trying to figure out what could be considered obscure. For example would
if (!btc_amount_import(&x, obj->u.string.ptr))
return 0;
be considered obscure? Cause if you know the structure of a json_value, then you get what the phrase says.Well thought out and cleanly written code, I'd say.
[1] https://github.com/chjj/mako/blob/master/src/json.c
[2] https://github.com/chjj/mako/blob/master/include/mako/json.h
Basic rules are: return values on the left side of params, and functions generally return a boolean for success or failure.
The function you mention is parsing a fixed-point integer string (from a json string) and returning an int64_t. It will return 0 if it's not a syntactically valid integer, or if there is some kind of overflow, etc.
Perhaps the GP means the general crypto nature of the project lends to opaqueness. For example, I randomly clicked into https://github.com/chjj/mako/blob/master/src/crypto/chacha20... and one could ask where the magic numbers on lines 41-44 come from. Maybe it's explained in the references, or just specified without explanation as part of the protocol. I looked at the reference implementation at https://cr.yp.to/streamciphers/timings/estreambench/submissi... which has the equivalent line at L66. After finding the implementation of U8TO32_LITTLE which does some bitwise-or'ing and shifting of the first 4 items ("expa" for either sigma or tau), in Lisp I quickly verified:
(format nil "0x~x"
(logior (char-code #\e) (ash (char-code #\x) 8) (ash (char-code #\p) 16) (ash (char-code #\a) 24)))
;-> "0x61707865"
So perhaps an argument could be made that your implementation could be slightly cleaner by using those sigma/tau constants instead of the magic numbers for the first four entries in the initial state, which would show how they're related, but that doesn't really take away the magic-ness of them, just moves the question elsewhere to why those magic strings. And it wouldn't surprise me if the answer is just the strings were arbitrarily chosen; the sigma/tau naming in the reference suggests a common domain convention to me (I'm not a crypto guy though so I don't know).Still overall this is a minor point, it looks like a clean piece of code (especially given what I'm used to seeing from browsing other bits of crypto code here and there) and random sampling of other parts of the program are at least as nice and nicer, especially when they're not doing anything complicated, which I've seen plenty of C code make a mess of. While it might not always be clear why something is done, it's at least clear what's going on and where to jump for more context, what the interfaces are, and things are well-named such that I could probably go find references for some of the whys if they actually needed to be answered. e.g. in mempool.c's btc_mempool_verify, I have no clue what "Annoying process known as sigops counting" refers to, but even without that comment, the simple if condition's pieces are well-named so I could go search for more info on sigops, which I wouldn't even expect to be part of the code (at least here) anyway.
(Edit: It also occurs to me that a low ratio of comments and explanation to code could also be what is meant by obfuscation. To me that's not a large part of the cleanliness concept, though it does factor into other qualities. I consider the sqlite codebase to be some of the most beautiful C code on the planet, wonderfully documented and tested, and as a whole its beauty more than makes up for some stylistic or organizational quirks I don't exactly like. Not to say that style is totally unimportant, but neither this nor sqlite are exemplars of ugly C with lots of insane typedefs, super macros, inconsistencies, syntax abuses everywhere, and mngld_nms like vowels cost $100 a pop. For a project I consider "not so good" C code, I reference Enlightenment Foundation Libraries...)
One way to do this, however is:
#define FOURCC(a, b, c, d) ... insert shifting-oring expression here ...
Then: ctx->state[0] = FOURCC('a', 'p', 'x', 'e');
That doesn't tell you where those characters came from but at least unmasks the ASCII connection.However, the code now breaks on EBCDIC compilers, where 'a' won't have the value 0x61.
https://news.ycombinator.com/newsguidelines.html
Would you mind reviewing the Show HN rules too? Your comment broke them as well, and we're especially trying to avoid the culture of shallow putdowns when people are sharing their own work.
Here's one reference: https://msrc-blog.microsoft.com/2019/07/18/we-need-a-safer-s...
So while other software that isn't Microsoft most certainly has memory safety bugs, this blog doesn't speak for those, only Microsoft's.
The only part that Mozilla is indicated is in reference to Rust.
It is just a show and tell
And now we also have a new client on the network!
To be clear:
Libbtc is a small bitcoin library meant for everyday things: signing transactions, maintaining an SPV wallet, etc. It's good at what it does, but you wouldn't be able to actually sync and validate an entire blockchain with it.
I suppose picocoin is the most similar to mako, but I think it falls short of being a "full node" in that it doesn't provide a mempool, miner, or RPC (from what I can tell). I'm also not clear on how it is storing UTXOs. It looks incomplete right now, but maybe jgarzik could give more insight on that.
Mako is a _full_ reimplementation of bitcoin. It's an alternative to bitcoin core.