It seems like the disk usage will become untenable, but maybe someone has a plan to deal with it.
It seems like the disk usage will become untenable, but maybe someone has a plan to deal with it.
https://anduck.net/bitcoincore_vs_geth_full_node_stats.png
Ethereum's blockchain has a growth rate 2X bitcoin's. If it continues to grow at this pace, it'll eclipse 1 TB in size in 2018. If you aren't verifying the blockchain that you're reading and writing to, then you're using a trusted third party. If you're using a trusted third party, why are you even using a blockchain to begin with?
Even BitGo, a commercial payment company, has trouble maintaining datacenter servers that can keep up with the io demands of running an Ethereum node. [1] Doesn't bode well for the long-term viability of ethereum.
Bitcoin's philosophy has been, rather than broadcast every single state update to every single network participant (an unscalable network design approach[2])
[1] The Challenges of Building Ethereum Infrastructure https://medium.com/@lopp/the-challenges-of-building-ethereum...
[2] Lightning Network enables Unicast Transactions in Bitcoin. Lightning is Bitcoin’s TCP/IP stack. https://medium.com/@melik_87377/lightning-network-enables-un...
*Storage is only one part of the scaling engineering issues. Also important is the CPU and time requirements to keep a blockchain node synced.
If that's correct, then doesn't widespread usage of pruning threaten the ability of new nodes to get up to date?
I wrote a blockchain explorer on Bitcoin as a technical exercise, and the intermediate state thing caught my attention, too. If all copies of a block disappear and all you're left with is a hash, what do you have? A complete blockchain?
Archive nodes in Ethereum are almost completely superfluous.
An Ethereum archive node is not the same thing as a full node.
Archive nodes add zero extra security to the network. They are only useful for historical analysis of past states.
100% of the blockchain information is stored in non-archival full nodes.
The Ethereum system is a series of transactions that modify a state. The system is completely defined by the transactions. The latest state can be derived from them.
An archive node is a node that stores all past states for efficient retrieval.
This statistic keeps being trotted out despite having been repeatedly shown to be extremely misleading:
https://www.reddit.com/r/ethereum/comments/7gcheo/the_ethere...
In short: the intermediate states do not need to be stored for full trustless validation. Ethereum nodes now prune intermediate states by default. Bitcoin nodes don't even have an option to store intermediate states. Intermediate state storage is an entirely optional feature and comparing Ethereum nodes that do it to Bitcoin nodes that don't is disingenuous.
It's 2x slower than leveldb (if that matters to you), but given how slow a blockchain really is ...it is plenty fast enough.
https://github.com/RedCarpetUp/go-ethereum/tree/dev?files=1
I believe that it is time for developers to start thinking of single node scalability and infrastructure that is maintainable..rather than convenience.
Please feel free to comment on our code and point out some obvious flaws. We would love to learn.
Single node scalability is often overseen, but in my opinion has to come first, before you can expand scalability network wide. It's like standard computer science practices are not being applied in the crypto world, yet.
https://github.com/zsfelfoldi/go-ethereum/wiki/Light-Ethereu...
That simply isn't true. Accepting a blockchain blindly without validating is the Bitcoin SPV model. In ethereum it amounts to an undiversified risk that the two mining pools needed to constitute a majority hashpower (ethermine and f2pool) will not violate the protocol, or-- if you are accepting state with few confirmations as final-- that _any_ miner will not violate the protocol.
This is a common misperception of the benefits of a blockchain. It's like saying, if you are trusting a third party on science, why use science at all? You don't have to verify science yourself, but know that you could repeat the experiment if you wanted or needed to. The fact that the blockchain _is_ fully stored and verified by a number of distributed parties, and that you do have the ability of downloading and verifying if you want or need to, is enough to keep the system honest.
Now this statement does make sense if you are referring to your private keys. Don't ever let any third party hold those for you.
The point is that at this rate, quite quickly you won't be able to, the costs would just be too high.
How will you download and verify the chain when it hits double digit TBs?
How about 10 years from now, when the storage requirements are an order of magnitude greater than that? The rate of growth is far outstripping consumer storage device capacity...
>Even BitGo, a commercial payment company, has trouble maintaining datacenter servers that can keep up with the io demands of running an Ethereum node.
That company lost 120k of Bitfinex's BTC to a hack, the exact scenario it was apparently created to prevent. Their struggles reflect their competence level: especially the part about _directly_ using network storage, apparently genuinely not realizing that the main point of local instance storage is to function as a last cache level exactly for situations like these.
Problem solved for ~2.5 years.
Bitcoin's idea is to use Lightning Network to keep the extreme majority of transactions from ever having to hit the blockchain. And it's also why the blocksize shouldn't be increased like some forks are doing, because in 100 years (or even 5 years) the size will be unmanageable.
I believe ETH has something similar, but I'm not knowledgeable enough about that ecosystem to say either way.
It's clear to everyone that the blockchain as a "data structure" isn't scalable to anything resembling "widespread usage", and that it will be a hard requirement to figure out how to either quickly and safely "prune" information from a blockchain, or get most of the transactions to stay off of it somehow.
Also, I'm sad that most of the replies to your comment are jokes or comments about how it won't matter, or how you are dumb for even asking the question. It seems when it comes to cryptocurrencies HN is just as bad as other forums online that everyone resorts to insults and "witty" one liners.
µRaiden (two-party transfers only, no routing) is already live on Ethereum mainnet: https://raiden.network/micro.html
Kids these days. I think my head just exploded.
https://medium.com/@preethikasireddy/how-does-ethereum-work-...
Quote:
But unless a node needs to execute every transaction or easily query historical data, there’s really no need to store the entire chain. This is where the concept of a light node comes in. Instead of downloading and storing the full chain and executing all of the transactions, light nodes download only the chain of headers, from the genesis block to the current head, without executing any transactions or retrieving any associated state. Because light nodes have access to block headers, which contain hashes of three tries, they can still easily generate and receive verifiable answers about transactions, events, balances, etc.
The reason this works is because hashes in the Merkle tree propagate upward — if a malicious user attempts to swap a fake transaction into the bottom of a Merkle tree, this change will cause a change in the hash of the node above, which will change the hash of the node above that, and so on, until it eventually changes the root of the tree.
"Peter Todd, a Bitcoin Core developer, concerns what we can say is a fundamental technical argument from small blockers. They argue we can not scale on-chain, because we can not have light-client nodes, because we can not construct what is called fraud proofs."
Vitalik Buterin and Peter Todd Go Head to Head in the Crypto Culture Wars
http://www.trustnodes.com/2017/08/14/vitalik-buterin-peter-t...
<< Fraud Proofs explanation >>
How Fraud Proofs May Improve SPV Node Security in Bitcoin
https://coinjournal.net/how-fraud-proofs-may-improve-spv-nod...
Most people already have the compute, memory, storage, and network infrastructure to support that nowadays.
Multiply that by a thousand to get 1GB blocks and it will cost 5000$ a month to run a node.
Sure that's not a home Internet line, but it is a far cry from "2 or 3 corporations".
If you had said "10 thousand corporations" you'd be closer to the truth.
https://news.bitcoin.com/bitcoin-unlimited-reveals-gigablock...
The same would apply to Bitcoin scaling on chain.
https://bitcoin.org/bitcoin.pdf
Scroll to section 7 "Reclaiming Disk Space"
A Terra byte hard drive will sell for ~40 dollars these days.
In 2 decades maybe that will be a petabyte drive going for 40$.