Baidu File System – A distributed file system for real-time applications
github.com
github.com
Looks like a pretty good first attempt at a distributed filesystem. Initial impression is HDFS with a distributed NameNode/Nameserver. The first diagram also shows a Metaserver layer that's not mentioned at all in the more recent of the two design docs but "separate Metaerver from Nameserver" appears (unchecked) in the roadmap. All operations using access methods other than their own SDK seem to get funneled through the NameServer cluster, which will severely limit throughput. Not clear how they do replication, though weakly implied that it's driven from the client (like Gluster) or NameServer rather than the first ChunkServer (like Ceph, HDFS, everything else). No mention of how they handle consistency or repair. Likewise no information about performance or security. Not clear if it's anywhere near POSIX compliant (probably not).
FUSE support is in the diagrams, but not checked off on the roadmap. Slow-node detection and avoidance seemed like one of the most interesting features from the design, but is not checked off either. Other things not even on the roadmap, using Gluster not as a fair comparison but as a handy list of possibilities: multiple replication levels, tiering, erasure coding, NFS/SMB, caching, quota, snapshots.
As I said, looks like a good first attempt. Better than most I've seen, with lots of potential, but as of today it seems rather bare-bones. Many hard problems remain to be solved, and I wish them well.
(It is a pity that Chrome doesn't automatically translate github pages that contain different languages - not sure why that isn't happening.)
I think a low read/write latency dfs suitable for real time applications would be a game changer. I'm hoping they up the documentation from here and engage the English speaking community.
Untrue. Gluster is already deployed that way in many places. Yes, in production and at scale.
http://blog.xebia.com/persistence-with-docker-containers-tea...
Or do you know of clean, container only (no plugins or special external tools) solution ?
There is another distributed file system that support full POSIX semantics with well tuned FUSE client, called MooseFS [1].
My ex-employer used that in production for about 8 years, the biggest cluster has more than 2PB.
Disclosure: I'm a MooseFS fan and contributor :)
It all looks reasonable, but this takes self-documenting to an extreme.
That design is very reminiscent of CephFS's cluster monitors, metadata servers, object storage devices.
[1]: https://github.com/baidu/bfs/blob/master/docs/design.md, https://github.com/baidu/bfs/blob/master/docs/BFS_design.md
Where I think he did well was in de-escalating tension.
Sadly, in the past years, due to some political movements, the term "master/slave" has been declared problematic, and GitHub actively warns that projects using such language can and will be excluded from the service.
There have been previous discussions about this on HN.
There was actually a huge debate about this on Reddit caused by Swift merging a rename change PR into master. The Swift team was so excited about the change for some reason that they didn't even run tests before the merge...
Rename of what?
Here's my comment from the thread:
https://reddit.com/r/ProgrammerHumor/comments/3veu2t/comment...
The rest of the discussion is a good read too.
Haven't looked at any code, but what you describe is very common usage.
Is there any file system that would be able to sync what amounts to text/binary data across many hosts and allow me to aggregate the data off the network for more secure storage?
I was thinking about using IPFS for this but this also seems better. I'd hopefully like to have a private network for this use case so that other people can't post up a device on this file system and introduce fake data.
Using Raft over a dependency on an external consensus system is nice. Definitely makes the namenode architecture much better.
Another important aspect is that is using SSD + SATA(I suppose) , which could be a better option than standard SATA/SSD or LV cache using SATA + SSD.
Even if it's just a new thing, if it proves to be faster it may be implemented in Hadoop ecosystem in the future. HDFS has a lot of features being a mature piece of software but it lacks on the response time.
This is non sequitur. The conclusion does not follow from the premise.
Also, a C++ implementation is likelier to use far less memory than a Java implementation, assuming the skills of both programmers are roughly equal.
As for using less memory - you don't allocate buffers for file data on the JVM heap. You allocate them in native memory exactly as you'd do it in C++. Therefore it is possible to create a JVM-based file system that handles petabytes of data with just as little as 100 MB heap, used mostly for small temporary objects.
Also, the code here is using mutexes a lot to synchronize threads and lock out whole objects. Therefore I think these "realtime" claims are quite exaggerated.
Also if GC was such a huge problem, exchanges or HFT companies wouldn't use Java for their low latency stuff, and there definitely are companies which do.
Can you name one?
I meant the code size and heap allocations for data structures, not file buffers.
And 100MB is huge compared to many C++ programs. And that's on top of the Java runtime!
The fact that something is in C++ doesn't make it automatically efficient. And particularly, if we're talking about milliseconds, not nanoseconds here, in Java or C# you can do just everything what you can do in C++, performance-wise.
Eg. lylei changed the title from "cs启动太慢" to "cs start is too slow " ( on https://github.com/baidu/bfs/issues/376 )
lylei changed the title from "其他SDK写策略" to "SDK writing strategies(fan-out write for example)" (on https://github.com/baidu/bfs/issues/243 )
Also wonder if there will be larger skepticism toward integrating Chinese O/S in regards to potential influence by the government (like the NSA has tried to influence in the past)
"Once your code has passed the code-review and merged, it will be run on thousands of servers"
And the Chinese text below says tens of thousands of servers, which is it? :-)
There are several other discrepancies in the doc between the Chinese version and the English one. Some technical proofreading is needed.
For instance, Hindi has special names for 100000 (lakh), 10M (crore/karod) etc. so a similar translation to Hindi would use those even if it meant introducing a factor of 10 in the literal interpretation.
It's another point of failure, more infrastructure you have to keep alive...
std::string listen_addr = std::string("0.0.0.0") + server_addr.substr(server_addr.rfind(':'));"AFS volumes can be replicated to read-only cloned copies."
Seems (have not tried it yet) awesome because: - another big party offering such software, the more choices the merrier for the users/sysadmins - sandboxed - scalable to 10k nodes - no single point failure - ssd and traditional disk usage via the disk manager
Won't it be good if code has at least some documentation in English ? It's not that they don't want, they have some part in English.
- Chinese
- Spanish
- English
You mean in tech?
Do you want to learn 5 different languages just because 5 different kind of people write good software ?
It's easy to blame others.
while at the same time so do those complaining about English's hegemony in India (by furious Indians).
Strange is our world.
I feel like we should just make up random strings when we name things...
India is atleast a 100-200 years away from something of this kind happening; or more likely never at all.
Indeed, there is not a single research university worth its name that isn't also essentially an export hub of brains to the 5-eyes (& Singapore).
English is crucial for India's system of feudal slavery to work.
So as a team leader or project manager in China I would probably also stick with Chinese since it is much easier to find really good and not too expensive employees that way. Let them try to use English, support the ambition, but don't enforce it.
And I don't know much about India but from what I heard is that it is more a cultural issue that India lacks behind. Everything is (so I heard) still very traditional and backward focussed. While China as a country spent 20-30 years to become more open for new ideas and approaches. How true is that from other people's perspective here?
http://sankrant.org/2011/03/the-english-class-system-2/
There are systematic faults, which prevent much change, if the current policies are kept up (note: India's literacy rate is ~78 %, since literacy (except in English) brings no great advantage).
http://www.nytimes.com/2015/03/22/opinion/sunday/how-english...
http://www.forbes.com/sites/realspin/2014/11/06/the-problem-...
Imagine China, with only the expensive class of engineers, for instance. Or atleast, one with this class, and another class that was educated in English, but barely knows the language, let alone possessed of any usable skill.
There are now villages, driven by this economics, where rural-children are being taught in English. Considering how bad the Japanese/Chinese are with English, it shouldn't be hard to interpolate how disabling this is when everything is being taught in a foreign language.
What do you mean this statement? Do you even know GlusterFS (later became RedHat Storage) developed from India.
First try to understand the context before starting your racist rants.
Disclaimer : I'm Ex-GlusterFS dev.
English is the most common script in India, why would we use anything else. Unless you are Hindian trying to impose your minority language on the rest of us.
- English is hardly the "most common script". Just because the retainers in Delhi impose the colonial apparatus on us, doesn't automagically give it "statistical power" as well.
- Every state has a (poor, uneducated, illiterate) captive linguistic population more than that of Korea; no reason they ought to use Hindi, nor even Nagari (script != language, in case your education didn't tell you that).
- Among the rich, yes, English is most common, and this is really what matters in the end, aint it ?
This is precisely why India will never be able to work in its own language, and also precisely why it is doomed to eternal poverty and continued illiteracy. Probably will remain a hub exporting little other than people, for the next couple centuries.
(See: https://youtu.be/SJx0KFtm9Rw?t=21m56s)
And this is why I admire China. It's not democratic, it has a paranoid regime, but at least they aren't run by hypocrites who'd use a "socialist democracy" as cover for continued colonization.
> trying to impose your minority language on the rest of us.
Your skills in generating irony amuse me.
For better or worse, English is the common language for software development (and aviation, etc.)
You don't say ?
> For better or worse, English is the common language for software development (and aviation, etc.)
English in India is more dense than the feudal castes of medieval Europe; hardly a professional thing this.
The Human-rights wallahs don't complain precisely because the current state of affairs benefits the nations that control them; much as it does their native retainers.
std::string pad;
if (path[path.size() - 1] != '/') {
pad = "/";
}
Else?How would that be consistent with your own classes? "Oh no, you can't just use a 'Tree' object, you need to explicitly set that there are no leaves yet, no branches yet, no squirrels yet, etc… etc…"
Do you .clear() your vectors before you use them?
This sounds like newbies that do: #define TRUE (1 == 1)
Anyway the reason for which it would no pass review is that today you use one compiler, tomorrow you have to use another and then you have to review all these little details again. It's about saving money more than anything and you do that by not relying on compiler behavior.
If a different compiler breaks this behavior, it's not standard compliant and thus could do all sorts of stuff in every possible line, including in:
std::string pad = "";The '= ""' won't help you.
std::string pad = descriptive_name_here(path);
with the added bonus of being able to add "const" to that, for the benefit of the reader.
This is not relying on compiler implementation! Can you name one language that has strings that initialise to anything but a valid object containing an empty string?
This is not an obscure side-effect. This is like assuming "std::vector<int> v;" creates an empty vector, not a undefined-state vector container.
(I don't want someone coding C++ as if all objects are references. Coding in one language as if it were another is a well-known antipattern)
I ran into this article that puts quite nicely why the problem isn't the "else", but the "if" itself:
https://medium.com/@bartobri/applying-the-linus-tarvolds-goo...
This is what I meant in the other comment by preferring the non-branching.
Good thing then that it's mandated by the language reference, and not up to the compiler to decide. According to C++11, §21.4.2/1, an uninitialized std::string should be an object of class std::basic_string with non-null data and a size of 0.
It's clever, yes. The bad kind of clever that's also misguided.
That would break approximately ALL C code.
You're being ridiculous. You might as well try to protect against the meaning of "if" changing.
I've seen amateur code that tries to protect against "stdio.h" going away and therefore reimplementing everything in it. This is like that.
Believing that the meaning of everything can change means that you cannot use anything you didn't code yourself. You can't trust documented APIs, then that's some sort of programmer NIH nihilist.