259 karma · joined March 1, 2011
Email me at alex.sharp@orionvm.com.au
My first take on it was the following: Take the average of the differences between each pixel and the pixel adjacent to it over each column and store that as, say, X[]. Find some maximum number such that the sum of every nth column of X - the sum of all other rows of X is maximised. That's the column width. Split into a series of columns, use stable matching to match columns based on the sum of differences between rightmost male pixel and leftmost female pixel over all rows. Use stable marrage to give you a partial ordering, turn into a complete ordering and unscramble the image.
Any better/more elegant solutions?
Besides, search index crawling is comparitively cheap compared to the data processing required to make it useful for actual search.
Truecrypt would also work. Does dropbox do binary diff syncs?
How do you create the "folders" that people see? (ui wise?) And hide the encrypted counterparts?
At the moment, owning dedicated servers and coloing them yourself is orders of magnitude cheaper then a cloud provider. It's only "cheaper" if your cost of money is stupidly high, ie. you're a startup, or other factors dominate your TCO, ie. you're a startup.
What is interesting is that supply side there is no reason for this to be so. Furthermore there are reasons to expect that the equilibrium price for cloud to be below the cost of dedicated hardware/colocation.
Also in your "how it works" s/loosing/losing.
Keep in mind that you are breaking dedup, which is likely to make dropbox sad. Probably not sad enough for them to do anything useful though.
Remember that you can take node failure, as long as you have a copy of your state stored somewhere that you can restart from when required. The overall product isn't particularly effected from that.
0mq is probably what you want if you're implementing this the way i'm expecting you to. ie, some variation on: take current state, multicast that to all of the machines doing the optimisation take results, unicast back to a machine to determine the best state to seed the next round, repeat.
0mq over infiniband will iirc use the native infiniband multicast groups which is very efficient.
Keep snapshots of each iteration, and all should be good.
But like all things, it depends entirely on the specifics of what you're trying to do. Send me an email, you seem to have an interesting problem.
Yes, it is OOM $100/mbit to get international bw. But that's nowhere near $1/gig.
IMHO providers end up charging that much 1) Because they can and 2) To help offset a very high cost base due to a lack of automation and inefficient architectures. A million dollar san really doesn't pay for itself now, does it?
Telstra's cloud is rather sad. Manual day long provisions really arn't a cloud by any reasonable strech of the imagination.
That being said, not all of the new clouds are, well, backwards. Take a look at OrionVM (disclaimer, I work there). They have vm's that provision in seconds, that are HA under hardware failure, that have real distributed disks that are faster then comparable hardware disks, that have real private layer 2 networks, that have a remote out of band serial console, that has a damn pretty UI, ....
But at the end of the day, the market isn't there for cloud. US companies love us. US VC's instantly get us and our tech. But when we try and market here in aus, we've found that nobody here actually wants a cloud. What they want is a dedicated server without the hardware.
Which is depressing, if you think about it. Profitable, but depressing.
What much warm data data do you normally have/node?
Why don't you guys use a proper distributed database to handle container mappings/etc?
This is cheap and at qty 1.
A backblaze type box is ca 12K for 135TB of storage.
Assume an interest rate of 5% and 36 months worth of repayments and the server itself is worth $725/month
It's uses roughly 1kw of power and 4u of rack space, so say you have 6/rack with a 30A rack. You can get the rack for say 5k, giving us a total rack cost/server of $833/month
Total cost/server month is $1558/month.
Total cost/gb month is $0.011/gb month.
Add in parity replication (1 in 4, 25%), $0.014/gb month.
This doesn't include compression or dedup, both of which drops cost price dramatically.
Compare that to say S3's $0.14/gb and you can see why I'd say the margins are stupid, especially at the scale they're running at.
How are you calculating your latency? Also, what distribution do you assume your file accesses will come from?
There are many reasons why. First is that by splitting things into small blocks spread around the cluster you have more consistent load (why is left as an exercise to the reader), you can more easily read ahead from later blocks, etc.
You're going to shard a SQLite database into a series of objects to deal with "large" containers?
Just as an example, they store object listings using SQLite databases that were file based replicated between nodes for HA. Thus when you had too many files in one container your performance would sink like a stone. Assuming it was never corrupted/etc...
I'm all for people working in this space though. A monoculture is rarely good for anybody.
I'm unsure if anybody has a large scale RADOS based blob store though. It would be interesting to see how it holds up.
And yes, I know full well what it means :)
I've had a quick look around the website, and the most information i could sort of squeeze out was from the blog.
The arch page https://nimbus.io/architecture/ is devoid of architecture.
I don't see any mention of compression or dedup.
I don't see any mention of network level failover/redundency.
I don't see any mention of high level CNC functionality/db arch (ie. swift's notorious file replicated SQLite database....)
I don't see a download source button. Is that just me?
Overall, sounds very interesting and rather promising. Who are the people behind Nimbus.io?
The purpose of the salt is to defeat time/space tradeoff attacks by inflating the required space to the point of impracticality. ie. 20 bits of salt will increase the size of the rainbow table required a million times.