HNHacker News
TopNewBestAskShowJobs

asharp

259 karma · joined March 1, 2011

Director/lead dev at OrionVM, the fastest cloud computing platform on the planet (For storage, at least).

Email me at alex.sharp@orionvm.com.au

submissionscomments
asharp··on Instagram Engineering Challenge: The Unshredder
Now that, my friend, seems to be an actual challenge.
asharp··on Instagram Engineering Challenge: The Unshredder
Warning, spoiler?

My first take on it was the following: Take the average of the differences between each pixel and the pixel adjacent to it over each column and store that as, say, X[]. Find some maximum number such that the sum of every nth column of X - the sum of all other rows of X is maximised. That's the column width. Split into a series of columns, use stable matching to match columns based on the sum of differences between rightmost male pixel and leftmost female pixel over all rows. Use stable marrage to give you a partial ordering, turn into a complete ordering and unscramble the image.

Any better/more elegant solutions?

asharp··on Bechtolsheim: AWS, open source rewrite rules for startups
The point will probably be a few years off, at least for retail prices. Cloudsmithing is still a new art, and there are very few people outside of Amazon who can produce a functional cloud.
asharp··on Why CommonCrawl is a Disruptive Force in Big Data
Wasn't the major problem with commoncrawl that most of the index data was too old to be useful?

Besides, search index crawling is comparitively cheap compared to the data processing required to make it useful for actual search.

asharp··on Show HN: Safebox - for ur Dropbox
encfs also works really well with dropbox and it similarly includes many more options.

Truecrypt would also work. Does dropbox do binary diff syncs?

asharp··on Show HN: Safebox - for ur Dropbox
Cool.

How do you create the "folders" that people see? (ui wise?) And hide the encrypted counterparts?

asharp··on How To Log Bash History to Syslog
Interesting, but is there anything that can easily process syslog on the other end? ie. split data like this out into something useful?
asharp··on Bechtolsheim: AWS, open source rewrite rules for startups
It's interesting that Netflix mentions the cost of infrustructure.

At the moment, owning dedicated servers and coloing them yourself is orders of magnitude cheaper then a cloud provider. It's only "cheaper" if your cost of money is stupidly high, ie. you're a startup, or other factors dominate your TCO, ie. you're a startup.

What is interesting is that supply side there is no reason for this to be so. Furthermore there are reasons to expect that the equilibrium price for cloud to be below the cost of dedicated hardware/colocation.

asharp··on Show HN: Safebox - for ur Dropbox
Interesting. What crypto functions do you use? Openssl?

Also in your "how it works" s/loosing/losing.

Keep in mind that you are breaking dedup, which is likely to make dropbox sad. Probably not sad enough for them to do anything useful though.

asharp··on Researchers show how to break quantum cryptography
Also, I believe a good portion of the cryptosystems around DES's time were intentionally crippled to meet export restrictions so that they were not classed as munitions. So it would not be so much to say that the cyphers were weak insomuch as they were defective by design.
asharp··on Ask HN: Which meassage passing system for distributed computing?
Distributed optimisation is interesting.

Remember that you can take node failure, as long as you have a copy of your state stored somewhere that you can restart from when required. The overall product isn't particularly effected from that.

0mq is probably what you want if you're implementing this the way i'm expecting you to. ie, some variation on: take current state, multicast that to all of the machines doing the optimisation take results, unicast back to a machine to determine the best state to seed the next round, repeat.

0mq over infiniband will iirc use the native infiniband multicast groups which is very efficient.

Keep snapshots of each iteration, and all should be good.

But like all things, it depends entirely on the specifics of what you're trying to do. Send me an email, you seem to have an interesting problem.

asharp··on Ask HN: Help with writing a job-seeking article for sydney
If you're interested in getting in touch with the startup scene, send me an email.
asharp··on Researchers show how to break quantum cryptography
Very interesting, i'd like to see what physics hacks they end up using to patch this.
asharp··on Researchers show how to break quantum cryptography
Actually you would worry more about incorrectly implemented padding or a side channel attack or something similarly stupid destroying your cryptosystem. Sad, but true.
asharp··on New AWS region: US West (Oregon)
Bandwidth prices here are rather stupid.

Yes, it is OOM $100/mbit to get international bw. But that's nowhere near $1/gig.

IMHO providers end up charging that much 1) Because they can and 2) To help offset a very high cost base due to a lack of automation and inefficient architectures. A million dollar san really doesn't pay for itself now, does it?

asharp··on New AWS region: US West (Oregon)
Shipping data offshore has real legal implications that are anything but FUD. It's actually rather complex to unravel who has juristiction over what.

Telstra's cloud is rather sad. Manual day long provisions really arn't a cloud by any reasonable strech of the imagination.

That being said, not all of the new clouds are, well, backwards. Take a look at OrionVM (disclaimer, I work there). They have vm's that provision in seconds, that are HA under hardware failure, that have real distributed disks that are faster then comparable hardware disks, that have real private layer 2 networks, that have a remote out of band serial console, that has a damn pretty UI, ....

But at the end of the day, the market isn't there for cloud. US companies love us. US VC's instantly get us and our tech. But when we try and market here in aus, we've found that nobody here actually wants a cloud. What they want is a dedicated server without the hardware.

Which is depressing, if you think about it. Profitable, but depressing.

asharp··on Nimbus.io: Open-source alternative to Amazon S3
I'd suspect that their data temperature is very bimodal, so they'd be able to easily split out hot data from cold.

What much warm data data do you normally have/node?

asharp··on Nimbus.io: Open-source alternative to Amazon S3
But why SQlite? And why file based?

Why don't you guys use a proper distributed database to handle container mappings/etc?

asharp··on Nimbus.io: Open-source alternative to Amazon S3
Hmmmm. Downvotes without comments. Classy.

This is cheap and at qty 1.

A backblaze type box is ca 12K for 135TB of storage.

Assume an interest rate of 5% and 36 months worth of repayments and the server itself is worth $725/month

It's uses roughly 1kw of power and 4u of rack space, so say you have 6/rack with a 30A rack. You can get the rack for say 5k, giving us a total rack cost/server of $833/month

Total cost/server month is $1558/month.

Total cost/gb month is $0.011/gb month.

Add in parity replication (1 in 4, 25%), $0.014/gb month.

This doesn't include compression or dedup, both of which drops cost price dramatically.

Compare that to say S3's $0.14/gb and you can see why I'd say the margins are stupid, especially at the scale they're running at.

asharp··on Nimbus.io: Open-source alternative to Amazon S3
Latency may be able to be made low through caching, but depending on the distribution the point at which additional cache is uneconomical may be well before the edge of your performance envelope.

How are you calculating your latency? Also, what distribution do you assume your file accesses will come from?

asharp··on Nimbus.io: Open-source alternative to Amazon S3
Keep in mind that S3, like all other Amazon products, are priced with stupid margins. As such providing lower prices isn't difficult.
asharp··on Nimbus.io: Open-source alternative to Amazon S3
Pretty much, it's always faster to read from multiple disks.

There are many reasons why. First is that by splitting things into small blocks spread around the cluster you have more consistent load (why is left as an exercise to the reader), you can more easily read ahead from later blocks, etc.

asharp··on Nimbus.io: Open-source alternative to Amazon S3
:(

You're going to shard a SQLite database into a series of objects to deal with "large" containers?

asharp··on Nimbus.io: Open-source alternative to Amazon S3
A cloud hosting provider.
asharp··on Nimbus.io: Open-source alternative to Amazon S3
Swift is....... woeful.

Just as an example, they store object listings using SQLite databases that were file based replicated between nodes for HA. Thus when you had too many files in one container your performance would sink like a stone. Assuming it was never corrupted/etc...

I'm all for people working in this space though. A monoculture is rarely good for anybody.

asharp··on Nimbus.io: Open-source alternative to Amazon S3
Interesting. If they only open source their client side libraries it would be a rather sad development.
asharp··on Nimbus.io: Open-source alternative to Amazon S3
Ceph is a very interesting project. RADOS, their distributed block store is now mainline I believe and the project is coming along in leaps and bounds.

I'm unsure if anybody has a large scale RADOS based blob store though. It would be interesting to see how it holds up.

asharp··on Nimbus.io: Open-source alternative to Amazon S3
I think somebody is making unfounded assumptions about the mental state of another person. A common, albeit embarrassing mistake to make.

And yes, I know full well what it means :)

asharp··on Nimbus.io: Open-source alternative to Amazon S3
So it stores data using RS encoding on multiple pieces.

I've had a quick look around the website, and the most information i could sort of squeeze out was from the blog.

The arch page https://nimbus.io/architecture/ is devoid of architecture.

I don't see any mention of compression or dedup.

I don't see any mention of network level failover/redundency.

I don't see any mention of high level CNC functionality/db arch (ie. swift's notorious file replicated SQLite database....)

I don't see a download source button. Is that just me?

Overall, sounds very interesting and rather promising. Who are the people behind Nimbus.io?

asharp··on I wrote BozoCrack to show why plain MD5 is a horrible way to hash passwords.
How does it defeat the purpose of the salt?

The purpose of the salt is to defeat time/space tradeoff attacks by inflating the required space to the point of impracticality. ie. 20 bits of salt will increase the size of the rainbow table required a million times.

← PreviousPage 3 of 8Next →