HNHacker News
TopNewBestAskShowJobs

sehrope

1,980 karma · joined July 22, 2012

Founder of JackDB, the web-based database development tool.

Website: http://www.jackdb.com/

Email: sehrope [at] jackdb [dot] com

submissionscomments
sehrope··on App parking system shuts down in San Francisco
Does anyone know of a city that has an open API for accessing whether or not there are cars parked at it's metered spots[1]?

For cities that haven't switched to shared muni-meters, the existing parking meters seem like the natural place to add hardware that checks for availability. If the meters are network connected then at the very least they could report back whether or not they are in use or "expired". A lot of the newer meters allow for credit card payments so I'm assuming they have some kind of network connectivity.

It'd be awesome if a city provided this data to the public. Then anybody could just make an app/website[3] that shows the closest place you can park. Forget trying to make a buck off of reselling your existing spot, the savings on traffic, pollution, and frustration would justify it (less needless circling).

[1]: All spots would be even better but I'm assuming a spot that doesn't have a meter isn't going to have any hardware there that could tell whether a car is present.

[2]: http://en.wikipedia.org/wiki/Muni_Meter

[3]: Or even better integration with your smart phone's existing maps/nav app.

sehrope··on Priceline to Buy OpenTable in Deal Valued at $2.6 Billion
Besides the usual scaling they'll get from the larger org and cross marketing (e.g. I'm guessing people who book with OpenTable travel more than the average person), imagine Priceline combining its "name your own price" concept with OpenTable's inventory of restaurants to offer prix fixe meals. Pick an area of town, a type of food, headcount, and a max price and they match you with a restaurant that can accommodate.

Not sure if it's a workable idea (I honestly can't see myself ever using it) but it's an interesting angle to think about.

sehrope··on We'd lose our security certificate if we allowed pasting
Purely a guess but I think they only allow for numbers because it's a phone company. If they intend for people to enter their password/pin on their mobile phone then limiting it to only digits that you can type from any phone is understandable. Now I'm not saying it's a good idea but at least there is some sense to it.

In no particular order my usual gripes with passwords and auth in general are:

    * Disabling clipboard copy/paste (because now I can't use my password manager)
    * Length limitations (anything less than 32 characters is a a limitation)
    * Requiring punctuation or "special characters" (they're annoying and don't add real security ... just use a longer password)
    * Lack of two-factor (preferably TOTP)
The password limit is really the scariest one. Short passwords are much easier to crack. Also, I have a sinking feeling that every site that says a password must be a specific max length is storing it in plaintext. Otherwise why the heck would it matter what the length is?
sehrope··on Cakebrew: The Mac App for Homebrew
> Hence on top of command line applications. The power and flexibility are still there for people who want them, but so are the GUIs.

You don't want the GUI to be "on top of" the command line. A CLI makes for a terrible piece of code to integrate with!

What you really want is for both of them to share a common library. The CLI should be a client of the library just like the GUI. Usually the CLI and the library will be developed in tandem (so you have an endpoint to invoke a new library feature) but keeping them separate is a good idea.

sehrope··on Strong, Unique and Memorable Passwords: a Creative Approach
Yes the lack of syncing is notably missing. That's kind of the tradeoff you get for not using a central service or listing the sites themselves in plaintext (see my note about pass below). I handle it by having a separate private repo for KeepassX and syncing the repo whenever it's updated. It's an opaque blob so there's no real history but it makes it easy to keep my desktop and laptop in sync. I don't modify it too often (seriously how often do you create new accounts?) so it doesn't feel like much of a pain but maybe I'm just set it in my ways.

Pass looks interesting though the convenient way of using it (plaintext name for each site) leaks information. If you're willing to give that up then it sounds like a good idea.

sehrope··on Strong, Unique and Memorable Passwords: a Creative Approach
> The main issue I have is password that I don't use often at all. I usually can't remember them, or if I can, I cannot associate between password and website.

Yet another reason I love using a password manager[1]. Besides letting you have unique passwords for everything (which is a must), it solves the "What the heck was the password for XZY?" issue when you haven't logged into XYZ in 6 months.

People really should just use a password manager. Yes it's a pain some times (mainly using mobile) but that's just something you live with. The rest of the time though it's way better than trying to remember silly thing like "Capitalize the second letter of each word" or "Replace the last letter with a digit that denotes the number of words in the phrase"[2].

Long passwords are a solved problem and the solution is not reinventing the Caesar cipher, it's to have a single long diceware password and use a password manager for the rest.

Oh and enable two-factor auth everywhere that allows it and vote with your wallet to choose businesses that do. For example, if your bank doesn't support it, find a new bank.

[1]: I suggest KeePassX: https://www.keepassx.org/

[2]: The article suggests things like this to make sure your password unique/dictionary proof. Forget that and just use the password manager directly.

sehrope··on Nodemailer: Easy e-mail sending from your Node.js applications
Nice module. Logo is pretty nifty too. The default "well known" server list is particularly useful.

If you happen to use this with AWS watch out though. It's hard coded to US east. If you're not in that region then you should configure the server manually to a local end point so that you don't have to pay for bandwidth. Would be a tad faster too.

[1]: https://github.com/andris9/Nodemailer/blob/master/lib/wellkn...

sehrope··on Edmund Thomas Clint
This reminds me of the plot to a Mark Twain short story[1] about an artist that fakes his own death to profit from the posthumous increase in the value of his work. Been a while but I remember it being a a fun read.

[1]: Called "Is He Living or Is He Dead?" - http://www.gutenberg.org/files/3251/3251-h/3251-h.htm#link2H...

sehrope··on Third cryptocurrency exchange becomes hacking victim, loses Bitcoin
> The problem appears to have been more straight forward than any of that. The actual transaction system was a different system from the order entry/webs system. It ran orders in serial after a delay using some kind of message passing.

I'm thinking that they had both issues. If the system let you submit multiple withdraw requests (totalling more than your balance) then they definitely weren't checking at request submission before queuing them for later processing.

> Obviously you can't have a SQL transaction that spans multiple applications, so SQL transactions wouldn't really work in this case. It was just a straight forward mistake: orders could be placed before previous orders were finished executing.

XA transactions allow you to have transactions across multiple services. The most common use case is a database and a message queue. It's nowhere near as simple as dealing with a single transactional resource but it's not that uncommon in the finance world. I've used them quite a bit and they're really convenient. Getting it right from a dev-ops perspective is a bit pretty tricky though as you need to really understand how the transaction manager itself works and make sure it actually runs.

> Transactions wouldn't actually work with bitcoin in this case anyway. eg, transaction 1 checks balance, deducts amount, does BTC transfer, and then commits. Transaction 2 does the same, but the commit would fail as transaction 1 has already modified it. Transaction 2 can't actually roll back though - the bitcoin transaction cannot be rolled back.

Yes you can't have XA transactions with bitcoin as it doesn't include any concept of rollback. What you can do though is guarantee that a transaction is only run once and re-submit it if it failed to send. If you can uniquely identify a bitcoin transaction[1] then you can guarantee that you don't send it out more than once. Just add in a significant delay (couple hours? a day?) before any retries to ensure the transaction has been merged into the block chain.

[1]: And not by just using the transaction hash because we all know what can happen with that: https://en.bitcoin.it/wiki/Transaction_Malleability

sehrope··on Third cryptocurrency exchange becomes hacking victim, loses Bitcoin
> You can use an ACID compliant database and run into this problem, it's a classical race condition.

No it's not. It's just sloppy code and/or a misunderstanding of database transactions and isolation levels. If you use an ACID compliant database like PostgreSQL it's pretty straightforward to write proper code that would not have this issue.

The main concept to understand is that you need to lock all the rows that you'll be touching before you use them. In most case this method works fine with the default read-committed isolation level.

Code that queries your database in multiple round trips and then compares the values for validity on the app side will have this problem. Anything you read could have been modified by the time you "act on" the information you read. Heck even stored procedures running on the DB itself will have this problem if run in the default read-committed isolation level.

The high level solution is to use SELECT ... FOR UPDATE and lock everything that you'll be touching before you update anything.

> The solution is to implement double bookkeeping and bank reconciliation, but it's tedious and complex to implement.

This is a business requirement for any accounting/financial software but it's not a technical one. Having an audit trail, double entry accounting, and recon are business functions. You'd be crazy to design a system that didn't include them but it's not strictly required (again from a technical sense, not a business one!).

sehrope··on Valve DNS privacy flap exposes the murky world of cheat prevention
> It is faster paced games where you can't do everything server side that are the problem, first person shooters being the prime example.

First person shooters can be (and I believe usually are) done server side. Official movement, firing, health, and overall physics is handled 100% on the server. To smooth things out though, the client interpolates its view of the world.

For example lets say you fire your gun at a person right in front of you. The server would have final say of whether you actually hit the target. Rather than waiting for a round trip to the server, your client may choose to immediately show some blood/shrapnel effects to give you the illusion that you hit them. If you actually did, the server will handle the health deduction. If not, you probably wouldn't even notice.

sehrope··on PostgreSQL 9.4 – What I was hoping for
Yes my mistake on that. I'm mixing up 9.2 (where you had to do use plv8) and 9.3 (which has the native operator). The rest of the comment still stands though, a parsed representation makes the operator faster and more efficient.
sehrope··on PostgreSQL 9.4 – What I was hoping for
> JSON in postgres is just a container type isn't it?

At the moment JSON is stored in a text field so accessing a JSON field involves parsing the entire text value. If instead the JSON is stored in a parsed binary format then you can more efficiently access individual fields.

If you're only accessing a small set of known fields you can work around the issue by using plv8 to create function indexes on the fields that you'll be using. This doesn't work for arbitrary expressions though. If you want to filter a WHERE clause based on a non-indexed field the entire JSON text will need to be parsed.

Binary JSON storage makes all of this much faster with no change on the user's side. It's all transparent and just faster!

In theory it could also reduce storage space for tuples. Duplicate field names need not be repeated. I don't think it's in scope in the Postgres JSON improvements but it's a possibility.

> Why is everyone expecting it to become document storage?

Document storage in a relational database is really useful and there are plenty of use cases for it.

The standard example I use for JSON (or hstore) usage for Postgres is an audit table. You'd have all the usual audit fields (who/what/when) but you'd also have a "detail" field with event specific details. Using hstore or JSON for this is perfect. Improved support for JSON makes it much easier to provide event specific search.

sehrope··on Total security in a PostgreSQL database
Pretty good article but had to laugh when I read this:

> Common practice dictates that passwords have at least six characters and are changed frequently.

sehrope··on How I built an application in two days
> It's 2014. HA is part of the MVP when you are charging people money or providing a business-critical service (even if it's free).

Business critical and free are mutually exclusive. A free service might be business critical for a business that relies on it but that's the faulty expectation of the business, not the provider. You get what you pay for.

From the article:

> 1. After evaluating a bunch of cloud server hosts I decided to go with an Amazon EC2 Tiny Instance running Ubuntu 12.04 my reasoning behind this was influenced by price and the technologies used in the development of the web application (which will be explained in a section below) namely NodeJS and MongoDB this comes at a cost of roughly $15/mo we’ll also be needing 1 elastic IP address and a few entries in route53 + $1

I'm a big fan of AWS but this won't work. I'm assuming that the "tiny" instance is an m1.micro. If so then then it will CPU throttle at the slightest load. You can run a whole bunch on a small slice of CPU but m1.micros are a different animal. Anybody that has tried running a build server on one knows that a compile alone will hang the instance.

Again, you get what you pay for.

sehrope··on AWS Tips, Tricks, and Techniques
> Fantastic post with really helpful tips. Thanks a lot for that.

Thanks!

> I really want to read the details on the spot instance setup for a web service. Do you have a quick summary for that?

I don't have anything written up to point to but the idea is to have the spot instances dynamically register themselves with an ELB on startup. As long as they stay up (ie. spot price is below your bid) you get incredibly cheap scaling for your web service. The ELB will automatically kick out instances that get terminated as they will fail their health checks. Combine this with a couple regular (ie. non spot instances) and you get a web service that scales cheaply when spot prices are low and gradually degrades the QOS for your users when it gets more expensive.

I've got a couple other topics in the works first, but I should the write up for that one up pretty soon too. Haven't gotten around to adding an email signup form to my site yet (need to do that...) so shoot me a mail if you want to notified when the full post is up. Email is in my profile.

sehrope··on AWS Tips, Tricks, and Techniques
Yes they're a really useful feature. You can use them to allow both downloads and uploads. The latter is particularly cool for creating scalable webapps that involve users uploading large content.

Rather than having them tie up your app server resources during what could be a long/slow upload, they upload directly to S3, and then your app downloads the object from S3 for processing. If your app is hosted on EC2 you can even have the entire process be bandwidth free as user-to-S3 uploads are free and S3-to-app-server downloads to EC2 are free. Just delete the object immediately after use or use a short lived object expiration for the entire bucket.

I've made a note to add an example of this to my next post on this topic.

sehrope··on AWS Tips, Tricks, and Techniques
> Question about the underlying SaaS product. I'm not understanding how a database client on the cloud can connect to a database server on my own machine?

OP and founder of JackDB[1] (the SaaS product referred to) here.

The majority of people using JackDB connect to cloud hosted databases such as Heroku Postgres or Amazon RDS. As they're designed for access across the internet, no special configuration is necessary for them.

To connect to a database on your local machine or local network you would need to set up your firewall to allow inbound connections from our NATed IP. More network details available at [2].

[1]: http://www.jackdb.com/

[2]: http://www.jackdb.com/docs/security/networking.html

sehrope··on AWS Tips, Tricks, and Techniques
OP here. Really cool to see people enjoying the write up.

It started off as a list of misc AWS topics that I found myself repeatedly explaining to people. It seemed like a good idea to write them down.

I'm planning on listing out more in follow up posts.

sehrope··on Poll: Do you actually use the product you are working on?
We use out product, JackDB[1], everyday both for analysis and to develop the product itself further. It's a database client entirely in your web browser and we use it to run queries to analyze stats, research database features, and just general querying.

If you have the luxury of working on a product you use yourself (anybody work on developer focused solutions probably falls in this category) then I highly recommend it. Daily use of a product shows you let's you really see what parts of our product are not quite "smooth".

Just don't go down that rabbit hole too far. The features you use and how you use them don't always align with everybody else. As an example, when we first got started with it, what we considered a "low barrier" for a new user to try it out was actually quite a leap. Seeing real users struggle with that let us solve it. In your own bubble you might not notice things like that.

[1]: http://www.jackdb.com/

sehrope··on Rethinking the limits on relational databases
I may be in the minority, but I rather like the rigidity of a fixed schema. Yes, it's a bit more "painful" for initial setup (you have to actually create the tables) and yes, it's a bit more work during migrations (you have to actually add the columns), but I don't see either of those as getting in the way enough to give it up. It's just too useful.

Data structures are meant to last much longer than application code. Anyone that has worked on a long running system can attest to that; well defined data structures and table layouts will outlive any application code.

When I'm designing a system, I think I spend orders of magnitude more time thinking about data structures then actually implementing the CREATE/ALTER TABLE code for them. Planned properly, you can even do the ALTERs/CREATEs necessary to add columns in advance of any actual app usage (ie. "two stage" app deployment).

There is a place for "flex fields" or storing generic "documents" but legit use cases are pretty rare. When they are necessary, a single JSON column is usually enough. The example I generally use is an audit trail: The who (FK to user), what (event enum), and when (timestamp) are all strongly typed but you may want a JSON field for event specific data.

Oh and if anybody has every tried to do a data migration with a schema-less database ... well have fun with that. Either you bite the bullet and convert everything or you end up with a lot if/then/else logic littered through your app that will bite you down the road.

sehrope··on Use SQL subqueries to count distinct 50x faster
If the author is reading please post the EXPLAIN ANALYZE results for these queries. I'd much prefer the raw text psql output to the GUI representation in pgAdmin. More importantly it has the actual numbers for the intermediate steps to see which is the culprit.

Also, any idea how many times each variation was executed? 10M rows is not that big. At 1K per row (which is probably an overestimate) that comes out to 100 MB. That easily fits in memory. After the first full scan all of it would come from the cache. It's not a fair comparison unless you flush the cache between tests.

sehrope··on How I went from 100 to 0 things (or how I was robbed of all my stuff)
> Oh yes, being somewhere other than your house has very clear upsides, but an appropriate location is not always forthcoming.

It's not too hard to find one. Unless you're completely anti-social you probably have at least one tech-savvy friend that can understand the need for this kind of setup. Even better if you have more than one friend (hopefully not too be an "if") then you can have a "round robin" approach with a group of friends. An open source (so the crypto can actually be vetted) version of BTSync[1] would be great for this.

> For example, I would consider it pretty poor form to plug in a personal networked backup box at my desk at work. That kind of move can also pose a risk to my sustained employment!

Haha. Yes plugging in random networked boxes at the office might arouse some (just!) concern. When I wrote that piece I was thinking specifically of my company as I'm the boss :D

[1]: http://www.bittorrent.com/sync

sehrope··on How I went from 100 to 0 things (or how I was robbed of all my stuff)
> Depending on your threat model, they don't even have to be outside the house. If petty burglary is what you are defending against, a disk in a quiet corner of your basement is probably plenty.

Being outside of the house protects it equally well against fires/floods/earthquakes/pets too. Protection against burglary is an added bonus.

sehrope··on How I went from 100 to 0 things (or how I was robbed of all my stuff)
Being robbed sucks but when it comes to digital possessions there's no reason it needs to suck this much.

> I didn’t really trust file encryption because I thought I might lose files because of it and therefore I never enabled Mac OSX’s built-in FileVault hard drive encryption. I should have though. It’d save me from worrying about who’s going through all my files now.

This is a no brainer. I have yet to notice any real performance hit for enabling full disk encryption. Just enable it, make sure to have a long/strong password, and make sure your computer actually locks when you close the lid.

You should never be worried about losing files on a single computer. If they're important then they should be backed up to multiple computers/drives/services. If you're worried about accidentally wiping your laptop when you setup FDE then just make a backup before hand.

> My backup drive was literally NEXT to my MacBook. By sheer luck, I had just backed up my internal drive the day before and they didn’t take it.

Offsite backups are a must. It can be your own "offsite" (ie. a server at friends/parents/office) but it needs to be somewhere other than the primary site.

> I didn’t have a cloud backup because I don’t trust a third party with my data.

There's nothing wrong with not trusting third parties but that's exactly what encryption is for. Encrypt your data locally and then you can store it remotely without worrying about it being accessible to a third party. DIY scripting with GPG/S3 works well for a lot of situations. Or you can just use Tarsnap[1].

Honestly it makes a lot of sense to do the same with USB drives as well. My Linux machine is my primary computer (OS X laptop when roaming...) so the majority of my backup USB drive usage is done there. I have them setup with LUKS/dm-crypt[2] for full disk encryption. It's really easy to setup, plug-n-play on modern systems, and it almost falls into the "no reason not too" category. I just wish OS X supported it too.

[1]: http://www.tarsnap.com/

[2]: http://en.wikipedia.org/wiki/Linux_Unified_Key_Setup

sehrope··on Ask HN: Password best practices?
> Use a password archive (e.g. KeePassX) encrypted with one really, REALLY good password you'll memorize (plus a keyfile, should you feel paranoid).

One more vote from me for KeepassX with both a long, unique passphrase and a keyfile. Once you start using a password manager it's impossible to go back to trying to memorize individual passwords (which is a good thing) and you'll be hooked for life at the convenience.

One tip: make sure to increase the number of rounds of key derivation for the master password (should be under "Database settings"). The default is relatively low. There's a handy button with a clock on it that will perf test your computer and pick a number of rounds that is approximately 2 seconds of CPU time.

> For default, I have chosen 60-character password length, upper-lower-numbers-spaces-special characters mix (as a rule of thumb, a longer password is preferable to a more complex password).

Yep when it comes to passwords, long beats strong.

For answers to "secret questions" (which arguably are one of the stupidest concepts in account security) I follow the same rule. Long, random gibberish answers. Ex:

    Q: What's your first pet's name?
    A: gVr5nlLy0BLPu6OpW5YUPVMtHXqy5xxr2UMUbCoSTIR3KmPTfdl8Ml8ckfbDC7Pw
    
    Q: Where did you go to high school?
    A: Wdn3hohFxq2H4f7y8n501gwlPGZLlKDIDusWA0dGm9NEaabwogh9DUKngDy2zabC
I think it'll be pretty funny if I ever have to read one of those over the phone.

>One of the worst offenders seems to be Skype, with its "maximum password length: 20 characters, no weird characters allowed."

Yes that's pretty stupid and it's not just Skype, it's also Hotmail/Outlook and "Microsoft accounts" in general (ex: Windows Azure). I don't remember if it was always like that for Skype or if it happened after the Microsoft acquisition though.

sehrope··on What Hard Drive Should I Buy?
> and in some cases...externals that were made internal! :)

I remember reading about this a while back[1] (fun read!).

Have you guys noticed a big difference in life expectancy of the repurposed external drives or do they generally match up with their internal equivalents?

[1]: http://blog.backblaze.com/2012/10/09/backblaze_drive_farming...

sehrope··on How To: Hosting with Amazon S3, CloudFront and Route 53
You're absolutely right about being able to MITM the HTTP piece and replace the content. That's true for any mixed content site. In this case though I disagree that having the HTTPS link to S3 is entirely useless. It's used specifically for an SSL link to download our GPG key, that additionally is available on a number of key servers and indexed by search engines like that too[1]. In that usage it's one of many ways of getting that key and, like all GPG keys, should really be verified before use anyway. For just about anything else though I agree that mixed content is a very bad idea.

[1]: https://www.google.com/search?q=jackdb+gpg

sehrope··on How To: Hosting with Amazon S3, CloudFront and Route 53
The same reason you'd use it for a non-static site; to ensure that visitors to your site are getting your actual site content and not something else(i.e. avoid a MITM).

If you had a purely informational site and it listed phone number, address, or heck even a bitcoin address, wouldn't you want to make sure that your visitors got the actual site and not something malicious?

There was a story last week about a guy who's ISP was inserting content into web pages. I can't find the link as I think it got lost as a result of the HN server crash. SSL prevents crap like that.

sehrope··on How To: Hosting with Amazon S3, CloudFront and Route 53
Nice write up. We have a very similar setup (Jekyll generated static site + S3) for our website[1] and reading through this is kind of nostalgic of getting it set up (and a friendly reminder to go back and gzip some of CSS files).

The biggest plus of this setup is that once it's deployed you don't think about it. It just works and you never think about scaling. Oh and it's cheap (seriously it's like peanuts a month as all you pay for is bandwidth at $.10/GB).

The biggest negative is getting SSL. CloudFront supports it but it's expensive ($600/mo see [2]). Compare that to the pennies it costs to host the non-HTTP site on S3. In our case our cloud app is on completely separate domain (SSL-only) and our public site is informational only so the trade off works. The only SSL enabled link on our public site is for our contact GPG key and it's linked directly to the HTTPS S3 URL.

[1]: http://www.jackdb.com/

[2]: http://aws.amazon.com/cloudfront/pricing/

← PreviousPage 2 of 9Next →