HNHacker News
TopNewBestAskShowJobs

samdk

2,399 karma · joined November 29, 2009

very outdated website: http://samdk.com

email name: sam, domain: defabbiakane.com

submissionscomments
samdk··on Day Against DRM

    You can't unshift a paradigm like DRM for big media
Yes, you can.

DRM was pervasive in the music industry. It no longer is. DRM was pervasive in the ebook industry. We are beginning to see signs from publishers (Tor books being the most recent example) of a willingness to sell ebooks without DRM. "DRM-free" has become a selling point for smaller companies in the games industry.

It is not easy, and it is not fast, but it is possible, and it is happening.

samdk··on Day Against DRM

    [DRM] can't simply be "abolished" without some suitable
    alternative to protects both interests.
It can be and in the music industry largely has been. iTunes has no DRM. Amazon's mp3 store has no DRM. Bandcamp has no DRM.

The solution is getting rid of DRM. DRM does not prevent people intent on pirating their media from doing whatever they want. It hurts legitimate users.

The existence of DRM has nothing to do with preserving the freedom of content creators. Your presentation of it as something that does is naive. It exists to protect the interests and power of media publishers. The publishers--not the creators--are the ones making the choices about DRM. Its use is an attempt to prop up outdated business models, because that's easier--and, in the short term, safer--than change.

samdk··on Introducing DuckDuckHack
For me, there are a couple of benefits of having the functionality built into my search engine rather than my browser.

First, all I have to do is set up DDG as the default search engine in my browser, and then I get tons of search shortcuts working without any more effort on my part. I don't have to worry about syncing between different computers, browsers (I use Chrome and Firefox at different times), and operating systems.

Second, there are many more shortcuts than I'd ever think to maintain myself, and somebody else has already put the work into defining them. I can do !{relatively popular site} and even if I'm guessing, it almost always works. (In fact, I can't remember the last time it didn't.)

samdk··on Python iteration
I use it rarely in Python, although it's sometimes useful. I use it more often in other languages to replicate Python's enumerate, since zip(range(len(seq)),seq) is basically enumerate(seq) (although with this implementation it's not a generator).
samdk··on mruby
One issue is that it's easy to end up with things that aren't round numbers by mistake (if you do any sort of division, for example, or are getting results from an external library, etc., and strategies for reducing errors that rely on the programmer not screwing up are generally doomed to failure.) That has several problems: floats are imprecise, you introduce the potential for additional runtime errors if you're doing something like indexing into an array, etc.

Also, Python has separate integer and floating-point operations:

    >>> type(1)
    <type 'int'>
    >>> type(1.0)
    <type 'float'>
samdk··on First hundred days of Clojure
Other people's attempts to answer your question seem to have not done what you're looking for, so I'm going to try to be more specific.

First and foremost, Clojure is a Lisp. Many other people more qualified to describe the benefits of Lisps have done so, so if you're looking to be convinced about why that aspect of the language is valuable, go and read what they've written. Looking at Clojure and ignoring the Lisp-related arguments is a bit silly.

It's dynamically typed, which sets it apart from the ML family of languages (SML, OCaml, Haskell, F#, etc). It has immutable data structures, which sets it apart from some of the elements of ML-like languages. It's eager, which sets it apart from Haskell (although Tyr42 rightly points out that many of its sequence operations are lazy). It has good support for multi-processor concurrency, unlike OCaml, because it has STM, and so doesn't have the problem of a global lock.

It's functional, and its primary programming style is functional, and it has immutable data structures, which sets it apart from Ruby, Python, Go, and to some extent JS. It's higher-level than Go/C++/C. It doesn't rely on callbacks for absolutely everything, like JS does. It's not object-oriented, which sets it apart from Ruby/Python/JS. It's a new language, which means it doesn't have the many accumulated years of cruft that Ruby/Python have, and it's well-designed, which means it doesn't have the many, many problems JS has as a language. It has multi-processor concurrency support that Ruby/Python/JS don't have.

Compared to other Lisps, it's a very practical language. It has native syntax support for vectors and hash tables. It's built on the JVM, which gives access to a lot of existing tools and libraries.

samdk··on A primer on Python decorators
The Python community generally advocates an "it's easier to ask for forgiveness than permission" coding style. When faced with a condition of the form "if condition a holds, do b, else c", it's very often a better idea to do "let's try b, and do c in case b fails because condition a didn't hold".

In this case it's better because you can avoid computing an extra hash of the object in cases where it's already a key of the dictionary. This may seem like a silly optimization, but it can very easily add up if you're accessing existing elements most of the time--I once had a bit of code that went 60x faster when I replaced an if-else with a try-except.

In other cases it can be even more beneficial. Say you're opening a file. One approach to avoid errors would be to check if a file exists first. This is error-prone because the file might cease to exist in between the 'if' and the 'open' statements, and now you have no code written to handle the error. Using try-except will ensure that you actually handle the error intelligently.

This isn't to say there's never a good reason to use an 'if' to check things, just that if you can do it in one step instead of two, one is usually better.

samdk··on Show HN: Password validity hinting
I don't normally comment on articles unless I've read them in their entirety, and this was no exception. I've re-read that paragraph in context several times, and I don't think it's entirely clear that you only meant it can't be done browser-side. That paragraph reads to me like you're saying bcrypt can't be used at all.

I don't think using bcrypt for this is feasible in any case. From a UI perspective bcrypt hurts you in the first place, because your feedback is going to take at least a few hundred milliseconds, so you lose the instant feedback that makes this potentially.

From the perspective of the person hosting whatever service is using this there are problems too. The first step from your "How It Works" section is this:

    On the server side, it first hashes entered password
    the same way it's done when doing regular authentication.
I'd say this is likely to increase the amount of time you spend hashing by about an order of magnitude. When you're using bcrypt, which takes a lot of computational power, I think that's likely to be significant. Especially because, in this case, you really can't afford to be queuing up requests--if you don't have the computational power to process them immediately, you get even worse at giving instant feedback. And again, that defeats the purpose of having this as a UI feature.

You having 10 years of applied cryptography experience means very little to me. Cryptography is one of those things it's very easy to get wrong even with experience. And getting it wrong has potentially very dangerous consequences.

samdk··on Show HN: Password validity hinting
While this is possibly an interesting UI experiment, this paragraph says everything you need to know about real-world use:

    Utilizing something like bcrypt as a hash function
    is not an option, because the number of rounds needs
    to be reasonably low for it to work in real-time in
    Javascript. 
If you're not using a slow hash function like bcrypt, you're doing password hashing wrong. Your users' passwords are vulnerable to brute-force attacks as soon as your database is compromised. (And "my database will never be compromised" is not a valid response.) Before doing anything with your users' passwords, read this: http://codahale.com/how-to-safely-store-a-password/
samdk··on Giles Bowkett summons monsters
I like this post very much. HN discussions often end up being about minor details rather than the overarching point, and I think that's unfortunate.

Those details are often things that are a matter of preference. There are good and bad arguments for both sides, and most of those arguments have been made many, many times already. Different people will hear the same exact arguments, draw on their own experiences and values, and come to different conclusions. And that's okay.

Which isn't to say that details are always irrelevant: just that you're missing a lot when you focus on the details to the exclusion of the larger point. Like in this post: it doesn't really matter whether or not Giles is baiting people intentionally. It may be true, it may be false, but it's not the point. The point is that people are responding as if he's doing it anyway.

samdk··on Learn from Haskell - Functional, Reusable JavaScript
The map is unnecessary. Just reduce:

    (* assuming: val total_len : string -> int *)

    List.reduce seq ~f:(fun a b ->
      if (total_len a) > (total_len b) then a else b)
It's not necessary in this case, but fold is often a lot more useful than reduce. At least the way I think of it, the type of reduce is 'a list -> ('a -> 'a -> 'a) -> 'a, whereas the type of fold is 'a list -> init:'b -> ('b -> 'a -> 'b) -> 'b. The upside is that you can construct basically any type of thing you'd like, since 'b is a completely different type. The downside is that if 'a is different than 'b, you need some sort of initial value to give it.

edit: Yes, this does require you to compute string length multiple times...but keep in mind that getting string length is very cheap in languages with good strings. (That is, basically everything except C's null-terminated strings.) 99% of the time it's not going not going to matter at all. If you do care, you can map to a tuple of (original_struct,total_len) and then do the reduce and then another map to get back to your original structure, or use a fold, as I mentioned, or write a (tail-)recursive function that does it in slightly fewer operations. (Although I don't think JS has tail-call optimizations, so that's probably a bad idea if you're doing it in JS.)

samdk··on Online Python Tutor
Having been a teaching assistant for intro programming classes targeted at non-CS majors, I disagree. You don't need to understand what a 'stack' or 'heap' is to understand what's going on and for this to be useful.

And even if it's too complicated to be useful for people on their own, being able to use something like this as a teaching aid would be very helpful. The number of people who have trouble grasping even what most of us consider very basic concepts like iteration over a list is astonishing. Trying to explain something you find trivially easy to someone who doesn't understand it at all is an exercise in frustration for both people. I usually resorted to doing a manual step-through on a whiteboard or piece of paper, but this is much nicer.

I would discourage people from using this constantly, just because I think having a step-through debugger handy constantly will prevent you from developing necessary debugging skills. (The concept of wrapping something in print statements is another thing I take for granted that's not obvious to some people.) But I still don't see how something like this could be anything but a net positive.

samdk··on IcedCoffeeScript
I haven't looked at this implementation specifically, but the issue with many similar proposals/implementations is that the generated JavaScript is ugly. One of the nice things about CoffeeScript is that there's no magic--it's very easy to figure out what the generated JavaScript is doing, and there's a straightforward CS <-> JS mapping. jashkenas (CoffeeScript's author) has been reluctant to change that.
samdk··on [dead]
The first most effective thing was taking a data structures class. It was the second programming class I took in college, and I use the knowledge from that class every day. I think that you can be a pretty good programmer without a lot of the theoretical stuff you get taught in CS classes, but data structures is not one of the things you can skip.

Another very useful thing I did was was read The Little Schemer (http://www.amazon.com/Little-Schemer-Daniel-P-Friedman/dp/02...). Before reading it I'd had a lot of trouble thinking recursively, and since reading it I've had very little trouble. I've also found its question-and-answer style very useful for teaching other people how to program using recursion.

samdk··on DuckDuckGo gets a new look
The positioning of sponsored links makes me very sad. I use j/k/enter to navigate most of the time, and this means that I need to worry about accidentally selecting a link with a near-100% probability of being completely useless to me. (It's really annoying when it's the first link, because I can't just press 'enter' to go to it.)

It also wastes a lot of screen space on a smaller screen, especially when the zero-click info box is there. Being able to see only 1-2 useful results instead of 2-3 is annoying.

I've been using DDG for over a year now as my primary search engine, and I like it quite a lot. I understand the need to make money, but I'm going to be very sad if I have to go find a new search engine because ads have compromised the UI.

samdk··on Jane Street releases open source alternative to OCaml's stdlib
I'm really not the right person to be answering this. I haven't been at Jane Street that long, and I have no experience programming OCaml without Core. You'll likely get a much better answer on the mailing list: https://groups.google.com/forum/#!forum/ocaml-core. (I've also posted a link there to this discussion already, so hopefully someone will come by and give you a better answer.)

Probably the biggest difference between Core and the INRIA (default) standard library is that Core itself is mostly written in OCaml. This has many benefits: code is less verbose, easier to read, it's easier to verify correctness, etc.

samdk··on Jane Street releases open source alternative to OCaml's stdlib
While this is mostly true, it's worth noting that this is part of a push to make Core easier to build and use, which we hope will encourage more people to use it. (And enable others to contribute.) There's still a ton of work to be done in that regard, but this release, at least, includes a Linux build script, and the libraries which are necessary for building Core.

Also, this release includes Async (our concurrency library), which is a much more recent release (October 2011). If you'd like more information about Async, there's Ron's blog post announcing the release [1], and also a previous comment of mine with some very basic sample Async code [2].

[1] https://ocaml.janestreet.com/?q=node/100

[2] http://news.ycombinator.com/item?id=3278532

samdk··on The CIO's lament: 20-something techies who quit after 1 year
HN's response so far to this is disappointing.

It's easy to read the first page (or just the headline) and come up with the "you're not paying enough/your problems aren't interesting enough" response. That's what pretty much everyone who's posted here so far has said. (And that's what my initial reaction was, too.)

But that's not a terribly useful reaction, and it's especially not a terribly useful one to be posting here, where pretty much everyone agrees with you already.

First, this guy understands (or at least claims to understand) a lot of the points you're all making already, and describes some of the steps he's taking to address them. (Mostly on pages 2 and 3.) Maybe it'll work for him, maybe it won't; I know far too little about the details to have any idea. In either case though, repeating the same "more money/interesting problems" thing over and over doesn't really have any effect.

Second, some of the problems he's having are actually real problems. Any company that's been around more than a few months is going to have existing systems, and it often doesn't make sense to rewrite the entire thing every time you need a new feature...even if the existing system is a bit ugly. As a programmer, that's an important thing to understand, and it's something people who don't have a lot of practical programming experience aren't necessarily going to understand. Knowing that this is a potential issue is something that can help you both as an employer and as an employee.

edit: I'm not saying that this guy actually knows what he's talking about, and that the things he's doing will actually fix the problems he's having. I'm saying that reactionary "this sucks" responses aren't very useful, especially here.

samdk··on Technical Papers Every Programmer Should Read (At Least Twice)
I'd add Types and Programming Languages to that list.
samdk··on How big are PHP arrays (and values) really?
Python and Ruby both have separate arrays and hashes. They're completely different data structures in both cases.

Python's dict type is a hash table like you'd expect. Python's list is a pointer-array-backed list. (it may inline ints/similar things--I don't remember if CPython does, and exact details are implementation-dependent), and raw arrays are in the standard library if you need them.

From a very quick check, a list of 100k ints in CPython is ~1.5mb, and a dict of 100k ints -> other ints is ~6mb.

Ruby's hash and array implementations are similar, I think, although I don't know Ruby as well, so I don't know the specifics.

samdk··on You've Probably Read Enough
To add: stop reading the same things over and over.

A recurring theme of discussion on HN is whether or not the submissions are getting worse. I tend to think they're not. (Although the comments are another story.) It's just that the twentieth article you read on something like lean startup methodologies is a lot less interesting than the first one.

samdk··on Roy — small functional language that compiles to JavaScript
Async includes a module called Monitor that lets you correctly handle exceptions in async code. The signature of one of the most commonly used methods is this:

    Monitor.try_with : ?name : string
                       -> (unit -> 'a Deferred.t)
                       -> ('a, exn) Result.t Deferred.t
It takes two arguments. Name is used to tell you what monitor the error was caught by. The second is a function that returns some kind of Deferred.t. The whole thing returns a (deferred) Result.t (: [`Ok of 'a | `Error of exn]), letting you either do something with the results of the async call or handle the error. Usage might be something like this:

    (* query_or_print : Db.t -> Db.Query.t -> Db.Result.t option Deferred.t *)
    let query_or_print db query =
      Monitor.try_with (fun () -> Db.query db query)
      >>| function
      | `Ok result -> Some result
      | `Error e -> print_endline (Exn.to_string e); None
samdk··on Roy — small functional language that compiles to JavaScript
Jane Street's Async library for OCaml does this: http://ocaml.janestreet.com/?q=node/100. It's excellent, and I now miss it in every other language I use.

Code looks something like this:

    ...
    (* long_running_call : unit -> string Deferred.t *)
    long_running_call ()
    >>= fun result ->
    print_endline result
    ...
Running calls in parallel is easy. Say we want to run a couple of queries against a db at once, and only perform an action once both return:

    ...
    (* Db.query : Db -> Db.Query -> Db.Result Deferred.t *)
    let d1 = Db.query db query1 in
    let d2 = Db.query db query2 in
    d1 >>= fun r1 ->
    d2 >>= fun r2 ->
    ...
If you have questions about getting it set up/using it, the mailing list is the place to ask: https://groups.google.com/forum/#!forum/ocaml-core
samdk··on Hiring advice for startups from Hackruiter (YC S10)
I finished up a job search a month and a half ago, and there were several places I didn't bother to apply to because I found their job posts annoying.

Pretty much anything than mentioned being a "ninja" or a "rockstar" was ignored immediately, and there were a few others I discarded out of hand, too for similar reasons.

Other than that, I'd just like to echo the point about responsiveness. There was one company I'd have loved to have worked whose job description I fit very well, but they took almost three weeks to respond to my application, and by then I'd already had an offer I really liked and I'd already written them off completely. There were a couple of others that said to bug them if I hadn't heard back in a while, and I mostly just stopped paying attention to those too.

samdk··on Advantages Of Being A Polyglot Programmer
There's really not a whole lot abstracted about it. It is really hard and impractical for actually accomplishing anything, but it's also pretty easy to see how it maps to a Turing machine if you know what the actual definition of a Turing machine is.

That kind of thing I learned in a class: a textbook on computational theory is probably what you want, and as far as I know there are only a couple of those. I used this one: http://www.amazon.com/Introduction-Theory-Computation-Michae...

samdk··on Ask HN: How do you "read"/understand/interpret existing code for software?
It really depends where you're starting from. If you don't have any idea what a particular piece of code does, finding that out is your first goal. If you can, ask someone: a two-sentence summary of what a module does can save you hours of digging through things, and takes almost no time for the person you're asking.

From there, I have a couple of things I usually look at immediately.

The most obvious thing to do is to try to find entry points into the code, and figure out what happens when those calls are made by reading sequentially. This is often the best way to work, although not always.

The second thing to do that can be very helpful is to look at the external library calls that are being made. Those can give you an idea of what the code is actually affecting, which can be very helpful for getting a high-level overview.

Another entry point is state that's being kept and modified from multiple locations. (Figure out what's being stored and why by looking at the code around where the state is modified.)

One last thing you might try to do is reformat the code if it's particularly ugly. I try not to refactor unless I understand what's going on so I don't introduce bugs, but just reformatting can make a big difference sometimes if the style of the person who wrote the code is significantly different from your own.

(And truth be told, 1000 lines of code in one file isn't that bad. When you start working with multiple files, some sort of method of jumping to function/method definitions becomes extremely useful. Since I use Vim, I often use something like ctags. Knowing how to use Unix tools like find and grep is also extremely helpful.)

samdk··on Why I hate the Kindle
These are all very valid points. Several of them (the ones related to DRM, mostly) kept me from buying a Kindle for a very long time.

However, I love my Kindle.

First and most importantly, it means that books can now compete with the internet in convenience, and that's meant that I've read more in the last few months than I have in the several years before that. Thinking "I want to read <X>" and being able to start reading minutes rather than days later is awesome.

Second, I find the Kindle to be much nicer to read than cheap trade paperbacks. The text is higher-contrast and crisper, and only having to push a button rather than turn a page means I can read more comfortably. That there's less text on each screen was an issue at first for me (because, like the author, I read very quickly), but I've since gotten used to it and don't even notice anymore.

It has its downsides, and I really do hope they're resolved at some point in the future, but my Kindle has become something I really wouldn't want to live without, at this point.

samdk··on GoDaddy SSL Cert Scam
Not with my money. You taking my money without specifically asking means that I will never, ever do business with you again. (And that I'll specifically warn other people against dealing with you.)

"It's better to ask forgiveness than permission" is not a universal truth to be applied to every aspect of startups/tech work. It is often a good idea when there's a lot of bureaucratic red tape and/or you need to get approval from other people about things. This is a very different situation: you're charging your customers ~4x more than they originally paid for something automatically. Are there other companies that do this sort of thing? Sure. Maybe it makes sense financially, but it's not a good way to treat your customers.

samdk··on I don't understand why anyone would use Node.js
1. I think you're wrong. I would agree that JavaScript's syntax is not great, but it's still better than Java's. (And dynamic typing, whatever your complaints with it, does significantly reduce how verbose the language is.) Also, first-class functions are a huge benefit for me.

Also a lot of JS's syntactical issues/verbosity can be alleviated with libraries like Underscore. Start using CoffeeScript and you're miles ahead of Java in terms of syntax and verbosity, and you make some of JS's more annoying issues a lot easier to avoid.

2. Yes, you do care about performance, to a certain point. If everything I was doing had to immediately be available at Google scale, then yes, I'd consider Java a lot more seriously. However, most of the stuff I do doesn't have to scale to millions of simultaneous connections on just a handful of servers. Absolute performance is not my highest consideration, and NodeJS absolutely is more productive for me than Java.

I think you're missing a couple of very important points, which are why I think NodeJS is the best tool for some tasks.

First, and by far most importantly, it has Socket.IO, which is by far the best choice library when you're doing anything in real time. It is fantastic, and is a huge part of the reason why I use NodeJS for certain classes of applications.

Second, there is value to being able to share code between the server and the client for nontrivial web applications. Most code doesn't get shared, but not having to have separately maintained models for the server and client is really helpful.

For me, the 'cool factor' doesn't come into it at all. I enjoy programming in JS and even moreso in CoffeeScript. I dislike programming in Java. And when I'm not doing something that has to be hugely scalable, Java's only real advantage is raw speed, and NodeJS has vastly superior tools for doing what I'm trying to do (real-time webapps), NodeJS is the clear choice.

samdk··on Captcha chaos
This looks very interesting on a technical level, but I'm not sure it really solves the problem of people using bad passwords, and it has some serious usability problems.

As far as I can see, this method is likely good enough to help in cases where people choose short but otherwise good passwords, but not in cases where people just plain choose bad passwords. Computers are getting increasingly good at solving captchas, and you can get human-solved captchas done very cheaply [1]. When the most common passwords are in use by more than 1% of your users [2], just trying those and using a combination of heuristics and cheap human labor still gets you a pretty large number of compromised accounts.

Even in the cases where it works, though, it comes at the cost of a really awful usability problem: users need to read and retype a long random string from the captcha every time they log in, and people hate catpchas. That alone is likely to prevent this being used in any application that depends on getting traction with a large number of users. Slow hashing algorithms like bcrypt or scrypt do a good enough job of protecting short-but-good passwords for most purposes.

The only cases where I can see this being actually used are in applications with very high security requirements, where people have no choice about whether to use the application, and where there are mechanisms in place preventing very bad passwords from being used. However, high-security applications are likely to have much lower traffic, which means that setting a very high bcrypt or scrypt work factor will accomplish much the same thing. That might take a bit more computational power for the application server, but if the traffic isn't very high, that cost isn't likely to be an insurmountable obstacle, and then you get to avoid all of the usability headaches that complicated captchas come with.

[1] http://motherjones.com/kevin-drum/2010/08/price-captcha

[2] http://blogs.wsj.com/digits/2010/12/13/the-top-50-gawker-med...

← PreviousPage 3 of 13Next →