HNHacker News
TopNewBestAskShowJobs

davisp

173 karma · joined August 20, 2009

submissionscomments
davisp··on The origins of the steam engine
We could, but it would certainly be less efficient than other approaches. I haven’t seen it brought up recently but there’s a simple quote that makes it easy to evaluate any sort of energy related discussion based on the laws of thermodynamics:

“You can’t win, you can’t break even, and you can’t quit the game.” - C.P. Snow [1] (No idea if that’s the correct attribution, I heard it from a professor 20 years ago)

My personal shorthand for this is to focus on the efficiencies of energy transformation and transport. Converting or moving energy means we lose some fraction of the total energy. All we can do is minimize that fraction.

The main point being that with that mental framework, I can easily assume that your posited approach won’t be more efficient than a solar panel or some other process with fewer transformation steps because each step magnifies energy loss inefficiencies.

Basically, the entirety of the energy debate is two things: How can we harvest energy cheaply and move it with minimal waste.

This is the basis for the whole room temp superconductor story that captured a bunch of our nerd imaginations, what happens if you get super close to zero energy loss while moving electrons?

Also, one more thing to keep in mind is constraints. Your proposed solution may be way more efficient by some metric if we change the constraints. For instance, my brain mostly goes to transportation which is transport and storage with immediate demand requirements. However, if our requirements were different along the lines of “lets store as much energy as possible over months, to slowly release it over months”, inefficiencies of scale change things. Then maybe slowly moving water uphill during summer while releasing slowly during winter is a valid thing to consider [2].

[1] https://gohighbrow.com/laws-of-thermodynamics/

[2] https://medium.com/drax/how-to-turn-a-mountain-into-a-giant-...

davisp··on Cats that fetch are an evolutionary mystery
Not necessarily. There’s an entire class of dog breeds literally named “retriever”. A quick search shows the AKC saying the first retriever breed was from the 16th century which means humans have been selectively breeding dogs for this specific behavior for nearly 500 years.

Whether or not we consider the basal instinct as retrieval vs some other similar instinct that can be easily used to motivate the retrieval behavior [1], I’d wager most dog owners that have trained (or attempted at least!) would agree that some dogs are just more likely to become good retrievers which to me suggests that there’s more than just a basic pleasure stimulus response.

Similar dog behaviors based on breed that I would argue at least point towards instinctual behavior being at least a partial explanation for behavior are herding and live stock protection breeds. Take a quick watch of any of the sheep herding competitions or other fun related videos [2] to see what I mean.

However, at the end of the day I also think that most “instincts” are more of a shorthand/stereotype for some set of abstract probabilities that may or may not result in a dog being good or bad at any specifically desired behavior.

[1] https://www.wildfowlmag.com/editorial/tips_strategies_wf_ins... [2] https://www.youtube.com/watch?v=D2FX9rviEhw

davisp··on The Problem of West Bank Settlements
First, thank you for giving me the benefit of the doubt on the internet as I know that can be hard these days. I have no idea who OP is so its totally possible I got caught by propaganda.

My read of the article was that it seemed like a historically based account of the current situation with some fairly obvious bias/editorialization mixed in so I was interested to see what other folks would post in response.

I’ll absolutely look into both of those authors for more background.

davisp··on The Problem of West Bank Settlements
Why did this get nuked? I’m no expert on the region, but it seemed about as unbiased of an opinion on the conflict as is currently possible.

If there are members of the community that feel strongly about this article misrepresenting something I’d be very interested in hearing those points of view to help educate myself.

davisp··on SciPy builds for Python 3.12 on Windows are a minor miracle
If you’ve ever wished that CMake’s ExternalProject pattern could be lit on fire, launched into the sea, fished out of the sea and then incinerated via tactical nuclear strike before bundling up the ashes that are then launched into the sun, vcpkg is definitely something you should look into.

I’m no Microsoft fanboy. I’ve been in software dev for roughly 20 years at this point so generally view anything from Microsoft with genuine suspicion, so I get the hesitation to take it seriously. But it works across the big three (Windows, Linux, macOS) and is MIT licensed so I’d definitely recommend giving it a whirl.

The only serious knock against it is that they went the OG Homebrew route with a single Git repo containing all of their ports (equivalent to Homebrew Formulae). And then whoever designed the Git repo approach also knew slightly too much about Git internals and leveraged tree-ish refs as part of the versioning design which is just weird and confuses anyone that’s not spent time tearing into Git’s object model.

So basically, vcpkg is honestly a good tool that does what it does fairly well. It may not do everything you need, but if it can it’s amazing.

Also, the buried lede here is how vcpkg handles binary caching. Think of it like sccache but at the dependency level. I’ve seen it drop CI runs from over an hour to 10m purely because it helps skip building dependencies without resorting to bespoke caching strategies.

davisp··on The Accept Header
As in all things web, this is related to user agents (specifically IE) having a horrendous misunderstanding of what the Accept header is for. If you google "Accept header IE" you'll see a long list of hits about its particular brand of gnarly.

An example IE Accept header might look like such:

  image/jpeg, application/x-ms-application, image/gif,
  application/xaml+xml, image/pjpeg, application/x-ms-xbap,  
  application/msword, application/vnd.ms-excel,  
  application/x-shockwave-flash, */*
Notice the obvious missing entries like text/html, text/plain etc etc. By default, IE is saying "I prefer all these other things to HTML."

The second follow on here is that if a resource provides two content-types for the same resource, and an accept header matches both with equal priority, then the server gets to decide which representation to return. The first example the author gives is technically correct in this regard. His probing shows that Netflix is being a bit naughty, but not without just cause.

Basically, they're making the assumption, "Anyone that asks for JSON really wants it." which does violate the spec, but given the long history of user agents that have historically broken this spec I can only say that its not a question of whether the spec is broken, but how it's broken.

For instance, the spec mentions in passing that given an Accept header of:

  text/html, text/*
That text/html should be preferred to text/plain even though both would end up with a quality of 1.0. Although there's no actual algorithm defined for dealing with this situation. The entire decision tree is left implicit based on the examples provided.

Anyway, getting in a huff that your browser extension breaks a website that tries to cater to developers while providing reasonable content to people that still use IE is a bit awkward. Especially if you've written a book on HTTP and haven't realized that maybe there's a certain amount of spec breakage to deal with weirdo user agents.

davisp··on Isaacs: try/catch is an anti-pattern
The issue with the JSON spec is that it's deceptively simple. For instance, this particular issue comes up relatively often. As you point out it's quite specific that a valid "JSON text" has a top level object or array. This assumption is even required for the sections on "detecting character encodings" (which makes me wince to this day).

The particular issue is that the spec refers to "JSON text" (which no one ever uses in practice) as well as "JSON value" which is what everyone expects. The difference being that a "JSON value" is any of the 'bare' types (null, boolean, number, string, object, or array). A quick survey of JSON parsers will show you that most will accept any of these types at the top level. The one notable exception I'm aware of is Ruby's JSON library (the default one? I'm not so hip to the Ruby).

Notably, though, is that the major JavaScript interpreters don't enforce this constraint (Nor do Python or YAJL (kinda) or Erlang (even the ones I didn't write)).

Don't bother asking me about invalid combining characters as \u escapes. You wouldn't like me when I get the rage eyes.

davisp··on PouchDB (Portable CouchDB JavaScript implementation)
Yes and no.

The most important part of replication is to have an algorithm that's capable of merging two document histories that are not identical. This core bit of CouchDB is quite important and often overlooked in terms of its replication scheme. Randal Leeds is currently hacking this into PouchDB and I'm quite excited to see the algorithm in a non-functional language so that more people can see its simplicity without worrying about learning a new programming paradigm.

On the other hand the atomic update of the two indexes is quite important for efficient replication. Without the atomic update of a by-update-sequence index it would require a full table scan for each replication. Its definitely a necessary optimization, but not a sufficient optimization. The per-doc revision history merging is the special sauce that makes things work.

In unrelated news:

http://twitter.com/vmische/status/100956289387077633

davisp··on Cloudant (YC S08) Releases In-Database, Distributed Search
The issue there is that you're still anchoring your index to one end of the string which means you're not solving the general problem, only a specific manifestation of it.

A general example would be given the string "foo bar baz", your solution could find "foo%" or "%baz" efficiently, but not "%bar%". Its not out of the question if what you really want is a suffix search, but the general problem of finding an internal substring is still less than optimal.

Edit: Formatting

davisp··on Cloudant (YC S08) Releases In-Database, Distributed Search
Its hard to say. There are quite a few ways in which this could play out.

Firstly, there are two important points to consider. Currently, BigCouch is more or less a superset of Apache CouchDB. The only patches we have to CouchDB sources can and should be back-ported but require that we solve a couple possibilities for bugs in non-clustered deployments. Secondly, Erlang is a language which allows for an easy mish-mashing of code so that once we have back-ported these patches there's no real requirement for a merge at all.

There are also a few things that we're discussing in the CouchDB community that could very well contribute to not needing to merge the projects. Specifically rearranging our source tree to be more prototypically Erlang as well as some tools like a couch-config script that could allow plugin-type extensions to CouchDB.

In the end, its hard to tell how things will shape up. It could be a full on back-port, or it could just be a general improvement to CouchDB's source tree and build system so that BigCouch is strictly "CouchDB + Other Erlang Apps" if that makes sense. And with my CouchDB committer hat on, it really depends on what the community wants. Its easy to fall into think of the trap of "it's obvious" but we also have to consider that others are taking CouchDB and porting it to mobile phones. What we end up with in "core" CouchDB has to consider a lot of use cases.

davisp··on Django Setup using Nginx and Gunicorn
As you say, its not so cut and dry.

Gunicorn has only had a mailing list since about May. Most of our dev work is based in IRC or via the GitHub issues tracker. The three committers, @benoitc, @tilgovi, and @davisp have a combined 1709 twitter followers. There are also 479 watchers and 64 forks of the main development repo on GitHub.

Bottom line, numbers like this don't really mean much. In the end, people should investigate any project they think is a candidate and make a decision based on their specific criteria rather than try and use some proxy metric to make the decision.

davisp··on Django Setup using Nginx and Gunicorn
If you have Gunicorn crash please make sure and report any tracebacks in the logs to the issue tracker so we know that's something is broken.

https://github.com/benoitc/gunicorn/issues

davisp··on Propellerhead Software - Application Programmer Test
Its been awhile since I C++'ed heavily, but I think this isn't quite right:

"But even there, the code (most likely) dereferences a pointer one past the end of an array, which is (1) invalid C++ and (2) not on their list of possible bugs."

There's a weirdo thing with array allocation in C++ at least which makes it valid to use arrays with all of the std:: algorithms that take a start and end parameter. Unfortunately I can't remember it well enough to find a good reference but as I recall it was more than just a pointer comparison as the std:: algorithms wanted to be able to dereference it occasionally.

Then again its been about four years since I've touched C++ so some of these details have leaked out and gotten mixed up.

davisp··on LevelDB: a fast and lightweight key/value database library
Managed to get it working last night:

http://code.google.com/p/leveldb/issues/detail?id=2

davisp··on LevelDB: a fast and lightweight key/value database library
I managed to get an initial patch up here:

http://code.google.com/p/leveldb/issues/detail?id=2

davisp··on Why programmers are not paid in proportion to their productivity
I see quite a few comments addressing the fact that there's no way to directly measure 'productivity' of a programmer. And there's also the questioning of whether the 10x figure is made up or not (Its not. There's research that attempts to measure productivity, but its always by some proxy measure, usually time to complete a task. There was a post not too long ago that collected various articles on the origin on the 10x number).

So far I haven't seen anyone point out that there are lots of abstract concepts that we measure and rank without having a specific measurable quantity. Art is the first thing that comes to mind. What makes someone a 'good' artist? Even in things people might normally consider numbers based like college football rankings aren't purely objective.

The to deciding some ranking and relative comparisons is to ask people familiar with the subject to compare to instances of the set. In this particular case I could see having the team rank each other member or assign some numeric score representing their opinion. Giving this and some brief written reasoning for the scoring it seems like management would be able to get a clear picture of who the most 'productive' programmers are. There'd be obvious chances for pitfalls and politics in such a system, but if it were applied with some sanity it seems like it could work.

davisp··on On moving from CouchDB to Riak
The main two reasons that CouchDB doesn't use multiple files per database are system limits and increased complexity.

CouchDB has a bit of an alternative design in that it accepts that people might be running a large number of databases on a single node. I don't remember the exact numbers but I think we've heard of deploys using 10-100K (small) db's on a single node.

As to complexity, with a single file, there's no magical fsync dance to coordinate when committing data to disk. Its not unpossible, its just more complex.

Its not out of the question that CouchDB will move to using multiple files per database, but as its open source, the biggest road block so far is that no one's needed it badly enough to implement it.

davisp··on Quora Raises Questions
I'll preface this with a mention that I haven't taken the time to check out Quora yet. The only interaction I've had is to read the occasional answer someone links on twitter.

> Following questions and topics makes perfect sense to me.

But he didn't mention questions as a source of confusion. Those would be almost understandable (new message for new answer). As to topics, is it new questions, new answers? What type of events trigger a notification in my Quora stream? As to following users its the same question, what events trigger a notification?

For things like twitter its fairly straight forward to the point that normal people look at you funny. "Why would I want to know what someone had for breakfast?" is a perfect example of great UI (though, perhaps a deficit in marketing). The reason its great is that the entire system is easily and directly communicable to non-techies. The "What are you doing right now?" (which I just noticed is now "What's happening?") is such a simple thing that most people don't realize that its an example of genius UI.

I have just formulated a theory that technical folks are at a disadvantage for identifying and creating a great UI experience. A stereo type of technical folks is that they like to solve puzzles. So a UI that presents tiny puzzles in the form of "What do you suppose this does?" are like tiny injections of happy directly into the blood stream. So if the annoyances are unnoticeable they can be a good thing.

The fact that you switched from "confusing" to "innovative" means that while they may have a good idea hidden in that particular UI device, it is not good UI because it requires the user to figure it out. Good UI is not something you figure out. Good UI is "I know what this is and what it will do without ever having seen or used it before."

davisp··on Mustache 2.0 and the Future of Mustache.js
Yes to else sections. The lack of else is just plain annoying after awhile.

Don't care on dotted access.

Context addressing is fine if we can agree on a method and keep the spec simple. Perhaps using Git's SHA addressing with things like HEAD, HEAD^, HEAD^^ or similar.

Yes to booting dynamically changing delimiters.

Yes to stealing helpers.

For cross-language compatibility, we really should have a collection of test inputs and outputs.

We should adopt a directive thing like you have for the 'this' notation with {{.}} or whatever for lists. The original is a bit funny looking but I think this sort of thing would fix the delimiter issue.

Another pattern that I run into that's a bit weird is when I need to add markup around a list when its not empty. I end up having a context that returns a context with a rows object when there is data, or None when there are no rows. So it looks like this in mustache:

{{#genes}} open table {{#rows}} row stuff {{/rows}} end table {{/genes}}

I keep having the feeling there's a better way to make that work.

If I get another project finished up shortly I'll write these up in a branch of pystache.

davisp··on Programmers think differently than non-programmers
In contrast with the author's point of view, there is research that suggests that the ability to program is dependent on an ability/desire to form consistent mental models. The actual paper is at [1] while Jeff Atwood's write up that's been linked to from HN before is at [2].

  [1] http://www.eis.mdx.ac.uk/research/PhDArea/saeed/paper1.pdf
  [2] http://www.codinghorror.com/blog/2006/07/separating-programming-sheep-from-non-programming-goats.html
[edit: formatting]
davisp··on Cassandra vs MongoDB vs CouchDB vs Redis vs Riak comparison
They're not guaranteed to be the primary revision for a document (ie, the last write could be listed as the conflict) but the data is still there so that client code can resolve the conflict.

Also of note, if the edits that caused the conflict are replicated to other nodes, each node will independently choose the same revision to use as the 'primary' document response.

Bottom line, the choice is deterministic and is guaranteed to be preserved, but the choice may not be the last written revision.

Also, bugs that result in corrupt dbs are treated as major bugs as opposed to a part of the design. I've also not seen reports of index corruption under load, if you have logs or any more information we'd definitely appreciate if you could put that info into a ticket on JIRA [1] or even just mail dev@couchdb.apache.org with details.

[1] https://issues.apache.org/jira/browse/COUCHDB

davisp··on This is not the Freehackers Union...
The idea about voting isn't to select a "best" hack each night, its to say "take some time to find a better hack". That way if someone does get through the pre-screening you have a way to say, "take a couple weeks to think about your next Tell."

Granted the nation wide voting idea kinda gives the wrong impression on that part. I still think it'd be a good idea to keep videos somewhere.

Granted it was mostly just hand waving on how to keep the non-hacker attendees from overwhelming the hacker presenters.

davisp··on This is not the Freehackers Union...
Was I the only one reciting the rules to Fight Club while reading the article?

Rule 8: If this is your first night at fight club, you have to fight.

http://en.wikipedia.org/wiki/Fight_Club#Plot_summary

Basic rules that come to mind to keep it hacker centric could be:

  1. If this is your first night to Hack and Tell, you must
     Tell.
  2. Before leaving Hack And Tell you must vote yay or nay
     on each Tell. If your Tell has less than X votes, you
     are not allowed back for Y weeks.
  3. If you're just returning from a Y week layoff, you must
     Tell.
  4. You must have a Tell every Z weeks.
I'm not at all certain what the correct values for X, Y, and Z are. Z would obviously depend on the group size and time allotment. I think voting should be anonymous and not announced until after everyone's left so voters are more likely to be honest.

If it gets popular I could even see having a Hack and Tell website that hosted videos of "successful" Tells and so forth. Going out on a limb and having national votes to try and get a group of hackers together for a nationally voted "Best Hack" event or some such.

Could be fun.

[ed: formatting]

davisp··on No Java 7, The End Game
Let me preference this as I'm not a Java guy, but I've been reading a bit since Oracle filed suit because of the ASF angle.

From what I understand, the Apache Harmony project was doing a clean room implementation of the Java standard library. There was a clause somewhere (I haven't researched the exact text of this) that said something along the lines of "You must pass the TCK (Java's language test suite sort of like rubyspec http://rubyspec.org/) to be a 'compatible implmentation'," or some such.

Apache Harmony people started writing code and when it got time to say, "Ok, can we have the TCK to validate this?" Sun said, "Yes, as long as you promise no one will run your shit on mobile phones."

If you know the ASF, this sort of thing is wildly against their core mission: to make unencumbered software. So they gave Sun-Now-Oracle the finger.

Now, we need to cover a few bases here. Apache Harmony is not a VM, its an implementation of the standard library. The patent issues Oracle sued over were related to the Dalvik VM. The important point here is that the Dalvik VM is not Java. Its an interpreter for byte code. The byte code that runs on Dalvik will not run on any other VM. People think that Android Java is Java, but its not. Android Java is Java that's been compiled to JVM bytecode and then been translated to Dalvik bytecode (again, IIUC).

The basic point is that 'The Java Language' is not a JVM. It is not a standard library. But there's some weird ass 'Certified Java' thing that can't be obtained unless you give up rights, even though Java is open source.

My reading of this IBM/Oracle tryst is that IBM decided that it could get a "TCK certified" Java implementation by backing OpenJava.

Which is a long way of saying, no, Google/Apache probably can't just create a new Java spec. The JCP et al will prevent such things. Perhaps they could fork and make a RangerRick language that was source level compatible, but something tells me these people employ too many lawyers for that to go over well.

davisp··on Scratched glasses give perfect vision for any eyesight
I don't have access to the full text article so I can't say for certain but the article makes it sound like the effect is accomplished in a fashion similar to phase contrast microscopy.

http://en.wikipedia.org/wiki/Phase_contrast_microscopy

The basic concept of phase contrast is that certain objects (notably, monolayers of biological specimens) don't greatly affect the amplitude of light passing through them. IOW, viewed under "normal" optics, they would appear transparent.

On the other hand it was noticed that these objects do cause the phase of light to shift as it passes through the specimens. The basic idea is that you create a system that has light passing through the specimen in two paths. The major light path is dimmed and accelerated by half a wavelength. This gives us a situation of theoretically perfect destructive interference at the point of observation. When the reference light waves pass through a phase object, it changes the phase of the light and introduces perturbances to the completely destructive interference which results in an observable image of the specimen.

The reason these glasses remind me of phase contrast is the concentric circles milled to specific depths and widths (which is one of two important pieces of a phase contrast microscope, the objective). Also, the article describes the effects based in terms of phase:

    The rings shift the phase of the light waves passing
    through the lens, leading to patterns of both 
    constructive and destructive interference. Using a
    computer model to calculate how changes in the diameter
    and position of the rings alter the pattern, Zalevsky
    came up with a design that creates a channel of 
    constructive interference perpendicular to the lens 
    through each of the 25 structures. Within these 
    channels, light from both near and distant objects is
    in perfect focus.
I never took the optics courses in school so I'm not able to connect how the phase properties are leveraged for focus. The only thing that comes to mind is that they're using phase differences to selectively destruct out of focus rays and reinforcing in focus rays (hand wavy factor: 8.6 of 10).

[ed: speling]

davisp··on You Spilled Your CouchDB in My Riak
Most interesting. Though I tend to wonder about how compaction might work. Perhaps the root node links to the previous version and use a merkle tree to find nodes that can be deleted.

Also, I'd have to agree with [p]. We need to spend some quality time cleaning things up. Though I'm not sure about Rebar here. I quite like it, but one thing I've learned watching Noah do his thing is that downstream packagers expect lots of esoteric Autotools features.

davisp··on Tests for programmers Part IV: comparing routines
Pretty neat stuff. The method used for decimating representations is pretty awesome. Seeing the aligned routines makes me think that quite a few of the bioinformatics algorithms could be useful for such an analysis.

The two that spring to mind would are the MCL clustering algorithm that could be applied quite easily to the similarity matrix. As a more academic endeavor, it'd also be interesting to see what kinds of differences in similarity you'd get by applying Needleman-Wunsch.

http://www.micans.org/mcl/ http://en.wikipedia.org/wiki/Needleman%E2%80%93Wunsch_algori...

davisp··on Ask HN: Python framework advice
WebOb, Routes, SQLAlchemy, Mako. My favorite framework is no framework.

The bottom line is that its a sliding scale. Django will give you a large amount of functionality for free but requires that you do things the Django way. As you move along the scale towards projects like Werkzeug you'll inevitably lose some of the free stuff but you'll be able to make a lot more choices on how things work.

You'll want to check Ian Bicking's blog for some good reading on the no-framework approach. [1] is a great place to start.

[1] http://blog.ianbicking.org/2010/03/12/a-webob-app-example/

davisp··on Benchmark of Python web servers
I'm glad he updated it. A lot of people would've just left it. That said I also would've liked to see at least a note that he tried more than three workers and there wasn't an improvement. When you have an almost linear improvement its a bit confusing why you wouldn't keep increasing it.
davisp··on AGPL != Awesome GPL
So you're saying that I can license my code under anything as long as I still follow the rules of the GPL...

http://www.fsf.org/licensing/licenses/gpl-faq.html#IfLibrary...

Does mention "or compatible license".

And I'm really confused about dynamic languages that aren't explicitly compiled before distribution.

← PreviousPage 2 of 3Next →