HNHacker News
TopNewBestAskShowJobs

sb

521 karma · joined June 19, 2008

submissionscomments
sb··on Living outside America
That's an interesting point, but I guess that this is mostly due to the immigration friendly policies the US has. IIRC, I read somewhere in the Economist that the US is actually the only OECD country that is not going to suffer from a demographic shift. But if those policies would change, or research universities need to cut down their expenses due to budget problems, I'm not sure how long this would last. (IIRC, when Ireland's economy came to a halt, almost all foreign blue collar workers went back to their home countries in no time.)
sb··on France excels at R&D, but Academia is overdue for disruption
Agreed. Besides the length issue of a tweet (which don't apply to blog posts, as we all know from Steve Yegge) I also think that there is no incentive for a researcher to participate in the social web--at least none that I can think of. Actually, quite the contrary is true: yesterday's post on academic blogs linked to the "serious engineering" blog, which contains the following line in a post about time management:

  Do not read blogs and above all do not write them.
AFAIR, Matt Might had a somewhat different point of view, where he elaborates on strategies how to publish blog posts out of presentations, etc. To me the social web compares very well with the TV on the time-sink scale: you save a lot of time by not doing it.
sb··on The Future of Science
Well, actually I was not only having attribution errors in mind, too. Peer review ensures that colleagues will tell you about related work that you don't know about. Sometimes, people will tell you that something is related, even though you yourself don't actually think it's related work. Only with some time and acceptance, you will see that the remarks are really related, probably not directly to your own contribution, but to the bigger field that you orient yourself in.

Come to think of it that this is probably the most important detractor for having an NLP based "recommender." Personally, something like this might be interesting, probably even a great help, but at the end of the day, people need to really read a lot of papers, follow the proceedings of their target conferences, journals, and asking colleagues for their bibliographies. This has the added benefit of teaching them how to present their own work in contrast to others, do meaningful evaluations (in the best of all worlds, of course!) and figure out who is doing interesting work and might be valuable to get into contact with. Of course, some parts could be automated, but there is currently no incentive for scientists to do so.

IMHO, it would be a much more important step for CS researches to publish their code, too, because I frequently come across papers that have no implementation or evaluation at all--and that's really bad, because then the least-publishable unit becomes an idea with nice pictures. Researchers can be very successful using this "publication strategy." Come to think of it, there should be another approach to rank scientists by the number of publications, or their impact; unfortunately, I have no idea what could work instead.

sb··on The Future of Science
I am not sure that this will work anytime soon, even if NLP gets better substantially. I think the problem reduces to examination of patents and their novelty. I think that having algorithms to find related work is going to be hard to the point of it being unreliable. For example, in CS there is a trend of changing nomenclature every n years, which I think makes it hard to find related work.
sb··on The Future of Science
I think that your impression of doing related work is overly simplified. Since you are submitting to a conference/journal in your field, chances are high that the reviewers are knowledgable in the subject area and will point out errors in attributing due credits to related work.

While I agree that there are systemic problems with peer review and how the science "enterprise" works, there is fitting analogy from politics by Winston Churchill: "The worst form of government except for all the others."

sb··on How to cope with the Gmail redesign
Hm, it seems strange though that they're not keeping the old interface, which is what many people were using Gmail in the first place (even if their target group prefers the new interface.) Come to think of it, it's also kind of sad that HN users are obviously not their target group...
sb··on Why I Use Safari Instead Of Firefox
In addition, I think they completely failed with the HD themes, too. Some of them have such a contrast difference that you cannot even read label captions anymore, i.e., the complete left part below the Inbox becomes essentially useless...
sb··on Why I prefer scheme to Haskell
Assuming that Debug.Trace offers similar functionality as "observe" does, I have to briefly mention that this tracing might not be overly helpful, precisely because of the lacking control over evaluation order, as you write later on. I had to write a parser using both, combinatoric and monadic style, and debugging was a major turn off because I could not easily see what went wrong (though I seem to remember that using observe the ouptut was reverse to actual program flow, at least the one I have in my head.)
sb··on Introducing Gmail Tap
Agreed, what a perfect April's joke prelude.

(unrelated, but still interesting: I guess the guy with the mustache looks kind of familiar, too...)

sb··on Books Programmers Don't Really Read But Recommend
I wonder how often you looked at a compiler's implementation to make such claims. First of all, there are many parts in a compiler that are not fully automated (lexer and parser generators making frontend development easier; basically BURS systems for backend automation (w.r.t. instruction selection), but then in optimization you have to deal with instruction scheduling, register allocation, etc.)

Then, I also think that your comment about people having written top-down recursive descent parsers for a long time without concrete historical reference is suprising, particularly when we all know of a prominent example (C) that you cannot parse with LL techniques (I rememeber a professor at university remarking on K&R probably not knowing about LL -- though I can't attest to this being true or not.)

I can also not figure out how you connect the Dragon book to Parser combinators. Having implemented parsers in both, monadic and combinatoric style (both of which btw. are a mess to debug) the best connection I can think of is translating formal grammar descriptions using parser combinators. Is this what you are referring to?

I like your characterization of compilers being just calling functions and loops plus switches and pointers to text. While we know about Church-Turing thesis, I am positive that the intricacies of database system implementation, operating system implementation and programming langauge implementation deserves distinction (which is supported by many of them having dedicated special interest groups.)

sb··on Germany's unheralded computer inventor
Hi,

thanks so much for linking this. I have always been a fan of the history of computing (e.g., Moshe Vardi mentioned Charles Peirce in a talk once http://en.wikipedia.org/wiki/Charles_Sanders_Peirce, and I have seen a reconstruction of one of Leibniz' calculators [http://de.wikipedia.org/w/index.php?title=Datei:Leibnitzrech...), but I did not know that Konrad Zuse got a patent on the concept of pipelining in 1949. (AFAIR Hennesy and Pattern's "Computer Architecture: A Quantitative Approach" cites Tomasulo's algorithm from 1967)

sb··on What’s New In Python 3.3
Supporting your statement:

I am positive that Python 3 could be a lot faster than Python 2. If people are interested in what's possible, they should follow the corresponding mailing list (python-dev).

sb··on Parsing Techniques - A Practical Guide
There is also a second edition, which updates some chapters with much more recent resulst (AFAIR, the book is from 1992.) In addtion, the author (Dick Grune) also co-authored a book on compilers ("Modern Compiler Design"), which I like a lot as it has a sound treatment of non-imperative programming language concepts, too. (Plus I think the treatment of BURS plus usage for instruction selection is best described therein as well.)
sb··on PyPy 1.8 - business as usual
In yesterday's "Fast VM" thread I asked about memory consumption of PyPy vs. CPython, because most of the benchmarks focus on speed and there usually are no memory consumption figures (iirc, they were huge for Unladden Swallow [~800megs on one of their benchmarks, probably the django benchmark]). While I agree with fijal that there are no benchmarks stressing the memory sub-system, I think it would also be interesting for many potential early adopters to know about expected requirements.
sb··on Fast enough VMs in fast enough time
Hm, I agree about having benchmarks with memory impact is more compelling, but wouldn't it be at least interesting to show the memory impact as it is right now? (i.e., how much more memory does PyPy need?)
sb··on Fast enough VMs in fast enough time
Regarding the hints: without those hints the compiler would also trace the interpreter dispatch loop. If you trace them, too, then your generated code would contain unnecessary branching code. Hence, in a sense, these mechanisms allow the trace recorder to record the sequence of interpreter instructions executed without the interpreter dispatch interfering.
sb··on Fast enough VMs in fast enough time
While I find it good that the article explicitly addresses issues with trace-based compilation (usually this is not the case), a completely fair account needs to present the additional memory requirements for using the PyPy tool chain. Quite recently, somebody here has addressed this by mentioning that he does not really care for all the performance speedup he gets, if the memory requirements become outlandish at the same time.

It would also be very informative to know what the differences in automatic memory management techniques are (i.e., what did the previous implementation do?) Personally, I am also interested in interpreter optimization techniques, and it would therefore be interesting to me what--or if at all--the previous VM used for example threaded code or something along these lines.

sb··on Best Papers in Computer Science up to 2011
IIRC, expanding with programming languages (PLDI, CGO, ISMM, OOPSLA, etc.) was already discussed when this link was first mentioned on HN. Unfortunately, it seems to not have been done yet...

Anyways, PLDI has the "Best of PLDI" papers from 1979 to 1999 (http://www.informatik.uni-trier.de/~ley/db/conf/pldi/pldi200...) and they establish the most important paper 10 after its publication. AFAIK/IIRC OOPSLA/SPLASH does this now as well. I think it's good, but it would also be very interesting to know, whether there is an intersection between the set of "Best Paper" awards and the set of "Most important/impact" awards.

sb··on Technical Papers Every Programmer Should Read (At Least Twice)
Queinnec's Lisp in Small Pieces is a superb book, if you liked that you might also like Henderson's book on implementation of functional PLs (particularly SECD machines.)

Regarding the list: I think it's a nice effort, but lacks several excellent and remarkable books. Just quickly browsing, I found the following essential books missing:

  * In algorithms: 
    - Aho, Hopcroft, Ullman: The Design and Analysis of Algorithms.
    - Wirth: Algorithms and Data Structures.
  * In Compilers:
    - Wirth: Compiler Construction.
    - Grune, Bal, Jacobs, Cerial: Modern Compiler Design.
    - Muchnick: Advanced Compiler Design and Implementation.
  * In Lambda calculus:
    - Barendregt: Introduction to Lambda Calculus.
  * In theoretical computer science:
    - Kozen: Automata and Computability.
    - Davis: Computability and Unsolvability.
  * In concepts of PLs:
    - Turbak, Gifford, Sheldon: Design Concepts in Programming Languages.
(In general, the list does not mention any of Wirth's books, which is a shame. The Project Oberon book should also be mentioned in OS stuff, I guess...)

Regarding compiler construction: As has been previously mentioned several times on HN, Cooper's and Torczon's "Engineering a Compiler" is a more recent and (IMHO much more readable and accessible) choice.

sb··on Tornado twice as fast with PyPy 1.7 compared to Python 2.7
Which is actually what Python is doing, too. There is (IIRC) a mark and sweep collector that collects cycles. Another interesting technique for dealing with this problem is called "trial deletion."
sb··on Tornado twice as fast with PyPy 1.7 compared to Python 2.7
I think that's a good point and was also heavily commented on when Google's Unladden Swallow released its benchmark numbers (IIRC, for the django benchmark its binary size grew to 800 megs.) Probably that was even a reason they stopped working on it (there was a link somewhere, but I cannot find it right now.)

Furthermore, I think this "problem" is attributable to jit-compilation in general, since you have to store the code somewhere. The situation was/is somehow similar to the JVM's memory requirements. An interesting alternative to code generation is to optimize interpreters instead.

sb··on Smuggling Data in Pointers
Just for the record and to provide additional references, the technique of reusing a pointer for integer representation is rather old, for example fixnums refer to the same technique in Lisp implementation, IIRC, in Smalltalk this refers to SmallIntegers.
sb··on PG's Rarely Asked Questions
If you don't mind reading German then you might be interested in the the four volume series "Deutsche Gesellschaftsgeschichte" by Hans-Ulrich Wehler. Friends from Germany keep telling me that this is the reference about Germany's society. Parts II and III should cover 1848 and I figure there are similar texts for other European countries.

(Since you've read Faust and the intial post mentions Gibbon's Rise and Fall, you might be interested in Theodor Mommsen's "History of Rome" as well.)

sb··on Deft -- easy note taking for Emacs
seconded! I made the same beginner's error, too, and then decided that it was just the wrong thing to do. Emacs does not have a modal concept and cursor movement is not done with "hjkl" (and "ew") in Emacs. As Steve Yegge said, you should use incremental search forwards and backwards instead. I guess that eliminates many of vim's keystrokes (stats would be interesting), so staying with its keybindings might be prohibitive to learning Emacs...
sb··on Deft -- easy note taking for Emacs
Hi,

I have been using vim extensively for about 8 years and used a boring phase during my previous work-life to start learning Emacs. My motivation was not because I liked anything particularly well in Emacs or disliked vim, but more that I wanted to see first-hand what the difference is really all about. (So that I know what the flame wars are all about, without ever needing [and also never wanting] to participate in one...)

I am still using vim from time to time, but cannot imagine going back to vim full-time and leaving Emacs. The learning curve is steep and getting a nice setup takes considerable time (thankfully, there are very enlightening articles, for example the one from Steve Yegge, as well as the excellent emacswiki; plus, many people post their ".emacs" file on the web.) IIRC, it took me about 6 months until I felt proficient, and now, after almost 5 years or so, I couldn't actually be happier. There are many reasons to my happiness with Emacs (TRAMP, ido, yasnippet, auctex+reftex, vcs-interface, dired+, org-mode, macros, breadcrumbs, etc.) but I don't want to get into that, let's just finish this by saying: if you're generally interested and have some time at your hands (it's far less cumbersome as you might think, and mechanically codifying is usually [for me at least] not really the time consuming task in programming), just do it and stick with it for a couple of weeks!

sb··on Bill Dietrich gives CMU $250,000,000
Thanks for that list, it seems to me that this is the list all of the articles refer to.
sb··on Bill Dietrich gives CMU $250,000,000
I wondered what the other biggest gifts were, but could not find a canonical list. A good starter is the following:

http://web.mit.edu/newsoffice/nr/2000/neurogifts.html

I guess that its top 3 still hold. In addition, I found the following:

- $1b endowment to found Vedanta University from Anil Agarwal Foundation (2006)

- $454.5m to National Taiwan University from Terry Gou (2007)

- $400m to Columbia from John Kluge (4th largest in 2007)

- $360m to RPI from an anonymous donor (page mentions largest in US history in 2001)

I could not easily find the official list all of these pages refer to, anybody has an idea?

sb··on Miguel de Icaza: Learning Unix
Hm, instead of MC I prefer dired+ within Emacs, I have never used anything more powerful than this (particularly with TRAMP and the regex features.) So, if you are already learning Emacs, I think it pays off to at least take a look at dired(+).

(Minor remark: for smaller tasks [and instead of launching a terminal window] I prefer to use the DirOpus clone "worker" on UNIX.)

sb··on TinyVM 1.0 released; adds 16 lines of code, registers, a VM stack and more
Well, at least in the area of virtual machines and interpreters, the usual argument goes like this:

A stack-based virtual machine architecture has compact code representation (only bytecodes) where operands are pushed onto and popped of the corresponding argument stack. Register based VM-architectures require you to encode source and destination registers into the operations. IIRC, for Java bytecode, going from a stack-based representation to a register-based virtual machine grew the code size by more than 40%. But on the other hand interpretation got more efficient, since you have less instructions overall and thus fewer instruction dispatches. If you want more details I'll gladly point you to the excellent and canonical reference for this: Shi, Casey, Ertl and Gregg: "Virtual machine showdown: Stack versus registers." TACO http://dl.acm.org/citation.cfm?id=1328195.1328197 (There is also the journal article's predecessor from VEE05, https://www.usenix.org/events/vee05/full_papers/p153-yunhe.p...)

sb··on Want to Write a Compiler? Just Read These Two Papers.
I agree and there used to be (maybe it's still the case) problems with parser generators if you wanted to have good error recovery and reporting to the user. It's also very telling that sometimes--contrary to what people might expect--parsers pose a substantial problem in production systems: http://cacm.acm.org/magazines/2010/2/69354-a-few-billion-lin...:

Law: You can't check code you can't parse. Checking code deeply requires understanding the code's semantics. The most basic requirement is that you parse it. Parsing is considered a solved problem. Unfortunately, this view is naïve, rooted in the widely believed myth that programming languages exist.

← PreviousPage 2 of 7Next →