HNHacker News
TopNewBestAskShowJobs

reginaldo

1,400 karma · joined October 12, 2008

Reginaldo Silva

Contact: reginaldo at ubercomp.com

Links: http://www.ubercomp.com/ https://github.com/ubercomp/

submissionscomments
reginaldo··on Douglas Crockford on Fat Arrow Functions in JavaScript
The new functions do not have names, so you will need to use the old functions to write self-recursive functions.

Or you can use the Y Combinator for extra fun...

https://secure.wikimedia.org/wikipedia/en/wiki/YCombinator#E...

reginaldo··on Effective examples of startup commercials
Did you see the bald guy with cream on his head is reading "The Lean Startup" at about 25s?
reginaldo··on Show HN: Open Source emulator in Javascript that runs Linux
That would actually work, I guess :). I'm pretty sure the memcpy, memmove, etc. implementations I've seen for this architecture all work by copying one byte at a time (argh!) though, but that problem can be solved with some smart-if-ugly hackery.

But, as you said yourself, it would be kind of messy, in that it would be hard to be certain if it works in 100% of the cases.

I think it would be easier to just change the toolchain code and make it think the processor is little endian.

reginaldo··on Show HN: Open Source emulator in Javascript that runs Linux
Ok, one way to do interpretation is to do a loop like

  while True:
    instr = decode_instruction(pc)
    execute_instruction(instr)

Where execute instruction is usually a big switch that has every opcode on it, or, in my case, is a call to a function in a table whose indices are opcodes.

The problem is you decode the same instructions many times (in a loop, for instance). A slight more elaborated version is like this:

  while True:
    if not pc in cached_code:
      block = generate_block(pc)
      cached_code[pc] = block
    run_from_cache(pc)

The generate_block function decodes and generates code for every instruction until it finds a jump. When it finds it, it finishes up that block, which is then added to the cache.

Now, whenever there is a jump to an address that is in the cache, the cached code can be executed directly.

Finally, the captures up to three backwards jumps thing:

My first try was to have the cached code just be a regular function call without loops. So a block of code is executed and then control is given back to the main loop.

But a significant number of jumps have targets that are inside the same block of code (anything you run on a loop, for instance). So, to get better performance, I generate, for my block of code, a function that has a loop in it.

It has a rather characteristic look:

  function code_block(cs) { // cs is the "cpu state"
      var count = 0;
      cs.next_pc = BLOCK_START;
      block_loop: while(++count < 3) {
          count++;
          block_choice: switch(cs.next_pc) {
              case BLOCK_START: cs.r0 = cs.r0 ^cs.r0;
              case BLOCK_START+4: // regular instructions fall through
              case CONDITIONAL_BRANCH_ADDR:
                  if(branch_condition) {
                      cs.next_pc = BRANCH_TARGET;
                      break block_choice; // this captures backwards jumps targeting the block
                  }
              case  UNCONDITIONAL_BRANCH_ADDR:
                      cs.next_pc = BRANCH_TARGET;
              default: break block_loop;

          }
      }
  }
The "captures up to three backwards jumps things" is controlled by the ++count < 3 loop condition. 3 was obtained by trial and error and looks like a good number for my machine (I could, for instance, put 100 in there, but then my javascript functions would take too much time to return and the javascript engines don't like this).
reginaldo··on Show HN: Open Source emulator in Javascript that runs Linux
Yeah, it is a known bug. The hterm terminal emulator runs on its own iframe, so you'll have to click anywhere on the screen (outside the terminal). If that doesn't work, then it's a new bug...
reginaldo··on DuckDuckGo is blowing up
pg's analysis was spot on, then [1]. I recently realized DuckDuckGo would be a hit when a hacker friend who does not read HN started using DDG as his primary search engine.

To me, this means it is out there. I switched too, at least for a while, and now I use both google and DDG, constantly having the impression that DDG's results are getting better and better.

[1] http://paulgraham.com/ambitious.html

reginaldo··on Show HN: Open Source emulator in Javascript that runs Linux
> I don't understand this..

Well, I won't speculate on the workings of the Chrome profiler, but what I can say is that profiling in Chrome and Firefox gives different results (which may of course be caused by differences in the Javascript engines).

> I wonder if your RAM accesses are mostly 32bit

All accesses on the LM32 processor are aligned. Half word accesses have 2 byte alignment, and word accesses have 4 byte alignment. So in theory it would be easy.

I use ArrayBuffers for the RAM array (where supported). You can have different views for the same array buffer, so in theory I could have the ram ArrayBuffer with an 8bit view, a 16bit view, and a 32bit view, where each would be used for accesses of their size. In that sense, you're absolutely right and I think a good speedup would be achieved by that change (something like 2x speedup as most accesses are in fact 32 bit).

I don't do that for endianness reasons. The LM32 is big endian, and most devices nowadays are little endian. Unfortunately, the state of ArrayBuffer is still kind of chaotic, in that the endianness of the ArrayBuffer is the endianness of the host (they do this for performance reasons). So I can't do the v32[rpc] thing. Theoretically, I could if the programs have Read-Write consistency, i.e., if they only used the same size for reading and writing some data, which they don't. A program that tests the endianness of the processor, for instance, does not have read write consistency.

I could use ArrayBufferViews to access memory in an endianness-independent way, but I tried it and it's actually much slower.

To make memory faster, two things can be done either: 1) emulate a little endian processor instead. 2) have a fast byteswap operator in javascript, without a function call (never gonna happen).

So since 2 is never going to happen, the solution is to tackle 1.

reginaldo··on 10 years of Linux and Red Hat Linux 9 (Shrike) is the best OS i have ever used
It looks like HN's ranking algorithm is giving a boost to submissions by new users. I submitted a perfectly fine story (shameless plug about my an emulator I wrote) that got 48 points in 9 hours and didn't make it:

http://news.ycombinator.com/item?id=3769498

reginaldo··on Show HN: Open Source emulator in Javascript that runs Linux
> - how did you find what to optimise? - what approach did you use?

I used profilers: Chrome developer tools and Firebug. The Chrome profiler is very nice and easy to use, but it looks like it is a statistical profiler. When the usual invocation takes less than 1ms, it might get the times wrong (there might be a bottleneck that Chrome profiler is not showing).

The Firebug profiler is more precise, I believe it traces every executing instruction, which also makes it much slower to run you program in, but does sometime gives you interesting insights.

Also, there's intuition. I wanted to keep the code readable. The interpreter loop is basically something to the effect of

  while(true) {
    op = decode(mmu.read_32(cs.pc));
    cs.pc += 4;
    (optable[opcode])(op);
  }
which I believe is very readable, but it counters everything one expects about writing fast javascript code that runs in a tight loop. Bellard told me he thinks my "dynamic interpreted" version might be slower than a throughly optimized interpreter version, and he might very well be true. But I found the readability of the interpreter to be of the upmost importance, at least until I had a working "dynamic recompilation" version.

To optimize the interpreter version, the profiler showed me that basically three things take most (about 65, 70% of the time: * Instruction decoding * Memory access

There's not much I found to make instruction decoding faster , but the memory access time could be improved. In the beginning, I used the form mmu.read_32(addr) to read from any region of the memory, including RAM. Most of the accesses are of course from RAM, so it made sense to optimize for this fact (despite it being bad for readability).

So nowadays I do: op = (v8[rpc] << 24) | (v8[rpc + 1] << 16) | (v8[rpc + 2] << 8) | (v8[rpc + 3]);

For memory accesses.

Some other things I found randomly. The processor in the board I'm emulating has a 75Mhz clock. With that emulated frequency, Linux was spending a lot of time on the "idle task". Whan I told the emulator that the frequency was about 8Mhz, the problem went away.

Then, I thought to myself: it would be neat if I could implement parts of the libc in Javascript. That's how https://github.com/ubercomp/jslm32/blob/master/src/lm32_libc... came to be. So I implemented some string.h functions in javascript (memcpy, memmove, and memset), and voilá, the boot time went down about 3 times.

But in the end of the day, lm32_libc_dev is a hack with very negative consequences. I would have to modify every program I wanted to run on the emulated board to use my libc functions instead. I did this with the linux kernel and it was a bad experience. So I focused my energy on the dynamic code generation side of things.

I can tell you I have succeeded as the jslibc thing is now deleted on my local source three and I didn't experience any slowdowns whatsoever.

reginaldo··on Building the worst Linux PC ever: 6 hours to boot Ubuntu
Wow. Contrary to my expectations, this picked up some interest. A lot of people actually visited my demo.

As I don't want to kidnap this thread, which is about an awesome feat of hackery by dmitrygr, I submitted one about jslm32: http://news.ycombinator.com/item?id=3769498

I'll check the jslm32-specific thread sporadically to answer any questions.

reginaldo··on Show HN: Open Source emulator in Javascript that runs Linux
Since there seems to be some interest on the subject, and I don't want to kidnap the great thread about the worst Linux PC ever [1], I decided to submit on a separate thread.

Some things do keep in mind:

1. Very early stage. Runs better on Chrome.

2. If you think this sucks, you're probably comparing my code to Fabrice Bellard's Jslinux[2], and as Bellard is a genius his code (which is unfortunately still not available in readable form) is probably much better than mine. But progress is being made.

3. Still, the potential is there, soon it will be possible to have a somewhat standard programming stack and write GUI applications that run off a canvas, as I said in [3].

4. For Chrome there is a nice terminal emulator (hterm) I lifted from the Chromium source tree. The rest use textarea.

5. The code, of course: https://github.com/ubercomp/jslm32/

If anyone has questions, I'll be happy to answer them.

[1] http://news.ycombinator.com/item?id=3767410 [2] http://bellard.org/jslinux [3] http://news.ycombinator.com/item?id=3769486

EDIT: Added itens 4 and 5.

reginaldo··on Building the worst Linux PC ever: 6 hours to boot Ubuntu
Working on it...

The LatticeMico32 toolchain (gcc, gdb, ar, ld) is very bad, seriously... I'm thinking of doing a MIPS or ARM emulator, just because the toolchains are so much better to work with.

Anyways, I do have a framebuffer demo that runs at a very decent frame rate on my machine (at this moment it is Chrome only and I unfortunately don't have the binaries for the demo on github). But this is not on Linux, it's on a barebones newlib environment (no Operating System).

Writing a mouse interface is trivial, the only reason I haven't done it yet is I'm thinking if it is worth to continue investing on the LatticeMico32 architecture, as the toolchain makes me want to pull my hair out.

Besides running Linux, the emulated system also runs RTEMS (actually it runs anything as long as you can get the toolchain to produce working binaries), and it might be easier to get an RTEMS system running with a simple graphical environment, but then there wouldn't be man y libraries do choose from.

So that's the status. I believe we are on the verge of having a viable option for making "GUI" programs on a canvas screen. If I were to work full time on it, I could pull a prototype off in about a month or two (literally), maybe a little less as this thing is so addictive I would easily work 16h days on it if I could.

IMHO, the ideal thing would be having an easy to use environment with a mature toolkit on it (I was thinking Qt) and let the user choose the language of choice to develop in. Possible choices would be C, Python, Ruby and Lua, which are fairly easy to have on the web.

Technically, it can be done, I just don't know if there would be enough interest, and how to have a sustainable business model around this idea. What do you think?

reginaldo··on Building the worst Linux PC ever: 6 hours to boot Ubuntu
These crazy projects are the most fun. As Richard Feynman notoriously said: "What I cannot create, I do not understand".

shameless plug below:

Last year, inspired by Bellard's jslinux, I too wrote an emulator that can run Linux on the browser. Only I was lazy and emulated the vastly easier LatticeMico32 processor.

Anyways, the result was very intellectually satisfying.

After writing the interpreter, I went ahead and wrote a version that generates Javascript code on the fly (and captures up to 3 backwards jumps to the same block), for massive speed ups.

Anyways, it doesn't serve any purposes, but boy was it fun...

The code is at: https://github.com/ubercomp/jslm32/

And there's a demo running on: http://www.ubercomp.com/jslm32/src/

BEWARE: It only works well on Chrome (takes download time + 10s to boot on my machine).

If anyone is interested in this stuff, just ask and I'll write a post describing what I did to take boot time from 2.5 minutes to 10 seconds.

reginaldo··on A Pinterest spammer tells all
The crawler better be undetectable as such. For instance, it better send "expected" headers (User Agent, Accepts, etc.), and it better have cookies enabled, and also operate from many distinct and perpetually changing IP addresses.

Otherwise, the spammer will be able to run his/her own URL shortener service in a 5USD/month VPS and be able to show a spammy link to the users and a regular-looking link for the crawler.

BTW: a "crawler" implemented with Mechanical Turk workers would be a little bit harder to detect, but would also have its downsides.

reginaldo··on Why do we need "? extends" in Java
Yes, you're right. We're talking about variance when it comes to the "? extends, ? super" situation, not erasure. And you're also right the things would be better with definition-site variance.

On the other hand, sometimes I think an unsound type system with List<String> being a subtype of List<Object> would be better than what we have today, for pragmatic reasons. Of course, I think this makes me a non-type-theorist, as what I'm saying is considered heresy in some circles [1].

[1] http://lambda-the-ultimate.org/node/4377 (search for unsound).

reginaldo··on Why do we need "? extends" in Java
Not true. You'll get, for instance, people who read Effective Java [1], by Joshua Bloch (who is, IMHO, THE man when it comes to java), and remembered the PECS mnemonics. PECS stands for producer-extends, consumer-super.

So, to cite the textbook example, say you are implementing a Stack<E>. It will probably have the methods:

  public void push(E element);
  public E pop();
and, for convenience:

  public void pushAll(Iterable<? extends E> elements);
  public void popAll(Collection<? super E> destination);
The pushAll has "? extends" because the elements Iterable will "produce" elements for the stack. The popAll has super because the destination Collection will "consume" elements from the stack. It is not that hard, is it? Let's note that guard-of-terra is talking about proficiency, not mere familiarity. I believe reading "Effective Java" is a nice way to get closer to the proficient level.

Let it be noted that this whole mess exists because generics in Java were implemented with type erasure so their introduction wouldn't break legacy code. I personally think this was a bad idea, but it does show that when a language is evolving, there are a bunch of constraints the designers must be aware of.

[1] http://www.amazon.com/Effective-Java-Edition-Joshua-Bloch/dp...

reginaldo··on ZeroVM: lightweight containers based on Google Native Client
Thanks for the info. Seriously I should have said "almost nobody" in the first place. In fact I thought I had said that :)

I didn't know nfshost was doing it. From what I see, they're probably using FreeBSD jails in this case, which is nice if they are.

Anyways, we have to agree that this space is largely unexplored. I never thought there was a need for this kind of service, but just after reading the ZeroVM pages I think it's a very good idea. With a "little" more initial effort, it would enable writing systems in a very interesting way: self-healing (when the other end has failed, make an API call to provision another copy of it), self-provisioning (when traffic is high, make an API call to provision another copy of a worker), etc. Of course we can already do this already, it would just be more natural, and if you combine this with the idea of Mobile Agents, then the cloud suddenly becomes much "cloudier".

reginaldo··on Poll: What's Your Favorite Programming Language?
I feel likewise (except for the Perl part. I have the most fun working with PLT-Racket tools).

I believe this happens because language design involves a lot of trade-offs, so every language incorporates these trade-offs.

When one is comfortable in many kinds of languages, one is in a special position to see the trade-offs in the design of the language one is currently programming it.

The way I see it is slightly different then the way you describe. Say I'm programming in Prolog, for instance. In the beginning, I'll be like "Wow, such expressiveness". But inevitably I'll do something that is outside the scope of Prolog and it will be as slow as hell. That's when I think to myself: "if I could just fix this little part... I miss C". So when I'm back programming in C I'll think "now my code is fast", but then inevitably: "if I could just find a way to do this without so much repetition".

In the end, it comes down as a "right tool for the job" thing. One possible ideal would be do the things Prolog is good at in Prolog, the things C is good at in C (and also throw in some Python, Erlang, Lisp, etc). Except most of the time this is infeasible. Working with FFIs suck. Sometimes there are no FFIs, sometimes there are many incompatible ones, and the dream of calling any language from any language is just further and further apart. And don't get me started on the pain that it is to have multiple runtimes with slightly different semantics...

Even if it were possible, a developer would have to learn all these little languages enough to be working on them comfortably, and we all know this is not going to happen. So hardly anybody does this, as hiring someone for a polyglot project is a complete nightmare.

reginaldo··on ZeroVM: lightweight containers based on Google Native Client
From what I read it's something like this:

It gives a bare-bones environment for you to run your programs that is presumably very low overhead. Think of it as an embedded system where programs run without an OS. This is the environment a program running inside zerovm will see. All you have is libc and the zerovm-provided APIs. If you want more, you'll have to statically link your programs.

The thing is, you can run many, lets say thousands, of these little programs inside a single machine in such a way that each one can never see the other ones (as long as it's impossible to break out of the ZeroVM sandbox).

Such a technology would enable neat stuff, like renting a server for someone to run a single program for some period of time and have the results sent back. Nobody does this for unrestricted programs today, for many reasons, a very important one being the fact that it would be very hard to do this in a secure way.

The "run a C program for some period of time thing" would work kind of like the AWS dashboard, but instead of having to spin up a machine with linux on it and running your program inside that, you would only upload your binary and a manifest file. Kind of like what app engine does, but with less restrictions (you'll probably be able to do anything as long as you're able to compile a "safe" binary that does it).

reginaldo··on Facebook: Legal action against employers asking for your password
IANAL, and I don't live in the US either, but I can tell you that at least here in Brazil the network traffic is the property of the employer, and you have no expectation of privacy while working, so they can do whatever they want with the traffic that is going to their routers.
reginaldo··on Space Monkey Dropbox Competitor Wins Launch, Has Already Raised $750K
First, let me say that I'm not convinced this new wave of peer-to-peer based systems is the way go forward. I believe poorly implemented yet popular peer-to-peer systems might make the internet worse (as in slower) for everyone and, more importantly, cause unpleasing surprises when peer-to-peer traffic that is unknown by the users starts hitting their data caps.

But, for the sake of argument, let me play devil's advocate and take the "yes, of course you should do it, dugh".

1. Yes, users care that their files are available. Given the recent megaupload events and the outbreak of copyright-related-madness that has contaminated the US government, I don't discard the possibility that some random raid might cause Dropbox files to be unavailable. The reason for the raid might be because "users are using Dropbox to illegally store copyrighted files", which I'm sure it's true. For availability, distributed is better than centralized, period. Of course, I have 4 or 5 fully synced computers, so my Dropbox setup already is kind of distributed and I would be relatively immune, but I sincerely don't know how many users use this kind of setup.

2. I don't go over the free quota because I'm constantly policing my files and removing everything that is not essential. What I would really like is to use Dropbox as a backup-everything-sync-everywhere tool, not as a sync-essential-stuff-everywhere tool, but I currently do my on backups, because I can't bring myself to paying 10 bucks a month for 50GB space if I'm only using 10GB.

3. I agree. Most people don't care about anonymity unless they are affected by the lack of it in a way that is both personal and perceived as negative, which most aren't.

Ease: there is nothing that makes the Space Monkey way fundamentally more difficult for the user. Of course, the user's internet connection might fail a little bit, and probably any contender that wants to compete in this area will have to plan for that scenario, but this is not a deal-breaking thing (for instance, they can store user A's data on user B's device when A's connection is not working properly).

Taking of my devil's advocate mask, while I believe there might be some space for companies in the "Dropbox competitor" space, I believe being just a "Dropbox competitor" won't cut it, even though I would like someone else to succeed even if just for the sake of diversity.

I also believe Dropbox is more like a feature than a product, and putting a special-purpose device inside someone's house offers you the possibility of turning said device into a kind of command-of-control thing, such that it can do many things more than being just a backup-and-sync device. For instance, I could be on the street, hear about a movie, and use my cell phone to make my device buy and download this movie so that when I'm home I'll be able to watch it. This kind of stuff would bring "just works" to a whole new level.

By the way, I buy a movie using the device, Space Monkey might take a cut. God, the hardware these days is so cheap that it might be possible to offer it for free and make money by taking a cut on every transaction.

reginaldo··on How to solve supposedly intractable problems
This is a great post. It basically says that when you have a problem that seems intractable, look for the conditions under which it is intractable and see if the conditions apply for your specific case. Many times they don't. And even if they do, you can often pretend they don't and still get solutions that are good enough.

For a small pearl on problem solving, I recommend "How to solve it"[1], by the great mathematician George Pólya[2]. It teaches you simple techniques you can apply when you are stuck, like "draw a picture", "think about a similar problem you already know the solution for", and "solve a relaxed version of your problem". It all looks pretty much like common sense, but it is not. It's one of those few books I think everyone should read (whenever they are stuck).

[1] http://www.amazon.com/How-Solve-Mathematical-Princeton-Scien...

[2] http://en.wikipedia.org/wiki/George_P%C3%B3lya

reginaldo··on Books every self-taught computer scientist should read
You provided a good definition: solves problems scientifically rather than by brute force. I believe the essence of science comes from being able to think (and report your thinking) in a structured way. Meaning it is about content, but it is also about form.

And just so we can get this out of the way, I'm not in any way offended by people applying labels to themselves in good faith, I'm only slighted offended by people who are offended by that.

reginaldo··on Books every self-taught computer scientist should read
I have a set of honest questions to make:

I understand what it means to be a self-taught programmer, which I was from ages 13 to 17, before I got into college and, in some sense, still am today, as many of the things I learn and have learned for the past few years were "self-taught" (or, as I prefer to say, taught me by the authors of books, papers, and blog posts I read).

Some people like to call themselves self-taught hackers, software engineers, etc. And that is ok too.

But what does it mean to be a self-taught computer scientist? Is there some criteria, e.g. do you have to publish a peer reviewed paper or something like that?

reginaldo··on TypedJS: Sanity Check your JavaScript
Sometimes, and I mean just sometimes, it is better to just use ===. For instance, I recently wrote some code where a variable has a value of type Number, but I use "false" to signal the absence of such value (in the sense that the value is not available yet). In that case, I have to use stuff like if(x !== false), because "0" is a valid value in that context.
reginaldo··on Mitmproxy - an SSL-capable man-in-the-middle proxy
I have used stanford's ssl-mitm, Paros and WebScarab. The one being shown here is better than all of them, actually.

The textual interface is actually a curses-based interface that lets you do stuff as you would normally do in paros, such as capture a request or edit a request and replay it. So it is not just like you're watching logs. You can take many kinds of actions from within that text interface, and very quiclky. The scripting API is also very good.

reginaldo··on Super Cheap Virtual Private Servers - the Wild West of Hosting
I am also happy Quickweb customer, although I only use their VPS for OpenVPN and to compile stuff on Linux. Stuff compiles more quickly on the VPS than on my 2010 MacBook Pro, which means processor and IO performance are good.

They've lost my data twice in a 2 year period, but for 4.95 a month I was kind of expecting it. But now they moved my VPS to another node and all should be fine. Which brings us to customer support: it is absolutely great and they respont to tickets very quickly.

So: use for non-critical stuff, backup, and be happy.

reginaldo··on Devops for Everyone - Beanstalk's hosted deployments now can run SSH commands
Congratulations on the new release.

However, if we're talking about Devops for Everyone, I consider myself pratically obligated to tell people about the Salt Stack[1], even though I have no association with the project, just because it is awesome.

It is, IMHO, Remote Execution Done Right. Also, it does not use ssh at all, and I believe it will be able to handle thousands of simultaneous machines in no time. To see what I'm talking about, just take a look at the FLOSS Weekly episode about Salt[2], where the main author himself admits this is his fourth iteration in trying to make a remote execution engine that does not suck.

[1] http://saltstack.org/ [2] http://twit.tv/show/floss-weekly/191

reginaldo··on Poll: How many Raspberry Pi boards do you plan to buy?
I'm buying 3 as soon as they're out. I was going to buy only one, but again the price and the temptation to make them talk to each other are more than enough reasons for me to buy 3.
reginaldo··on Poll: HN readers, where's your residence?
I'm in São José dos Campos and would go to a meetup either in Campinas or in São Paulo (apart from SJC, of course).
← PreviousPage 6 of 8Next →