Chrome, multicore and embarrassing parallelism
blog.luckycal.com
blog.luckycal.com
Not to say it’s a bad idea, just that we’ve got to remember that multiple processes is nothing revolutionary. A single thread for each tab would probably work just as well from a parallelization standpoint, but you’d lose the ability to kill one without bringing down the entire browser.
But don't undervalue the difference in mindset. At least for me, threading makes me think "what can I peel off of my main task to run in threads" versus share-nothing which makes me think "here are the 3 shared resources that I anticipate will be bottlenecks".
Expensive lines make systems less flexible, not more. They lead to copies for efficiency, aka denormalization, which is another word for "bug waiting to happen". While this can be dealt with, doing so involves lots of nasty tradeoffs, with bugs a common outcome.
More to the point "more modular and more flexible" is neither necessary nor sufficient for producing good software.
I'm glad I'm not the only one that thinks this way. I wish I could upmod by 1 million.
Don't do that.
> "here are the 3 shared resources that I anticipate will be bottlenecks".
Do something like that.
The difference between threads and processes is in the mechanisms for sharing. While those mechanisms affect how you organize computation (different things are cheap), they don't mandate an organization.
They have a common document cache, a common cookie store, a common history store, probably a common dns lookup result cache, and maybe some other things. These all need to be carefully synchronized between multiple tabs.
In fact, I would argue that multi threaded would be better than multi process so that even more things could be easily shared. For example, I imagine that css stylesheets are parsed into some big fat data structure inside browsers. If I open two tabs from a website that share a stylesheet it would be optimal if they share this same internal representation (no locking would be required for the sharing since it's read-only). This has the obvious savings of memory, but it also increases speed since the css file only has to be parsed and processed once instead of multiple times. And things like sharing keep-alive connections between tabs are virtually impossible with multi process, while very possible with multiple threads.
Uhh, from the OS they get that for free.
This actually works quite well since the OS tends to schedule a process on the same core, so processes tend to always access the local memory partition.
Why? Because the "working quite well" is only really valid for single threads with no shared data structures. In other words, when you pretend a thread is a process.
As soon as you have more than one high load thread, the OS will want to split them across multiple processors, which means that you're now trying to share the same chunk of memory between processes. If the OS tries to keep them on the same core, though, then you've got 2 processes competing for CPU time and leaving another core free.
Then again, even turning them into processes on a modern OS wouldn't distribute the memory contention all the time; memory is usually copy on write across forked processes, which means that unless you've written the memory, reads are still contending for the same bus.
When the alternative is using a shared bus all the time, it works out nicely this way.
Most programs are single threaded, and even most multithreaded programs don't share that much between threads. Of course the scheduler is going to schedule across CPUs in a reasonable way, but if it makes sense to keep it on the same CPU, it does that. The point is to keep bus contention to a minimum, and this does that in the average case.
Again, this architecture makes sense for the "lots of totally independent processes" case. The problem is that this case isn't as common as you'd expect. on Linux, if you fork a process, you're sharing memory between them until you write to it. in threads, you're sharing all read-only data unless you've explicitly duplicated it.
1. Fewer synchronization points. You now have a two-level hierarchy (threads & procs) vs just 1 (just b/w threads). You can reduce the scope of the hardest-hitting syncs (e.g. malloc) to contend within a smaller arena.
Anyone well-versed in Win32 know if there are any GDI-related advantages to different processes? Synchronizing graphics context accesses is another PITA, which may be avoidable here.
2. Resiliency: a single bad plugin or V8 bug need not take down the entire browser. Just a self-closing tab.
3. Security. (not in this version of the browser) The supervisor proc may be able to set per-process access controls (e.g. RBAC) to keep the tabs in check. E.g. keeping activeX in control, while allowing the in-house controls to function with as much access as they need.
A lot of OS functionality is process based, and using them allows for a lot of open space for new possibilities.
In both Safari 2 and WebKit nightlies, GIFs don’t animate unless they are being painted somewhere. If an animated GIF becomes invisible, then the animation will pause and no CPU will be consumed by the animation. Therefore all animated images in a background tab will not animate until the page in that tab becomes visible. If an animated GIF is scrolled offscreen even on a foreground page, it will stop animating until it becomes visible again.
Many plugins do animation and work based off being pumped “null events” in which they do processing. The faster you pump these events, the faster animations will occur, and the more CPU will be used. Safari 2 actually throttles these events aggressively to background windows and background tabs.
As well as a background Pandora tab.
The hard core javascript guys I know talk about developing interesting client side code compare it to developing the old 8 bit computer games, and trying to do something really cool and interesting with 64K of memory. They're trying to do something interesting and bundle it up in 64K of code, so the download speeds don't cause people to bounce to another site.
I see this occasionally due to browser bugs. (This is tautological, since I define pegging the CPU as a bug.) In some sense Chrome is the first postmodern browser: instead of trying to eliminate bugs it lets you kill -9 tabs.
Be interesting to see some evidence that's likely to happen.
As long as you clear ~100 cores, I think you're done.
Why do you think so? The 2007 book "Future Directions in Processor Design" says (p483, Ch21):
http://www.springerlink.com/content/xp811205j8523537/
If we think of the technology development as predicted in the ITRS roadmap [208], it is clear that more and more processors will be crammed onto a single chip. The prediction for the roadmap is that we will see hundreds or even thousands of precessors integrated within the next ten ... fifteen years, e.g., 424 precessing elements per chip in 2017. The trend can be confirmed by looking at some ambitious high-end projects in multi-core and multi-processor development. For example, Rapport Inc. is shipping a chip with 256 processing elements on board and is developing a 1,024-core processor, however these are only 8-bit elements [281].
...
My bet is that we will see more specialized processors for different specific tasks. The "one size fits all" simply cannot provide enough cost and power efficient enough solutions for the embedded sector. Thus, we will see various special-purpose off-the-shelf cores emerging.
It isn't an argument about the technology, rather about what humans will be able to manage.
If you really think Larrabee is only a GPU, think again. Each core is a full 64-bit x86 that can quite happily run anything that´s running on my Core 2 Duo.
If you believe it will continue, and that clock speeds won't increase, then you get exponentially increasing #'s of cores (or some other use of silicon estate).
An alternative scenario is that there will be no mainstream demand for more computing power, and only niche markets will buy them. The mainstream will favour cheap, less powerful CPU's. The rise of eeePC and clones is suggestive of this scenario.
Only if you believe the wiring connecting the cores wouldn't explode with the number of cores.
A single bus works fine for moving data to/from one of many memory locations. Another method of using just one bus is TCP-IP packet-style communication. Of course, there's limited bandwidth, so you want most of your computation done within a processor, not between them.
For many problems, bisection bandwidth is a key constraint.