Web Browser Engineering
browser.engineering
browser.engineering
Developing for the browser is a real challenge. I think working with html / css /js has been a neglected skill for a long time - most software engineers look down on that type of work and its rarely covered in comp sci course work.
Still, its good to see a lot of progress has been made, this book included.
My only critique - why use python instead of node.js?
Basically: server-side JavaScript is just not as widely known as Python, and it'd be additionally confusing when our browser starts running JavaScript. And in-browser JavaScript is a bit too restricted (by things like the same-origin policy) to do the whole thing inside a browser.
I am resisting the urge to disagree, but since you did (literally) write a book about building a browser, I will defer to your expertise and try to learn from you :)
Respect
Does it have RSS?
I don't think either is a super compelling choice anyways for this type of work. I think you want to use a systems language here. However Python is completely fine as a teaching language. Pretty much anyone who knows a similarly structured language can read it. And there is very little noise. So it can serve as a good reference if you want to follow along with a different language.
IMHO, those are job skills and not comp sci topics. They shouldn't be part of a degree program (except the most superficial treatment required to get some ugly UI up that may be required for something else). You have your whole career to pick them up.
This made me start digging into whether this was considered a "safe" way of executing untrusted JavaScript in a sandbox.
It's not completely clear to me if DukPy currently attempts safe evaluation - it's missing options for setting time or memory limits on executed code for example: https://github.com/amol-/dukpy
There's a QuickJS Python wrapper here which offers those limits: https://github.com/PetterS/quickjs
I'm pretty paranoid though any time it comes to security and dependencies written in C, so I'd love to see a Python wrapper around a JavaScript engine that has safe sandbox execution as a key goal plus an extensive track record to back it up!
The time limit and memory limit support looks good too: https://github.com/sqreen/PyMiniRacer/blob/f7b9da0d4987ca7d1...
The funny thing is we can have the MINIX microkernel discussions all over again :)
EDIT: Reading you comment again I suspect you might have been joking :)
There's IPC. There's memory management. There's process management. There's network management. There's security. There's device management.
They all happen at a slightly higher layer, but they all exist similar to an OS. (I'm not sure if the higher layer makes it easier or harder to understand - but in terms of what you need to know, an OS class or three is definitely helpful)
I imagine the non-trivial parts of these are done for the JS VM (correct me if I'm wrong), and therefore a VM design course would have more intersection with Browsers with respect to these disciplines than an OS course.
> security
This one is everywhere, it has no special connection to OSes.
> They all happen at a slightly higher layer
Slightly?! That's an understatement of the week! The difference in abstraction levels is huge here and the specifics of the two levels are very, very different.
> in terms of what you need to know, an OS class or three is definitely helpful
Sure. But I think it's as useful as any systems programming course. I can agree that systems programming is a good preliminary for both Browsers and OSes, and learning either of the two will teach you a good deal about systems programming, but I doubt they will repeat each other.
Multi-process architecture requires you to think deeply about IPC.
Memory management is all over the place - there isn't a browser without custom allocators, investment into GC, etc.
Process management -> see multi-process architecture.
Network management: Browsers need to handle a tremendous amount of network issues. I mean... that's what they do. Outside of the VM, too.
As for "security is everywhere" - the whole point of an OS (and a browser) is to make it possible for security to be everywhere. To provide the primitives that you can securely build on.
> Slightly?! That's an understatement of the week!
Not really, no. I've worked on embedded networking stacks, on full-fledged OSs, and on browsers. I stand by "slightly". Yes, granted, a browser doesn't get quite as bit-fiddly as a on-the-metal OS, but it's a matter of degrees, not quality.
Can you work on many areas of a browser without ever touching OS-like code? Absolutely. This particular book has a good chance of avoiding most, because it focuses on the rendering part.
But a browser, as a whole, provides an abstracted platform just like an OS. And it echoes many concepts, if in slightly different forms.
You need to execute potentially malicious code in isolated containers (processes / pages). You need to protect the system from these processes, and protect these processes from one another. You need to run some code with elevated privileges / additional capabilities (setuid / permissioned APIs), including some with superuser privileges (root user processes / browser extensions).
OSs are built on hardware, browsers are built on OSs. Even though they "sort of" do the same types of things, the abstractions are very different.
- OS memory management is more about managing hardware page tables, copy-on-write, page sharing, page protections etc. Browser memory management is typically custom allocators optimizing for different types of structures.
- OS process management is about using CPU primitives to multitask segments of code, context switching, timers, wait queues, interrupt management etc. In the browser, it's more about managing webpage/process relationships, application threading, rendering pipelines, and the JS runtime.
- OS network management is all about interfaces, packet processing, buffer management, low level protocols (ethernet, IP, etc.) Browser network management is all about protocols, and not really much about the hardware.
Sure there's a lot of broad-stroke similarity between OSs and browsers, but for a university course, where one typically gets deep into the details, they're entirely different.
That's exactly what we are hoping for.
http://browser.engineering/preface.html
So far, my co-author Pavel has taught from this book multiple times (including this semester). In the spring at least one other university will offer a course. We'll list all known courses offerings on the website.
Also, if anyone would like to teach from this book, please get in touch!
Looks like RemoteHQ launched something called Remote Browser as well, which is a SaaS browser, but more for collaborative purposes with more than one person using the same browser at the same time.
https://www.html5rocks.com/en/tutorials/internals/howbrowser...
That article, along with a number of other resources, are listed here:
https://browser.engineering/bibliography.html
In my view, a critical part of really learning how something as complicated as a browser works is by trying to build it yourself. That's why our book is oriented around building a browser as you go.
edit: it's a great article! But nothing on rendering tables :-)
HTML 5 effort has cleaned up a lot of behaviors and specified how browser tags should behave. So it is, possibly, an approachable task now. Still daunting though.
I'm the rendering lead for Chrome, and know quite a lot about how it works. I also recently wrote a series of articles about the new rendering architecture of Chromium, see here:
https://developer.chrome.com/blog/renderingng/
Pavel is a professor at the University of Utah and has extensively studied CSS from an academic point of view. He also has a lot of experience teaching the material and making it accessible to students.
(This is a really fantastic resource; I'll probably use it as the basis of course once I get around to semi-retiring and taking a teaching job. Thanks, OP!)
W3C here is unfortunately a part to the problem.
Standardisation is good, but letting google pour streams halfassedly written RFCs onto other browsermakers is not good.
Non-enforcement of standards is also bad, and it's bad to extend W3C privileges to companies who themselves selectively implement their own proposals, so others' browsers can't match their behaviour.
https://www.fastcompany.com/90611677/flow-ekioh-web-browser-...
Also you should take a look at the WHATWG because it’s far more relevant than the W3C nowadays.
https://drewdevault.com/2020/03/18/Reckless-limitless-scope....
No idea how Flow does it, but building a browser is nearly impossible.
Flow didn't start "from scratch" recently, it's an evolution of a primarily SVG+CSS renderer for set top boxes. They also re-use Spidermonkey as their Javascript engine.
It's not meaningless. Because in order to implement a browser, you have to figure out which of them are dupes, deprecated, drafts etc.
And even that won't help you. Because a huge amount of "deprecated" standards are in the browsers. A huge amount of stuff in the browsers is still at the "community draft" stage, and yes, you have to implement that, too.
Microsoft simply gave up, forked Chromium... And they still can't keep up: https://web-confluence.appspot.com/#!/confluence
Even if it's overblown by, say, three times, that's still over thirty million words.
That contributes to "word bloat", but it's not necessarily a bad thing. Picking the right metrics is not always that easy!
This is true for HTML5 which defined full browser behaviour, including things like improperly closed and improperly nested tags.
Many, many other specs? Not so much. Especially the crap that Chrome has been pumping out the past several years.
> WHATWG [is] far more relevant than the W3C nowadays.
Which is arguably part of the problem.
Thanks for answer!