Beyond Source Maps
fitzgeraldnick.com
fitzgeraldnick.com
In practice they are only useful for less-aggressive minifiers and 'basically sugar' compilers like coffee script. Trying to use them for something like emscripten or JSIL output is a mess.
P.S. Variables live in scopes, so attaching variable information to line mappings is pretty gross... I suppose if it works it works, so better than nothing.
The 'names' field is a list of all original symbols in the program. Each mapping is 1, 4 or 5 fields long. If 5 fields long, the 5th field is an index into this names field, e.g.
Or see a more readable description on html5rocks http://www.html5rocks.com/en/tutorials/developertools/source...
It'd be great if they revised the spec to not use such opaque terminology, and directly describe what the 'names' list is used for. The only reference I can find that explains this is 'If present, the zero-based index into the “names” list associated with this segment', which I only understand how that you said what it means.
Since one of your other comments made it unclear: Does this apply to all variables, or only global variables?
Edit: I take that back. I just looked at the Blink source code, as far as I can tell, it throws away the name index. Therefore, Chrome Dev Tools won't deobfuscate symbols, but Google server-side JS deobfuscation framework will, which is a pitty. I guess it is time for a Chrome Dev Tools patch. :)
The misleading presentation of source map features is problematic because I think it leads end-users of transpilers to believe that it will solve their problems and ask for it as a feature addition, when in reality it won't do them much good.
I do think the people who specced source maps were correct to start with the narrowest possible feature set, though, even if the resulting spec seems strangely micro-optimized in ways that make it much harder to generate than it would be otherwise.
SourceMaps originally came out of a need for better stack trace deobfuscation I think, use inside of debuggers came way later.
I am totally onboard with fixing the deficiencies of the current format to make debugging better. Now that neither Firefox nor Chrome support the JSVM hooks we needed to make GWT DevMode work, we are stuck with sourcemaps for debugging, and Java programmers are not pleased with having to look at mangled symbol names in their scopes and watches.
I don't really care what the format of the file is, as long as it preserves a number of features:
1) can still ship and debug 'stripped' binaries, doing deobfuscation offline on the server 2) memory and network efficient. Don't make my JS debugger freeze loading these things in, or each up hundreds of megabytes 3) supports cascading and merging. We have hybrid apps at Google, that is, apps that are compiled with multiple transpilers and linked together. We 'merge' sourcemaps from these together in a unified sourcemap. So whatever the future format is, it should support efficient composition.
I think extending what we already have instead of reinventing the format is smart. All of your concerns listed at the end are things I care about too.
However, I don't believe browser debuggers handle this properly yet.
The way to understand the spec is that they wanted to enable breakpoints. Breakpoints in debuggers such as Chrome are line-oriented. If the debugger can find the corresponding position in JavaScript and set a breakpoint, the sourcemap did its job.
However, there are other issues with sourcemapped breakpoints: https://code.google.com/p/chromium/issues/detail?id=333569
This critically used in a large number of Google products to huge extent. We globally trap all unhandled exceptions, and deobfuscate those exceptions using sourcemaps before they go into an automated bug triage system.
The issue of handling locals and properties can be fixed without overhauling the format. Storing the debugging information in the JS itself would remove the ability to ship stripped binaries and be a non-starter for production apps.
ClojureScript, for instance, supports dynamic bindings.
For example, for each symbol we record, we can simply introduce something called a 'scope identifier' and then define a language specific mechanism by which debuggers can compute scope identifiers.
Consider the following problem for GWT:
class Foo { String hello; }
class Bar { String world; }
Both of these can be compiled and obfuscated an an object with field 'a'. So when encountering a heap reference to an object with field 'a', does it deobfuscate to 'hello' or deobfuscate to 'world'?
We have many kinds of scopes: global scope, object scope, function scope, etc. In this case, its an object scope. If we uniquely identify each scope, we can perform the mapping as long as there's a way to compute runtime type, e.g.
Store 'a' as '2:a' (scope identifier '2') Put a mapping in the sourcemap that 'Foo' maps to '2', and now we can introspect the heap reference, determine it is 'Foo', map it to '2', and properly display it as 'hello'
Different languages may have different ways of setting up classes and defining runtime types. GWT for example, does not rely on JS constructors, but instead hangs a classId off of the prototype. I think Dart does something different. In any case, it is a simple extension to the existing format and backwards compatible. The only runtime API needed is a function getScopeIdForReference(obj) which takes a compiled object and gives its scope.
The AST approach I think has several problems:
1) if the annotations are present in the shipped obfuscated code, it leaks information that people don't want. Many people do not want to leak the source filenames or directory structure of production apps.
2) it increases the download size for consumers who are not developers and don't need the maps
3) if written out as a separate file that co-exists with a stripped binary, it increases memory requirements of the debugger.
The way Google deploys Gmail, for example, we use sourcemaps in production as well as development. When exceptions are thrown on the client, the stack traces are reported back to servers, deobfuscated via sourcemaps, and put into a triage system. A person looking at a bug report sees a deobfuscated trace, but the consumers aren't burned with having to download the sourcemaps in annotated js.
I think the idea of defining an actual sourcemap API for the debugger (getDisplayValues, getLocals, eval, etc) has a lot of value, because these function are source-language dependent, so having GWT/Dart/Emscripten/CoffeeScript/etc generate them is useful.
But I don't think annotating the actual JS, or switching the sourcemap format to an AST is a big win. Yes, it is a simplification, but it also trades off other useful properties of sourcemaps.
You would be able to continue using the approach you describe for GMail. There is no reason to publicly serve the debugging information unless you want to.
> 2) it increases the download size for consumers who are not developers and don't need the maps
No, the debugging information would still be an auxiliary file like it is now.
3) if written out as a separate file that co-exists with a stripped binary, it increases memory requirements of the debugger.
Debuggers need to keep the source map around now, anyways; this is no different.
Regarding the scope identifiers: it would solve scoping for the most part, but it fails to handle the last three requirements I defined:
- It should provide a way for the JavaScript debugger to display values in a meaningful way.
- It should optionally provide an eval capability, for use from a REPL, watch expression, or conditional breakpoint.
- The format should be future-extensible. That is, when SourceMap.next v2 rolls out, any SourceMap.next v1 consumer should still be able to parse and use instances of SourceMap.next v2 (although without the new features, of course).
Inspecting values would still be a pain, and you wouldn't have watch expressions or conditional breakpoints, etc.
Also, much of the value in what I described in the blog post is the future-extensible format. It allows us to fix our mistakes post facto.