The internal representation depends on the tool and application. Here are a few examples:
* For deeper analysis, we often use the same internal data structures that the compiler itself uses. This includes data structures that capture AST and type information. In the case of some languages, the compiler doesn't capture enough information (e.g., dynamically typed languages with no static typing). In those cases, the data structures are often enhanced/similar-but-not-quite-the-same as what the interpreter/compiler uses.
* In "shallower" analysis, you can use a language-agnostic parser to generate a generic AST. There's various techniques for doing so in a language-independent fashion, depending on what tradeoffs you want to make with speed, accuracy, and generality. The tree-sitter library is one such example. You can also use tools like CTAGS + textual heuristics in the extremely shallow case.
* Then, there's the Kythe approach, which constructs a taxonomy of source code graph nodes and edges that aims to capture every possible relationship in every language. Each language requires an integration to convert from the source code in that language to the Kythe schema; that integration can use either the "shallow" or "deep" approaches described above, and the right answer depends again on your tradeoff of speed vs. accuracy.
As for combining dynamic information like profiling and runtime info, this is something we've been thinking about for awhile. We think there's plenty of low-hanging fruit that combines static information (types, AST, source code) with runtime info (stack traces, profiling info). We've a couple of efforts in motion toward this end. One is the Sourcegraph extension API (https://sourcegraph.com/extensions), which allows injection of third-party info, often from runtime tools, into wherever you view code, whether it's in your editor, on GitHub, or on Sourcegraph. Another is allowing a user to plugin stacktrace data to heuristically select which jump-to-def/find-references options are most likely to be useful given what your currently debugging.
I'd love to hear what you're thinking of—care to share your ideas in this space?