1) foo.lang: class foo { func bar(a) { return b } }
2) baz.lang: import foo && print( foo.bar(1) )
Generally `foo` needs to have unit tests (self-tests), but `baz` needs to have "acceptance tests" for `foo` because `baz` imports `foo` (external validation / fit-for-purpose). Blah blah "law of demeter" aka: self.do_foo_bar() { foo.bar(1) }.
I've come to the mental model that code should test itself, and (some) code should test its dependencies. Code fundamentally can't test code that uses it, but can attempt to be "well-behaved" and share (via documentation or checked interface) how it can be / should be used.
eg: `foo_test.lang: assert( foo.bar(1) == 42 );`
But if `baz` calls `foo.bar(2)` or `foo.bar(0.3)` or `foo.bar("four")`, then `foo` cannot possibly be responsible for how `baz` interoperates with it in all cases.
"Software has no way to express its own compatibility with other functions or dependencies".
I've heard that at a certain point all internal behavior can turn out to "bake" and be depended on by an external user. An extreme example is `*.append(...)` having O(n), O(log n), O(n^2) performance characteristics in either memory or speed. Some (ab-)use means that you may get a bug report b/c your performance characteristics changed in a way that the _user_ of your code didn't expect, and I struggle to see how anything other than snapshotting the world [never mind network-call dependencies] would ever "fix" things?
I think it's an impossible problem:
1) code can test itself
2) by definition, code can't [exhaustively] test how it is used by "others"
3) therefore: any changes to dependent code are potentially "breaking" changes from the perspective of the user of the dependent code (even bugfixes!)
4) a generalized "really good" outcome is possible (eg: imagine something like TypeScript but for dependency management), but once you loop around to point #1 again... you either exactly match like a puzzle piece (eg: frozen dependency), or there's some sort of wiggle-room where any sort of change to the dependency has the _potential_ to break the user, or there is guaranteed slack/wastage in between the two dependencies (ie: size, performance, memory, operating range, etc) where changes can be relatively safely made.
5) therefore: code that uses a dependency should have acceptance tests (fit-for-purpose) of how the dependency behaves.
We're relying on version numbers and transitive dependencies at the moment, but I can't imagine anything other than "freeze everything" being a viable, generalized solution?