Cyclic dependencies are evil (2013)
fsharpforfunandprofit.com
fsharpforfunandprofit.com
Consider briefly the domain model of a bank. Customers rely on Accounts (I.e. have one or many). Accounts rely on Customers (i.e. have one or many). This is a simple kind of test you can apply to your language and architecture to see if you have what it takes to attack the problem domain. Most approaches lauded on HN would be a messy clusterfuck when attempting to deal with this. Now, if I can simply call CustomerService from AccountService and vice versa, there is no frustration anymore. This is the power of reflection. It certainly has other caveats, but there are advantages when it is used responsibly.
If you want to understand why functional-only business applications are not taking the world by storm, this is the reason. If it weren't for a few "messy" concepts like reflection, we would never get anything done. Having 1 rigid graph of functions and a global ordering of how we have to compile these things... My co workers would laugh me off the conference call if I proposed these constraints be imposed upon us today.
Not really rocket science but some times your model is telling you something if you’re willing to listen.
(see page 2)
E. F. Codd. 1970. A relational model of data for large shared data banks. Commun. ACM 13, 6 (June 1970), 377–387. DOI:https://doi.org/10.1145/362384.362685
Over the years I've found that these transposition problems are a huge drag on programs, and once you recognize them you start seeing them all over the place. They make simple and obvious relational tasks into dozens of lines of nested data structure traversal and manipulation.
The notion of transposition comes from the array programming world, where a multi-dimensional array/tensor can be seen as a tree structure, whose levels may be easily reordered by transposition. I sometimes use numpy arrays for general-purpose programming; it can be very powerful with a well-chosen representation. Unfortunately array concepts basically require the data to be rectangular -- `x[i]` has the same shape as `x[j]` -- which can be hard to adapt to general problems.
My dream then is a general-purpose data structure that takes the best of both the relational and array-programming worlds: the freedom and flatness of relational tables, enriched with tacit ideas like broadcasting that make array programming so ergonomic.
This type of relation shows up quite a bit, and either leads to application-level workarounds, or having to add many duplicate tables for the same "thing".
std::relation<Foo, Bar> foobars;
...
auto& bars = foobars.forward(foo);
...
auto& foos = foobars.backward(bar);
Edit: Boost apparently has it, see e.g. https://stackoverflow.com/questions/1128144/c-maps-for-many-...The problem is that this only lets you get customers from accounts, and not vice-versa.
And if you add a Map<Customer, Set<Account>> to the Bank, you now have to ensure the mappings are synchronized at all times. Which, to be fair, is very straightforward as long as Bank is immutable, but still kind of annoying.
But considering how incredibly common many-to-many relationships are in business logic and especially in SQL, I'm constantly surprised that there isn't a de facto standard, efficient data structure to represent them. That speaks to how entrenched cyclic relationships are in software engineering, I suppose.
Many-to-many relationships in the purest form is described as records of ItemAId/ItemBId pairs. `List<{ account: Account, customer: Customer}>`. There isn't any cyclical relationship in this form.
From `List<{ account: Account, customer: Customer}>` one can derive `Map<Customer, Set<Account>>` and `Map<Account, Set<Customer>>`. Bank can have copy of both mappings and use them as double index, synchronizing both mappings as operations goes by, while at the same synchronizing both mappings to the in-storage form. Or not!
We're talking mostly about data-modelling in a programming language, which applies to the in-memory form. `Customer { accounts: Set<Account> }` and A `Account { customers: Set<Customer> }` is its in-memory form, a layer of indirection from its actual form, the in-storage `List<{ account: Account, customer: Customer}>`.
If informations of Accounts, Customers, and the many-to-many relationship of both are laid flat in its purest, source-of-truth form, the in-storage database entries, why are people suggesting that there is/should be a cyclical relationship in its representative form, the in-memory objects?
Insightful. Once you learn to see dependencies as implicitly revealing abstract data types you haven't yet teased out into existence, you start seeing all kinds of places where behavior that ought to be explicit and handled by an abstraction is smeared across multiple other concepts.
There is nothing in this contrived example that F# or Haskell (gasp, pure functional) wouldn’t handle with ease. It just would be modeled differently from traditional langs. I’d argue the alternative modeling forced by FP langs would be superior.
Then I'd like to see that. The other example of the functional someone gave more above, was not convincing to me. Much more complicated then the straight-forward simple, but evil cycle approach.
Every engineer I've ever met who has been doing it for more than a couple years universally decries circular dependencies. Every one of them has come to that opinion via hard won experience dealing with real world problems they encountered. And yet comments like yours reveal there are other engineers working in places where apparently the opposite is true. Whenever I see that I wonder how this sort of self sorting manages to occur.
If a more experienced engineer does, it’s generally considered a necessary evil, but not a choice taken lightly.
They may not realize they have them - at runtime, in the object graph, possibly indirectly. There's little difference in principle between cycles in "static" code vs. runtime state, but we often can't express them in the former, because our languages don't specify the concept of dependency at enough granularity (and some rely on a linear compilation pass).
To me both are technically true, reality is messy, and we hide the mess under imaginary constructs.
If you've working with Java at Google (Guice) makes you learn to hate hiding circular dependencies behind dependency injection, since your binary gets injected by some huge tree of thousands of dependencies all that can create errors or interfere with each other. And without static checking trying to reason and fix those issues becomes really hard.
I strongly prefer using the languages actual type system to create circular dependencies that you can inspect using well known tools over that any day, better have a problem you can see than hide the same problem to satisfy some tools requirement.
I may sound very snarky, but it really feels like an over engineered mess and I suspect the reason they need all these dependency tricks is the fact that it’s over engineered, not necessarily that the problem they’re solving is so unique that they need circular dependencies.
And what has reflection to do with this?
You "wire up" the implementations at runtime, using reflection.
You get the other kind of interface when you want to loosely-couple your code, and so you define interfaces for many classes. Often, there is only a single class that implements the interface in your project, though there may be mock implementations in your test code. Even though there is a single "real" implementation, the IDE can't/won't jump to that implementation in the same way it won't in the first case. This is frustrating though, because it would have worked if you hadn't extracted the interface for improved testability.
The fact our solution has an interface for just about everything for testing reasons doesn’t slow me down even in the slightest when I’m looking through the code.
It certainly has its cost in additional complexity through indirection, but it's better than creating cyclic dependencies or giant balls of mud.
https://en.m.wikipedia.org/wiki/Dependency_inversion_princip...
If you can do dependency inversion without reflection, more power to you :-) We can't do classpath scanning in the project I'm working on because of the size of the classpath, and compile time configuration using direct imports would introduce cycles, so reflection it is for us, in one form or another.
class Customer implements ICustomer { ... } // Customer Class can be created after Account above
createAccount(new Customer()) // Use of both, with dependency injection
Scanning for classes dynamically using reflection has nothing to do with above. And you certainly don't need any xml or framework to do it either.
First of all, both data and functions are first class concepts in the functional paradigm and are not glued together as classes. So you’re not modeling your domain as „account service“ and „customer service“, but rather have data representation of these concepts and functions that query/reduce etc on those. In your example that would be a relational model, described as sets of tuples/maps, and relational functions do derive new data. There is no need for circular dependence, because your operations are generic over relational data.
I agree with the notion that the purely functional approach doesn’t dominate, for good reason. But more and more languages are incorporating FP concepts since about a decade or longer, at least the basic building blocks like closures, immutability and function composition have gained massive traction and eliminate the need for many of the complex patterns in traditional, class based OO.
> Customers rely on Accounts (I.e. have one or many).
I'm giving a domain relevant counter example to that premise.
And interestingly enough, that leads you to a clean way of modeling it without the cyclic dependency as you've just done.
There is always state somewhere, the question is where it should be.
Global state is the input to the program, and a new global state is the output. Given identical inputs, it will always produce the same output.
This is the entire principle behind ACID, and mutation of global state is a common approach for functional programmers regardless of whether they they are using a DB or not.
If a DB were not fully transactional, or if your queries of the DB are not deterministic themselves, then we’ve moved out of functional territory.
That said, I don't see what circular dependencies have to do with reflection, much less a functional style? For a trivial example, Rust lacks reflection but the following (functional-ish) code works just fine:
mod ModA {
use crate::ModB::B;
pub struct A {
ref_to_b: Option<Box<B>>
}
pub fn into_b(a: A) -> Option<Box<B>> { a.ref_to_b }
}
mod ModB {
use crate::ModA::A;
pub struct B {
ref_to_a: Option<Box<A>>
}
pub fn into_a(b: B) -> Option<Box<A>> { b.ref_to_a }
}
Am I missing something?- The definition of Account
- The definition of Customer
- A ledger consisting of: List of Account, List of Customer, List of Account-Customer relations
A clerk works on the records on the ledger with one hand and one pen, a metaphor for the service process working on the data in the database.
An act, such as money transfer, would be described as a function.
A customer creating a new account for himself is written as `fn createAccount(Customer, NewAccountData)` because from the perspective of the clerk/bank manager/service the customer, newAccountData, and the existing data in the ledger as objects which the clerk/bank manager/service must move around in a precise way.
The module which has 'fn createAccount` depends on the types `Account` and `Customer`.
In english it roughly sounds like,
"the success of writing the rule of creating an account for a customer depends on knowing the definition of Account and Customer."
The function is not written as `customer.accountService.createAccount(newAccountData)`, because the clerk doesn't schizophrenically pretend to be the customer and create an account for himself. The clerk just receive request from the actual customer, writes new account entry and customer-account-relation entry into the ledger, that's it.
There's simply no need to call CustomerService from AccountService and vice versa. There's no need for reflections because data types are all available at compile time.
I believe that most OO implementations would read
accountService.createAccount(customer, newAccountData)
Care to elaborate if that was the main point of the clerk's schizophrenia criticism? Or if I'm misinterpreting just call me dumb, haha.But my main point is that it should not be written that way at all.
Semantically, if we must include the subject, it should be
`theService.createAccount(Customer, NewAccountData)`
Or replace theService with CustomerAccountService or whatever.
Writing that way, with functions depending on types, avoid getting tangled from the so-called account-customer cyclical dependencies. Because account and customer don't need to know each other until a function needs to know both of them. There's no such thing as `customer.getAccounts()` because in the end the query would roughly look like `getAccountsWhereCustomerIs(CustomerId)`.
You're not dumb. It's just me mis-writing due to the fact that I'm writing this at 4 in the morning lol. Or maybe I'm misinterpreted parent comment and that confuses you. In that case, I'm the dumb.
A customer has an account. The customer is the domain root. Hence there would be no accounts if there were no customers.
A customer does need to to hve an account. An account must have a customer.
It's not cyclic in any way
Alternative model: the Customer relation couples a Person to an Account.
I think that’s a better model, as it allows Person P to be a Customer at bank B but not at Bank C, Person Q to be a Customer at both B and C, etc.
It also allows you to model Persons before you model Accounts, breaking the circularity, allows Companies to become Customers, etc.
Another way to think about this is, how would you do this in C++? You have header files, and they shouldn't be cyclic. Who #includes who? Well in this case it seems to me AccountService.hpp should #include CustomerService.hpp, but not the reverse. Maybe CustomerService.hpp would #include AccountID.hpp? Whatever I imagine I just don't see why you'd need cyclic #includes.
AccountMgr has many Account(s). Ditto for CustomerMgr. Interact with both to taste.
One pattern this nudges you towards is having a Domain.fs at the top of each project. Since everything in your project wants to talk in terms of the data types you've modelled, it makes sense to put them all in the same place at the top. Moreover, once you've done that, you're guided towards having your data types really being "dumb" - algebraic data types only - because to put actual behaviour into them is harder when all you have is definitions of data types (and no associated modules). You certainly can write OOsagne using the Domain.fs pattern, but it's much harder and the code really smells when you do (because the domain file gets super long).
The upshot is that your domain model appears explicitly at the start of the project, which is a big win for anyone who comes into the project and needs to learn quickly what's going on. By contrast, nearly all the C# I've ever come across has the domain spread across many files and all mixed up with implementation details.
This is certainly not the only way to write F# - I've written projects which have the domain spread across multiple files - but it's one nice way to handle a small-to-medium-sized project.
The only downside to hexagonal architecture is that you start to see how mediocre most architectures are, and how frameworks often cap how good your architecture can be by virtue of their structure. Frameworks tend to resist being put on boxes, and hexagonal architecture tries to put all external dependencies in boxes.
Compilers for languages like C# and Java can deal with (seemingly) circular dependencies within a set of classes to compile, because when A uses B they only have to consider the interface (type signature) of B, not its implementation. (Cyclic dependencies between modules are still a problem, in particular for build systems.)
I'm not really familiar with F# (so someone please correct me if necessary), but in languages that use type interference to determine function signatures, one consequence is that the interface (type signature) of the function then depends on its implementation, so you can't as easily invert the dependency as when you have explicitly declared type signatures.
On a related note, lately I have been looking into some biochemistry topics and their relation to computer science. It seems that cyclic dependencies are possibly a requirement for life, which makes faithful simulations of biochemical processes an interesting challenge.
But Lua and JavaScript satisfy their own niche just fine without having popular runtimes written in themselves.
Those that do, however, are immensely beautiful and there is more to be learned from them.
If you can bootstrap it, and you can, then it's not really a hard, unbreakable cycle, right? It's just an option that's quicker than starting from tcc or whatever every time.
https://bootstrappable.org/ https://bootstrapping.miraheze.org/wiki/
If nothing else, it encourages (somewhat) common ways of structuring projects and avoids a lot of bike shedding about separation of concerns and project structure I encounter in C#.
What I really wanted to say was: "F# lets us be lazy. You write way less code, you have fewer things to keep track of, you need fewer abstractions, you can write everything in the most straightforward way possible and then extend only when necessary because refactoring is super safe."
What I _actually_ said was: "F# is an extremely rigid language. It will not let junior developers take shortcuts, because you must write your code properly or it won't compile. Your classes and functions must be written in the correct hierarchical order, you must cover every corner case. I know it's going to be stressful or boring to be forced to do things right every time and never be able to just push some sloppy cowboy fix, but it's a sacrifice we're willing to make :)"
Why is this bad? It makes it impossible in many situations to locate your public API at the top level. Let's say you have a Widget interface and APIs to do things with those Widgets. Eventually your library gets complex enough that you want to separate functionality into subdirs. Now you have a problem: You can't access the definition for Widget from the subdir because importing the top level would create a circular import!
The only way to get around this problem is to have no code at the top level, and put your public API in a subdir. Just be sure to do this when first starting your library or you'll have to break things.
What would be interesting would be to organize modules into named layers--instead of just explicit 1-by-1 module dependencies (i.e. edges). I have not seen a language do that yet.
Another interesting thing that I haven't seen touched on in more mainstream languages is the idea of bidirectional interfaces. The somewhat DSL-like NesC language had this, primarily driven by the need to write device drivers that serviced interrupts (and events).
Correct me if I'm wrong, but I think you're describing namespaces here
It helps me understand what's included if I can see a nice list of all the dependencies. It helps me narrow down where the problem is if I only have to address dependencies I know have already loaded.
I learned to program in the 80's however, and it probably is a bit "old school".
mod widget
struct Widget
fn show(Widget) { ... }
fn join(Widget) { ... }
fn calc_box(Widget) { ... }
and one function is getting very long and complicated so I want to spin it out into its own module mod widget mod calc_box
import calc_box from widget import Widget
struct Widget fn calc_box(Widget) { ... }
fn show(Widget) { ... }
fn join(Widget) { ... }
The reason the article gives that this is "evil" is that there is no hierarchy between the modules; the cycle is really just one big supermodule. That's true... but that was basically the point. Where is the "evil" here?The section in Lakos on how to identify and remove circular dependencies was worth the price of the book.
https://erock.io/scaling-js-codebase-multiple-platforms/
In the article I argue that circular dependencies are a good thing because they uncover poor code organization.
https://wiki.debian.org/DebianBootstrap
Also the Guix and Bootstrappable Builds folks are working on this too, at a different layer:
https://bootstrappable.org/ https://bootstrapping.miraheze.org/wiki/
Tightly coupled dependencies are bad. Cyclic or not.
And if you do mange to to a cyclic dependency you are not thinking about your design property. Imo