Three Types of Data
brandons.me
brandons.me
It describes these categories in a lot more detail and the equivalents are:
Constants => User specifications
State => Essential state
Cached values => Accidental state
It’s a good model, and I think gets a lot of things right. But there are definitely nuances.
This would be a good caveat to add to the original post. :)
----------------------
Edit: I wrote this example in response to a comment which has since been deleted, so I'll post it here instead
Let's say your program stores the positions of two entities that can change arbitrarily over time:
let pos1 = { x: 0, y: 1, z: 2 };
let pos2 = { x: 1, y: 3, z: 0 };
And you also want to work with the distance between them: let distance = Math.sqrt(
(pos2.x - pos1.x) * (pos2.x - pos1.x) +
(pos2.y - pos1.y) * (pos2.y - pos1.y) +
(pos2.z - pos1.z) * (pos2.z - pos1.z));
When do you do this computation?If "distance" is thought of like any other state, it's unclear, and it's easy for it to get out of sync with the values it's derived from. Maybe you have some sort of core update or rendering phase and you re-compute it there. Maybe you try and re-compute it every time one of the two values gets modified, either by constraining their modification within methods or by somehow observing their changes. Maybe you have a data structure that allows you to easily compare them to the values the previous computation came from. Deciding which of these strategies to take is non-trivial, but you can simplify the question a little bit by seeing "distance" as not being a normal part of state.
If you pull it out into a pure function:
function getDistance(a, b) {
return Math.sqrt(
(pos2.x - pos1.x) * (pos2.x - pos1.x) +
(pos2.y - pos1.y) * (pos2.y - pos1.y) +
(pos2.z - pos1.z) * (pos2.z - pos1.z));
}
then "updating" it becomes a singular, clear action: let distance;
function updateDistance() {
distance = getDistance(pos1, pos2);
}
And then when it gets computed becomes an independent question from how it gets computed. You can update it eagerly, or lazily, or implicitly. You can use comparisons, or observables, or whatever.The benefit becomes more clear when the value isn't a simple number, but a whole object or object graph. By making it immutable, you have much more leeway when it comes to "refreshing" it, because you can guarantee you won't be losing any meaningful information.
None of these ideas are especially novel or profound, but as a mental framework they've shed a whole lot of clarity for me in my work over the last year or two.
Also the distance function can often be replaced by a distance to the power of two (since then the cost of running sqrt is not paid). This is often the case in my work - rearchitecting the application in such a way that you don’t need caching in the first place and being explicit.
But what about where distance is being used (read)? How do you know if updateDistance() got called already. It seems with this approach it would be important to still have a function for accessing it. But then that functions would need access to some sort of state that says whether or not the data is stale. Something like this:
let isDistanceStale = false;
let distance = 0;
let pos1 = { x: 0, y: 0, z: 0 };
let pos2 = { x: 0, y: 0, z: 0 };
function updatePos1(newPos1) {
pos1 = newPos1;
isDistanceStale = true;
}
function updatePos2(newPos2) {
pos2 = newPos2;
isDistanceStale = true;
}
function getDistance() {
if (isDistanceStale) {
distance = calculateDistance(pos1, pos2);
isDistanceStale = false;
return distance;
} else {
return distance;
}
}
function calculateDistance(a, b) {
return Math.sqrt(
(pos2.x - pos1.x) * (pos2.x - pos1.x) +
(pos2.y - pos1.y) * (pos2.y - pos1.y) +
(pos2.z - pos1.z) * (pos2.z - pos1.z));
}
Gross. This is why it's better not to made these kind of optimizations unless there really is a performance issue. Otherwise, just let it get computed every time or compute it every time pos1 or pos2 changes. No need to do lazy evaluation. If you really want that, there are languages that have it built in. Otherwise you'll just be fighting the language.The troubles begin when working directly in the data models, leading to coupling, dependencies, side-effects and narrow perspectives on how to accomplish better designs. In beginning it seems more powerful, until enough complexity creep attained to warrant headache examination.
I don't know if there's any other tricky implementation part, but from what I've seen you could simply define something like this:
struct Human:
name String
birth Date
derived age Integer
function derive age:
return (time.Now() - self.birth).Years().Floor()
And the compiler should have everything it needs. Sure, this example is pretty annoying because time changes all the time, so it doesn't look like you can cache much with such a naive approach, but I suck at examples (surely people working on languages could come up with more interesting approaches, like adding ways to schedule the cached value to be preferably kept until X time later or whatever).Others have stated that: "There are only two hard things in Computer Science: cache invalidation and naming things"
"There are two hard problems in Computer Science: cache invalidation, naming things, and off-by-one errors"
- Constant data
- Long-term configuration
- Medium-term configuration
- Client/account data
- Business transactions
- Transitory processing work