This is kind of a fake answer to the problem.
First off, using less memory is an effective way to reduce cache misses: if you shrink memory by ⅓, that allows you 50% more objects in the same cache size. And this applies to anything--it's the only way to reduce cache misses that is universal. So saying that it's not solving the "real" problem is really a spit-take, because it's a pretty effective way of solving that "real" problem.
Suppose you considered cases where a smart ordering could avoid hitting unused cache lines. If a struct is larger than a cache line, it's possible to put co-used values on one cache line and avoid bringing in the other cache lines. But this kind of optimization isn't going to work unless the struct is cache-aligned to begin with--otherwise, your clever ordering is only going to sometimes work and sometimes potentially cause unnecessary multiple cache lines to need to be brought in. As to whether or not cacheline-alignment is a good idea, well, the extra padding will increase memory usage (see point #1), and the potential benefit is going to be limited by how hot or cold field accesses actually are.
The other case that comes to mind is false-sharing, which is definitely a real concern. Except, we're talking about reordering struct fields, which means it's false sharing within fields of a struct, and that's a much smaller subset of where false sharing actually occurs--false sharing tends to be more of an issue when you have an array of objects, and you need to make the struct element a multiple of cacheline size to avoid it. The only reasonable cases I can think of off the top of my head are going to involve structs which have intrusive atomic reference counting or some sort of intrusive lock in them--and you can solve both of those cases by making large cacheline-sized versions of those structs that prevent any fields of the outer struct from being stuck on the same cache lines as those data structures.
So I rather expect that it is very possible to have a field reordering algorithm that would improve cache misses in all the obvious cases (as I mentioned in point #1) while not preventing the user from having sufficient control to optimize for minimizing cache misses in the rarer cases in the subsequent point.