If you're using Cassandra and you're also not a fool, it's probably because you actually need a lot of scale, so you'll generally do the denormalization up front and you will denormalize everything. If you're writing an event, you will write it several places. For example, suppose you're WalMart recording sales. You might write to: store transactions by store/year/day/hour (the "master" record insofar as you have one), user transactions by user/year/month, product purchases by manufacturer/year/day/hour... When you write the transaction, you write to all of these locations.
Each of these "by X" keys is a shard. Each can be located on a different set of ~3 machines (the number is configurable). Querying involves getting a copy of the ring topology, computing which integer shard-ID the key maps to, figuring out which machines in the ring own that integer, and then asking the machine for a whole bucketful of data, which should be a superset of what you're actually looking for. For something like a user's transactions you'll want to have basically everything there at once, so loading the "order history" page for the past month might be a single query that just returns a report: no joins at all, very fast, super scalable. Other lookup strategies might ask for a range within that bucket (the data within the bucket can be ordered by a single key; often this is a timestamp or time-based UUID). Anything that isn't a simple query of a few buckets like this has to be a map-reduce job and will be slow.
All of this is pain. You should generally not invite pain into your organization. However, if pain has already found you, something like this may be the least painful option.