One memory this project brought to mind for me was a hack I came across which allowed simultaneously running DOS 3.3 & ProDOS on a 128k Apple II, giving each 64k (well, a little less due to overhead) & a way to switch between the two with a simple command. Two programs couldn't run at once, but one could step between the two OSes to run programs made for each pretty seamlessly. If this sort of thing was possible on basic consumer hardware, ten or twenty years of development would have led to many far more interesting & useful things.
If you are absolutely limited to 6502 DIP chips, there would probably be more prevalence of large mainframe systems and single 6502-based "terminals"/"thin clients". The mainframes could use systems similar to the Transputer or the Connection Machine to use large amounts of (comparatively) low-power processors to make a single, more powerful computer. They both used custom processors, with the Connection Machine in the early 80s and the Transputer in the late 70s. You could probably reasonably easily create a "graphics card" style system, comprised of many 6502 cores in a SIMD configuration.
I don't know how easy it would be to implement wifi or ethernet with only 6502 chips, so communications with the mainframe might be quite slow
It could be useful for some sort of minicomputer for business applications.
And if you have a 16-bit CPU, you can do all kinds of silly stuff; for instance, you can have 4 16-bit MSRs, let's call them BANK0–BANK3, that would be selected by the two upper bits of a 16-bit address, and would provide top 16 bits for the bus, while the lower 14-bits would come from the original address. That already gives you 30 bits for 1 GiB of addressable physical memory (and having 4 banks available at the same time instead of just 2 is way more comfortable) and nothing stops you from adding yet another 4 16-bit registers BANK0_TOP–BANK3_TOP, to serve as even higher 16 bits of the total address — that'd give you 16+16+14 = 46 bit of physical address (64 TiB) which is only slightly less than what x64 used to give you for many years (48 bits, 256 TiB).
Even 4MB would take you hours to load from floppies with a 6502.
Terabytes with a 68000 would also be impractical.
Depends on your clock. Also, you could use some dedicated hardware, like a DMA controller e.g. 8257, or 8237. From 8257's datasheet:
Speed
The 8257 uses four clock cycles to transfer byte of
data. No cycles are lost in the master to master transfer
maximizing bus efficiency. 2MHz clock input will
allow the 8257 to transfer at rate of 500K bytes/second.
and I recall 8237 could do even better, if wired and programmed properly.Processing terabytes with a single CPU was impractical, but you could in theory connect it.
Due to the lack of support hardware in the C64 (no hardware RAM bank switching/MMU) this memory is not bank switched and then directly addressable by the CPU, it's copied on request by DMA into actual system RAM. But in some sense, a C64 with a 16 MiB REU is a 6502 with 16 MiB RAM.
But yeah, you want CPU addressable RAM with real bank switching. You couldn't really do 16 MiB, you wouldn't want to bank switch the entire 64 KiB memory space. The Commander X16 (a modern hobbyist 6502 computer) supports up to 2 MiB by having hardware capable of switching 256 banks into an 8 KiB window (2 MiB/256 banks = 8 KiB).
Let's say you design something with 32 KiB pages instead -- that seems kind of plausible, depending on what the system does -- you could then do 256*32 = 8 MiB and still have 32 KiB of non-paged memory space available. I think this looks like just about the maximum you would want to do without the code or hardware getting too hairy.
See The 8-bit Guy regarding what the world would be like if we were still limited to vacuum tubes: https://www.youtube.com/watch?v=mEpnRM97ACQ (video)