L1D latency with a pointer is mostly 4 cycles. Not sure, but I think having 1024 entries would increase that to 5 cycles.
Increasing cache line size from 64 bytes to 128 bytes would require more memory (DDRx SDRAM) bandwidth, because CPU needs to always fill a whole 64 byte - 128 byte with this change - cache line on a memory load. Even if you want just one byte. Wasted bandwidth for non-streaming workloads would be increased significantly. This change would force also L2 to have 128 byte cache lines, etc. LFBs and other buffers would of course also need to accommodate this.
Skylake will be able to process 64 bytes (512 bits) with a single SIMD instruction, with AVX-512. Perhaps Intel will need to soon increase cache line size, if 1024 bits wide SIMD instructions are desired.
SIMD gather instructions might be a hint that in the future memory loads will function in a different fashion. Maybe cache lines won't anymore be the smallest amount of data a CPU memory subsystem can deal with?