He's proposing to skip the whole computation of the parts[] array and within the kernel, replace
for (int i = 0; i < parts[bx].length; i++) { // bx is the global thread index
char c = buffer[parts[bx].offset-buffer_offset + i];
by something like long long split_size = size / num_parts;
long long offset = bx * split_size;
while (buffer[offset++] != '\n')
;
for (int i = 0; buffer[offset+i] != '\n'; i++) {
char c = buffer[offset + i];
(ignoring how to deal with buffer_offset for now).