for i,v in enumerate(open(f)): pass
now 'i' is the number of lines in your file and 'v' is the last line. This is also blazingly fast, especially useful when you have ~several GB large log files to parse.
Great article!
for i,v in enumerate(open(f)): pass
now 'i' is the number of lines in your file and 'v' is the last line. This is also blazingly fast, especially useful when you have ~several GB large log files to parse.
Great article!
tail -n 1 $f
Tail seeks backwards so it will only read the one line. Of course, this won't give you a line count.EDIT: I haven't tested it but you might be interested in this implementation of tail in python: http://stackoverflow.com/a/136368
Edit: Of course I misspoke, yes tail is much faster for getting the last line of the file! I meant for getting the line count the loop methods is typically just ~5% slower than wc on sufficiently large files.
In [73]: cProfile.run('for i,v in enumerate(open(fn)): pass\nprint(i)')
15496799
406215 function calls in 10.998 seconds
If you really don't want to do anything with the content of the file, better not spend time splitting lines, just read in the data in nice, big chunks. def file_block_count_lf(f) :
while True :
k = f.read(1<<20)
if k == '' :
break
yield k.count('\n')
In [77]: cProfile.run('print(sum(file_block_count_lf(open(fn))))')
15496800
10977 function calls in 6.479 seconds
On the other hand, I think sum(1 for l in open(fn)) is a little more pythonesque, and can also be written in one single expression, if you must. It's a little slower, though. In [70]: cProfile.run('sum(1 for l in open(fn))')
15903016 function calls in 14.409 seconds
Besides cProfile, also "timeit" is a useful python module for this kind of micro benchmarks.(the long output of the profiler is a little more interesting, I've put it in a pastebin http://pastebin.com/WXX93sS8 )
EDIT/ADD: There's an awful lot time spent in decoding the characters to unicode strings!
def file_block_count_lf(f) :
while True :
k = f.read(1<<20)
if not k :
break
yield k.count(b'\n')
In [48]: cProfile.run("sum(file_block_count_lf(open('more_logs.txt','rb')))")
4768 function calls in 2.969 secondsEdit; I just tested your code and my own and against a set of files of uniform size.
time wc -l GIANTFILE1 -> 16.7 seconds
cProfile.run("for i,v in enumerate(open("GIANTFILE2")):pass") -> 18.1 seconds
cProfile.run("sum(file_block_count_lf(open('GIANTFILE3','rb')))") -> 22.9 seconds
Would I have tested, say, different patterns of data access (reading files sequentially, parallel read from multiple disks/partitions), of course I would have made sure that either caches are cleared (/proc/sys/vm/drop_caches), datasets will exceed memory, and actual CPU load of the algorithm is according to the real-world-usage scenario (pure read() throughput is meaningless when CPU is bogged, and L1-3 caches swamped by actual algorithmic work).
Re: your additional tests: Interesting, what python-version and size of files are you running it on?