Don’t Slurp: How to Read Files in Python
axialcorps.com
axialcorps.com
And besides that, the author's solution is just good for one situation: when you want to read something line-by-line, which isn't always the case. For binary files, you may want to do something like this (untested code):
data = infile.read(256):
while data:
do_something(data)
data = infile.read(256)
Also, it seems like the author hasn't heard of enumerate: http://docs.python.org/2/library/functions.html#enumeratefor one, variance in the line lengths (by byte count) will probably force the usage of inefficient buffer sizes
also, the parent comment gave one universal example: binary files
Some formats, such as csv do make sense to read in line-by-line though.
The context of the question and answers was 'a LinkedIn group for professional Python programmers'.
2. Test that it works
3. Optimize (if needed)
If your file is small enough and is a trivial percentage of the overall run time, who cares how it is read?
My guess is that hardware improved and made slurping easier to do. Available RAM increased as hardware improved. This allowed slurping to displace filters as the choice way to work with files.
In a resource constrained environment, I still prefer filters to slurping. But how many developers or users today perceive their environment as resource constrained?
while (<FILE>) { print $_; }
is line by line and efficient. Perl programmers avoid my @entire_file=<$yourhandle>;
But even that loads the array line by line.As a side note, good Perl has always done things while there are things to do, in a list processing kind of approach feeling more functional than imperative.
// Slurping the whole file into a single thing is actually a chore: http://stackoverflow.com/questions/206661/what-is-the-best-w...
while (<>) {
...
}
By omitting the file handle, "<>" will go line by line through each filename specified on the command line. If no filename arguments were passed, it reads from STDIN. foreach my $line (<$FH>) { ... }
slurps the entire file because the foreach is forcing list context on the diamond operator.The
while(<$FH>) { ... }
that you describe works as desired though.http://stackoverflow.com/questions/585341/ appears to cover it in some detail.
[1] To me, anyway.
It has been submitted to HN many times: https://www.hnsearch.com/search#request/all&q=generator+tric...
for i, line in enumerate(sys.stdin):
print '{:>6} {}'.format(i, line[:-1]) import sys
for lineno, line in enumerate(sys.stdin, 1):
print('{:>6} {}'.format(lineno, line[:-1]))
The second argument to enumerate() is the 'start' (can also be a keyword). So by passing a 1 we start there instead of 0.I learned this the hard way when debugging a Python script that read from tail -f output...
for line in file:Then again, most of my files are three dimensional matrices. There's nothing I can really do on a line by line basis.
It's the wrong way in many cases, but it's the right way in a very large number of cases, if not in the majority of cases.
Often you need to read in a relatively small file, then do something trivial with it, or toss it through a couple of library-provided string processing functions. Or maybe the files are a bit bigger, and you're writing a script to automate some grunt work. You expect to run this script once. Or, more generally, maybe you're writing some code that gets called once a month and takes less than a second to complete.
In any of those situations, it would be silly to do anything other than slurping. String manipulation is easy to reason about. Stream processing is not.
Also note that most OSes cache files in memory, so if you are reading the same file often, the slowdown from reading the data into memory is drastically reduced.
The most obvious one is where you can exploit parallelism (either CPU parallelism or storage parallelism) by fetching multiple entire files at once and preparing them in memory. This allows you to start spending CPU time processing one loaded file, while other files load in the background. When you stream a file one line at a time, other than some basic optimistic lookahead, it's not really possible for the OS to do as much to help you there, so you're going to be effectively single-threaded. If the computation you're doing on the data is significant, you can end up being unable to even maximize your use of a single core on computation.
"A simple filter that prepends line numbers" # <-- Docstring
import sys
for fname in sys.argv[1:]: # ./program.py file1.txt file2.txt ...
with open(fname) as f:
# This reads in one line at a time from stdin
for lineno, line in enumerate(f, 1): # Start at 1
print '{:>6} {}'.format(lineno, line[:-1])
My way lets you pass as many files as you want to stdin, has a proper docstring, and uses the enumerate() function (so you don't need the silly `lineno = 0` and `lineno += 1` lines). with open("a.txt") as f:
c = ["{0} : {1}".format(x,y) for x,y in enumerate(f,1) ]
for x in c:
print x, <ipython-input-8-e2c5ebe72b17> in <module>()
----> 1 for x in c:
2 print x,
3
<ipython-input-7-9460e3a04a4e> in <genexpr>(***failed resolving arguments***)
1 with open("/tmp/foo.txt") as f:
----> 2 c = ("{0} : {1}".format(x,y) for x,y in enumerate(f,1))
3
ValueError: I/O operation on closed fileWhat's with all the nitpicking?
Edit: The idiom's mentioned in the Python docs on IO [1] as "memory efficient, fast, and leads to simple code"
[1] http://docs.python.org/2/tutorial/inputoutput.html#methods-o...
Another option is using mmap[1], particularly when the file is already in memory or you need more random access to it. It worked well when I was trying to parse some lines from the end of an open log file.
In CPython, since reference counting is used for GC, this occurs when the loop exits. However, other implementations (e.g. PyPy) may use different schemes that do not guarantee collection as soon as objects go out of scope. As an extreme, a valid and occasionally useful GC strategy is to never collect anything at all [0]!
Hence, if you want to portably ensure the file is closed, you should either use the context manager or call close() explicitly.
[0] http://blogs.msdn.com/b/oldnewthing/archive/2010/08/09/10047...
$ seq 1 10000000
I think a lot of people forget about the beauty and power of coreutils. $ uname -s
Darwin
$ type seq
seq is /usr/bin/seq
The man page says --The seq command first appeared in Plan 9 from Bell Labs. A seq command appeared in NetBSD 3.0, and ported to FreeBSD 9.0. This command was based on the command of the same name in Plan 9 from Bell Labs and the GNU core utilities. The GNU seq command first appeared in the 1.13 shell utilities release.
So the moral of the story is that Python makes it simple and elegant to write stream-processors on line-buffered data-streams.
I thought every language had this or a similar method as a best practice when processing 'large' files.Other than that, it's more memory efficient.
It can also be faster if you're searching for something in the file, because you can short-circuit reading the rest of the file when you find it.
By streaming the data, you can do both at the same time. (The OS will already fetch the next line from the disk while you are still processing the previous one.)