Copying stdin to stdout in Java
gist.github.com
gist.github.com
IOUtils.copy(System.in, System.out);
using the Apache Commons library.
That's the most apples-to-apples comparison to more "scripty" languages that have more stuff built in. With Java you pick the libraries/frameworks you want to use. Commons and Guava are extremely popular and robust libraries and really should be included in any comparison against "real world" Java.
I actually like how in Java you can get down closer to the metal and control the exact buffer size if you want to, or handle all the exception cases instead of just passing them up if you want to. If you're comparing intentionally lower-level more verbose Java code to higher-level Python or whatever, that's not really apples-to-apples. High-level Java involves frameworks.
You could build Awk and Perl in Java if you wanted and have the same result.
System.out.println("hello world");
Is the same as: printf("hello world\n");How is file manipulation High level. That sort of stuff should work right next after your Hello, World program.
If you are telling me Java intentionally makes easy things difficult. Its the last thing I want to use to solve difficult problems.
I don't hate Java, but "Java as it is spoken" does not often get used for Unix-style piping in a chain of small programs on a command line. It also isn't the easiest tool in the drawer if you want to do light (or heavy) string processing.
The lack of direct abstraction is not a valid argument, because as a programmer, you shouldn't be writing the logic in Java, Python, C or Scala but rather a higher order domain language implemented in the chosen host language and that's what the process of programming is all about. For the majority of real world programming tasks it's unlikely that a language that fits the domain perfectly already exist, so you have to create one.
In Java one can say:
copyStream(System.in,System.out);
And then one will have to implement copyStream but only once: long copyStream (InputStream src,OutputStream dst) throws IOException {
long bytesCopied;
byte[] buffer = new byte[8192];
int bytesRead = src.read(buffer);
while(bytesRead!=-1) {
bytesCopied+=bytesRead;
dst.write(buffer, 0, bytesRead);
bytesRead = src.read(buffer);
}
return bytesCopied;
}
I prefer programmers taking this approach of implementing domain specific language first and then expressing the logic in its terms instead of trying to express higher order concepts without resorting to available host language abstractions.Let's say Python or Perl let one express stream copying more concisely straight out of the box. However when faced with real life programming challenges one will very quickly encounter limits of what a language can express out of the box with one-liner. But as a programmer one has the power to create one-liners from scratch!
Disclaimer: I am not a Java expert, so the code above is just to illustrate the idea based on my very limited knowledge of Java.
First, constantly having to write implementations like this introduces a ton of friction as compared with reusing existing implementations. Languages do have an influence on that.
Second, languages vary in "expressiveness," determining how much code you have to write in order to make an implementation like this. In one language, copying stdin to stdout fits into a natural idiom which also handles other cases naturally. In another language with less care for ergonomics, every task might be equally un-idiomatic.
In other words, a language can offer its own "higher order domain languages" for the core tasks that everyone is doing over and over again as part of their general programming. Or it can choose not to do that, because what it offers is already Turing complete. But then the ergonomics are bad, and it makes a real difference.
It seems wrong that when I am paying for things like long JIT warmups and stop-the-world GC, I am still writing piles of functions with low-level idioms that are no more expressive than C's.
If Java ships with a 'copy from one stream to another' primitive then I'd find that a much more compelling argument than that I can treat myself to reimplementing things like stream copying, sorting and basic data structures on a regular basis. It's unbelievably tedious and wasteful to do this, there's just no reason.
My point was that limited expression means of a language are sometimes compounded by inability of a programmer to make a good use of the expression means already available to them.
I also understand the desire for more a expressive language, I am a programmer after all. One has to keep in mind, however, that the more expressive a language is the harder it is on the reader. Java code is trivial for a reader to follow (if not for the excessive verbosity sometimes covering up the true intent); much of Scala code base, on the other hand, is not that trivial to comprehend due to the high expressiveness of the language.
wget -O - http://audiostream | java inout.java | mplayer /dev/stdin
If I had to write this in C I'd use file descriptors in non-blocking mode and select(). With Tcl one could use event-based IO.[1] http://docs.oracle.com/javase/7/docs/api/java/io/InputStream...
The documentation does not answer how long it blocks. Blocking when there is no input is completely ok in this use case as there is nothing better to do instead of waiting for input. The relevant question is how long it blocks after there is some input available for efficiency reasons.
In case no more input is available it would be best to already output the data that have accumulated in buffer. Else in case the input program deadlocks you're never going to see the last (up to 8191) bytes output.
This is what non-blocking I/O is for. With non-blocking I/O you first wait until your input channel becomes readable, then reading functions return as soon as the data available for read is exhausted.
> Blocking when there is no input is completely ok in this use case as there is nothing better to do instead of waiting for input.
(Emphasis is new.)
[1] http://docs.oracle.com/javase/7/docs/api/java/io/InputStream...
package main
import ("os"; "io")
func main() {
io.Copy(os.Stdout, os.Stdin)
}Any ideas if this dest<-src order is by design following go design principles or was mainly authors' personal choice?
By the way, Pipe in the same package is the other way around: (PipeReader, PipeWriter)
dest = src // Variable assignment.
<-ch; // Read from channel.
// in net/http
type HandlerFunc func(w ResponseWriter, req *Request)
// in io
func Copy(dst Writer, src Reader)
// builtin
copy(dst, src) // Slices.
So the reason is consistency. I suppose the Pipe exception is to map Unix's pipes.At this point, Guava should just be added as part of the JDK. :) In Guava, ByteStreams.copy(System.in, System.out)
Could you elaborate? I want to understand how to fix it.
EDIT: I really want to know what is wrong with my code. It works. Is it a stylistic concern?
$ dd bs=1m count=10 if=/dev/random of=randomdata
10+0 records in
10+0 records out
10485760 bytes transferred in 0.878533 secs (11935535 bytes/sec)
$ python stdin_to_stdout.py < randomdata | cmp - randomdata
$ echo $?
0
EDIT2: I think I know how you came to the conclusion my code is wrong. Perhaps next time you should read the __entire__ post. Thank you. while( (bytesRead=System.in.read(buffer)) != -1 )That post, maybe?
while (<>) {
print;
}
Still Awk is shorter: { print }
And if you are allowed change the switches of the command line to Perl for your program that Perl program can have exactly 0 bytes.By default, unless you use the BEGIN block or something, awk will run your program on each line of stdlin. This is useful for programs of the type:
(/some regular expression/) { some action}
The default action if you don't specify one is "print $0" (the whole matching line). If your condition is a plain-ol' expression rather than a regular expression, and it always evaluates truthy, you thus get every line.
Still the shortest Perl is 0 bytes when the command line switch is allowed to be -p
Update: That's actually called "the sed mode of Perl." Moreover these two command lines behave similarly:
perl -pe 's/search/replace/g'
and sed 's/search/replace/' sed < input > output
:)[1] http://robertkotcher.com/sed.html
[update] s/set/sed/
perl -pe '' <input >outputI was referring to an empty SED-program that would then be run via 'sed -f empty-file.sed'. Of course you need >0 bytes to invoke the SED program.
Your example is very much just a special shortcut for one special condition. It's not generalizable in the same way. And let's be honest - most of today's code is not writing to standard output, it is writing to some iOS or Android GUI or outputting JSON to be used in some javascript webapp. That kind of shortcut just doesn't make sense anymore and is a relic of a different time.
http://computer.howstuffworks.com/cgi3.htm
The server opens the socket, redirects the stdout to his socket while calling your program fully unmodified.
As for serving files, you'd normally let nginx or apache take care of the sockets for you and not use output redirection.
In Plan9 it's actually viable to use network sockets as files which is pretty nice, but for Linux it's not really how people do it.
Please explain that to all the people who wrote CGI scripts for years. Seen the link at all?
The ability of a language to easily write to std out is simply not important anymore apart from shell scripting which is mostly used these days for setting up the runtime environment and building other languages to handle the grunt work.
I don't know of any companies that still pipe std output directly to sockets anymore, but I guess there must be some. If your company still does this then a language feature like that makes sense. For most people, having the ability to write code that can directly write to sockets without piping through a shell is better practice.
The major advantage that CGI gives you is the fact that you get a new, clean process on each request. This means you have fewer and simpler failure cases. This can be helpful when you have a line of business web application though it is certainly suboptimal for public-facing scalable systems.
I wouldn't say cgi is dead. It still has a role to play. It is not the ideal way for doing many things though. I.e. they are no longer the go-to tool, but rather one tool for relatively lower-volume applications where performance matters less than some other considerations.
grep $
That reads lines from stdin and echoes to stdout. #!/bin/sh
cat $ printf 'binarydata'|awk '{ print }'|hexdump -C
00000000 62 69 6e 61 72 79 64 61 74 61 0a |binarydata.|
0000000b