Also, it goes to show that no matter how bad aspects of bash are, the power accumulated over the years is hard to match, and a lot of that power is strongly coupled to the arcane and terse conventions.
Also, it goes to show that no matter how bad aspects of bash are, the power accumulated over the years is hard to match, and a lot of that power is strongly coupled to the arcane and terse conventions.
A lot of very simple things are harder to do (text is easier than objects up until a point) since Powershell treats everything as objects and LinuxCLI/Bash does everything as text being piped around to specific commands.
The biggest complaint I have with Powershell is that it is so obnoxiously slow at parsing text files, that it just can't do basic scripting tasks on even medium sized files ~100MB without tons of StreamReader[] overhead or without calling a pre-compiled C# binary directly. This is the biggest fail. Nearly all the obvious ways of appending to files or parsing them stop working once the file size gets bigger than tiny. This isn't a problem with Linux commands that just pipe text to highly optimized C binaries. I'd drop Python for 90% of my scripting tasks on Windows if Powershell were faster in this area, but it is just painfully slow (if all the obvious ways of doing something in Powershell take ~5min when even Python can do it in ~4sec, you have a problem).
However, you can usually write some very beautiful and easy to read code by using the built-in objects once you learn them. I also like how powershell's functions allow you to create commands with parameters in a super easy way so you can easily build DSLs in a way. Everyone feels differently, but calling my commands in this manner (InvokeCommand -param1 "blah") reads a lot more naturally to me than (Object.Method(param1)). Also, Powershell can handle things like filesystem interaction in a way that doesn't feel like a bolted on afterthought as Python's does.
I think Powershell "jobs" can help run things in parallel, but I've never done more than just read about it. The Powershell in Action book (3rd edition) covers that functionality I think. The first chapter to that book is pretty awesome.
My dream as someone forced to use Windows would be for Microsoft to use some dollars to make Powershell a one stop shop for data science, simple games (think commodore sprites), numerical computing...etc. That sounds a little wacky I know, but I think people would appreciate every computer shipping with a powerful environment like that. Since Powershell is built on .NET, surely someone could make a library for Powershell that abstracts away some of the .NET complexity to where I can just do something like:
Process-Data -data "blah.csv" | Create-PieChart -output "chart.png"
Not having to install Anaconda and import a ton of libraries on the target computer would be helpful. Note that I'm not saying to reproduce all of Matlab in Powershell, but putting some of the most popular routines as built-in commands would be pretty awesome.
It would be up to "the community" to make modules to do this kind of thing, which would be possible if anyone wanted to, but unlikely to ship with Windows.
> putting some of the most popular routines as built-in commands would be pretty awesome.
Microsoft's push is to avoid more "bloat" by leaning more heavily on optional modules installed from the PS Gallery. But `group-object` is a builtin for doing a kind of SQL GROUP BY operation, and Measure-Object just gained a -StandardDeviation option, and there's now ConvertTo/From-Json, and ConvertFrom-Markdown. If there is a basic routine that would be popular and fit many use-cases, you could request it at https://github.com/powerShell/PowerShell/issues/
There is a GPL licensed scientific computing framework called Accord.NET, it's not wrapped for PS but can be used almost directly, and some of these others might be viable as well:
https://www.reddit.com/r/PowerShell/comments/5ijcj7/kmeans_c...
https://en.wikipedia.org/wiki/List_of_numerical_libraries#.N...
I agree that there isn't much commercial incentive, but I honestly don't see why major OS providers can't include some basic graphics primitives in the OS. I'm not talking about embedding the unreal engine or anything. It could be used for a lot more than primitive games btw.
I understand the desire to avoid more bloat. I can't believe Windows 10 requires the storage and RAM that it does.
The XML & JSON objects are pretty cool.
I'll check out Accord.NET framework too.
Yes, that's a common complaint. And PowerShell 7 has just gained this with a new `foreach-object -parallel` parameter:
https://github.com/PowerShell/PowerShell/pull/10229
Which should help at least some cases. (PoshRsJob is a reasonably easy / fast way to use multiple processors inside the same overall Powershell process (with RunSpaces).
I agree with you that the "obvious things" are frustratingly slow, but you can get significant (>10x) improvements with small changes, even if not the fastest. If big.txt is a random test file of 100MB and 700k lines then I get this:
$x = get-content big.txt # 15 seconds (obvious and bad)
$x = (get-content big.txt -raw) -split "`r?`n" # 6s, regex split mixed line endings
$x = (get-content big.txt -raw).split("`r`n") # 2.2s, fixed line endings
$x = ${c:\path\to\big.txt} # 1.2s
$x = [io.file]::ReadAllLines('c:\path\to\big.txt') # 1s
$x = [io.streamreader]::new([io.filestream]::new('c:\path\to\big.txt', 'open'), [text.encoding]::ASCII, $false, 10mb).ReadToEnd().split([System.Environment]::NewLine) # 1.2s
PowerShell 7 is working on performance improvements throughout, but early design decisions mean this particularly is not likely to change.> This isn't a problem with Linux commands that just pipe text to highly optimized C binaries.
Python can do this easily as well, and PowerShell can't, and that is one thing I think doesn't get enough discussion and mention. For however good C#/.Net can be, it's still another layer between PowerShell code and low level OS stuff, a layer other shells and high level languages don't have, and a shell is supposed to be "OS Script" first, not "C# script". Python ctypes -> C library is just a ton easier than PowerShell -> C# -> P/Invoke -> Win32 API.
I didn't know that slightly modifying "Get-Content" could speed it up so drastically. As we've both stated, it isn't very obvious or intuitive without knowing the detailed inner workings of the cmdlet. Do you know why the "Get-Content" can't be more like the $x = ${'path'} option by default? Also, what is going on in that example and how is it so much faster than "Get-Content"?
The examples where you call out to .NET directly aren't too bad (well at least the first one is terse enough), but there are a lot of users like myself who don't write much C# and therefore don't know which namespaces to call for .NET voodoo. The reason I learned Powershell was so I wouldn't have to use C# for scripting as it really doesn't Excel there being statically typed and not having an interactive REPL. If I have to constantly write C# for performance, I might as well just write C# lol.
I'm glad to hear Powershell 7 will get performance improvements and am sad it won't change anything here, but at least I can try out your examples!
As far as Powershell performance is concerned, I still don't understand how some cmdlets can't call out to something pre-compiled. Does all the powershelly parsing and dynamicism really hurt that bad?
For the cause of the slowness, try a Get-Content file.txt | format-list * -Force on a single line file and have a look at what comes out; you'd think it would be a System.String with only a Length property but surprise, there is a lot of metadata about which file it came from, and which provider the file was on. A small addition to one string, but it adds up for tens of thousands of lines. That is the big reason Get-Content is so slow. PowerShell has Usermode filesystem / virtual file equivalents called PSProviders, so there is some overhead for starting up a cmdlet at all and then binding its parameters and then dropping through to the FileSystem provider, but the added string properties are the big overhead for file reading.
They should allow you to do convenient things further down the pipeline, e.g. filter the strings, then move the files which contained the strings you have left, although I'm not aware of any genuinely useful use case. I expect there is one somewhere. This is baked in from v1, so removing that would be a breaking change on a very heavily used cmdlet, so it won't stop happening.
The variations like Get-Content -Raw were the intended workaround, by outputting a single string, you skip much of the overhead. The other approaches, ${} and .Net methods, avoid get-content altogether. ${} is the syntax for a full variable name, and PowerShell has PSProviders which are a little like virtual drives, usermode filesystems, or /proc and so you can ${env:computername} for an environment variable, ${variable:pwd} for another way to get the working directory, and ${c:\path\to\file.txt} fits that format so they made it do something useful.
> Does all the powershelly parsing and dynamicism really hurt that bad?
The language is interpreted, but loops which run enough times get compiled by the Dynamic Language Runtime (DLR) parts of .Net, swapped in while the code is running, and run even faster. Parts which are a thin layer over .Net like @{} is a System.Collections.Hashtable and the Regex engine is just the .Net regex engine with case insensitivity switched on, they run pretty fast.
Function calls are killer slow. Don't use good programming practises and make lots of small dedicated functions and call them in a loop, for speed inline everything. Because functions can "look like" cmdlets with Get-Name -Param1 x -Param2 y, there is a HUGE overhead to parameter binding. You might have configured things so -Param2 gets the values from the metadata on the strings from Get-Content, there's a lot of searching to find the best match in the parameter handling.
The pipeline overhead is quite big, code like Get-Thing | Where { .. } | Set-Thing runs pretty slow compared to a rewrite without any pipes or with fewer cmdlets. (Get-Thing).Where{} is faster, and $result = foreach($x in Get-Thing) { if (..) { $x } } even faster.
The slow bits IMO are where PowerShell does more work behind the scenes to try and add shell convenience rather than programming convenience; the great bit is "it's an object shell", the downside of that is "making objects has costs". Compare SQL style code "select x,y from data" describes the output first, then the database engine can skip the data you don't want. Shells and PowerShell have that backwards, you "get-data | select x,y" and the engine has no way around generating all the data first only to throw it away later. Since PS tries for convenience, its cmdlets err on the side of generating more data rather than less, which makes this effect worse.
I don't want to end up writing C# for everything either :)
I had spent 2 hours yesterday trying to sort CSV lines in bash. Decided to do so even though I have a nice strongly typed framework to manage tuples, which is slower than bash.
I always forget that the -k argument requires 2 numbers if you want to sort by only one column.
Then LC_ALL.
These are not the things that are so easy to debug against.
So in the end it would have been faster to use Python.
There’s xsv-rs, but it doesn’t provide the same semantics or power I have in the framework.
Point is, I think there’s some space for novelty in UNIX tools. Especially given the research in DSLs for defining the transformations of tuples.
I did not know about CSVKit three years ago; but I don’t think it would take a prof. developer to build it too much time either.
I believe fish + python have a higher power/WTF ratio than bash and work in the same space. Perl is mentioned as well, but has fallen out of favor.