Computational survivalist
johndcook.com
johndcook.com
Here's an example:
I was working at a consultancy one day as a data engineer, and across the desk was a junior data scientist. We were preparing for a meetup and the manager said "this website has a bunch of email addresses on it I'd love to have them in a CSV" (they were in a format where it wasn't easy to copy + paste them)
The junior data scientist said "I can do that" and the manager said "only if it takes less than 2 minutes", the junior, thinking he'd have to: setup a python environment, find a parsing lib, write code to pull the page, parse the dom, search for the email addresses, write them to a file,
said it could be done in an hour or two.
I viewed the code on the page, and saw the html wasn't minified, and the "a" tags were 1 per line, so I :
use wget to pull the webpage cat webpage.html | grep 'mailto:' | cut -d '"' -f 5 | cut -d '"' -f 1 > emails.txt
^ unix wizards will tell me a better way of doing this, but it took me literally 60 seconds to go from "I wonder if" by the manager to "here's your CSV"
THATS WHY WE USE COMMAND LINE TOOLS, they are a swiss army knife
https://adamdrake.com/command-line-tools-can-be-235x-faster-...
Great story. I won't claim to be a wizard, but here's mine:
grep -Eo "mailto:[^\"']*" webpage.html | sed 's/^mailto://' | sort -u
The pattern means "mailto: and then everything up to a single or double quote", and the -o flag means "only show the matches". The main benefit of this approach is that it works even if you have multiple email addresses on a line.There is a whole heap of tools in the package, and they're definitely worth a look if you like to poke at the web from a term. Plus they handle a lot of the nasty corner cases that will pop up if you try to treat HTML as simple ASCII text.
document.documentElement.innerHTML.split('\n').forEach( (line) => line.includes('mailto:') ? console.log(line) : 0 )
But please push this further (is there a shorter way to say document.documentElement.innerHTML ?)
a = document.createElement('a');
a.download = 'emails.txt';
a.href = `data:text/plain,${Array.from(document.querySelectorAll('a')).map(a => a.href).filter(s => ~s.indexOf('mailto')).join(',')}`;
document.body.appendChild(a);
a.click();
`document.getElementsByTagName('a').map(elem => elem.href)`
[...($$('a'))].map(e => e.href)
can you share the smart line of javascript? You have 60 seconds, go.
More to the point, though: the junior data scientist might not have been entirely on the wrong foot there. There's a tradeoff between "I need to do this thing quickly" v. "I need to do this thing repeatably and robustly"; whipping up a quick shell one-liner to do it falls into the former category, and also can drive some ideas for implementing the latter category (where you'd almost certainly want to be properly parsing the HMTL, properly accounting for commas in the email addresses (and yes, email addresses can have commas per RFC 5321), etc.).
This is great! I can't believe I've never run into "cut" before. In similar spots, my workflow would have been filtering the lines with grep and then opening the file in VI to finish the task.
Cut has a mirror image program, "paste", but I've not yet hit a use case for it.
I mean, even my friend the electrician gets money ech year to buy tools. And he buys Makita and DeWalt, not that 10 euro for 50 tools package at Ikea.
Yeah I'm not forced to work there, anyone saying that is right. And I dream of starting my own business just because of this.
I think about quitting a lot, but I am pretty numb to it. I'm sure it could be worse, too.
Seriously, that kinda bs is time consuming and soul deadening.
I like the analogy of a photographer: Talents aside, a pro will always make better pictures using a cheap low-budget camera than any amateur will be able to do with the newest and best camera available.
IMHO the same holds true for any IT related job.
Unfortunately it doesn't work like this.
Photography, as a low threshold occupation has tons of professional hacks who are bad with any equipment. At the same time there are countless amateurs who are top talent.
And it's not just a recent trend. Aside from photojournalism, formal portraiture and weddings, photography was always dominated and advanced by talented amateurs. There's no such thing as "professional landscape photographer" for example, even if some of them are able to make a modest living off their works after years of effort. Ansel Adams is a great example.
Also, no professional photographer would use a toy camera for anything other than a gag/vanity/show-off project. They infallibly use reliable, quality equipment day in day out.
> You: "It doesn't work like this. [...] talented amateurs [...]
Uhm, okay.
Well, anyway, my point is pretty valid. If photography is not as much dependent on skills from your perspective (I wasn't talking about weddings and landscapes anyway) take something else. It really applies to everything.
Maybe music is a better fit? As someone that even actively plays piano I am still 100% confident that someone like Beethoven would run a better show on a $40 Casio keyboard than I would on a Yamaha C5.
For what it's worth, a lot of the best photos out there are from amateurs (probably the vast majority, there's just more amateur photographers).
The pro, however, should _always_ be able to at least get _good_ photos, while an amateur might, or might not, who knows? This is why people hire a pro for a wedding - you know that there's exactly one chance to get that first kiss, etc. and need to have someone reliable. My favourite wedding photos are from weddings I didn't get paid for but attended as a guest, though I wasn't worrying about running through a standard shot list (and the group portraits, which I never liked compared to photos of people having fun at the reception)
Small confession: I'm currently working from home, Win10 laptop runs out of juice, I forgot the charger. Managed to set up a reverse SSH tunnel from our company cluster to my home server, proceed to work via SSH and SSHFS but now on my personal Arch Linux laptop. Loving it, feeling a bit unethical though.
An "IT pro" should know about general risk management in IT departments or reasons why you don't give everyone full admin rights or every software they demand or everyone a machine with totally different specs from the other ones (read as: their favorite one). It becomes an unmanageable chaos at some point, too.
Who will be responsible if anything bad happens? Do you even want to be that one? (Note: bad can really be anything on a scale from 0 to hell and beyond in IT)
After so many frustrating experiences with modern GUI apps significantly changing interfaces and functionality, phoning home and creating privacy concerns, or being just plain buggy, I've moved to "slow computing" and have found it a huge relief. And incidentally, not slow at all -- I'm far more productive.
I don't think this is a helpful characterization, because this invokes post-apocalyptic associations of a world where there are somehow only text terminals left. In the article itself, the author goes on to describe a scenario where he used generic text/command line tools to achieve functionality that had not been explicitly granted to him by a restrictive software policy at work.
As opposed to survivalism, I think it's more fitting to describe it as user empowerment. I'd argue that the premier usage scenario for these tools is not as a fallback in hard times, instead it's an important type of toolset integral to the craft.
an SSH session?
I understand where he is coming from and why he may have gotten that impression but it is more like "accubus ergonomics" essentially for reasons I will explain at the end.
The point of the command line is that the shell is very transportable, efficient, and powerful. Compare updating a syatem - you can just write or pass about a single shellscript to do what woukd otherwise require clicking through a ton of menus, watching bars grow, and waiting for it to need clicked. Not that big a deal for one person one system, a bigger deal for 5000.
Its main weakness is the required experience and knowledge to use so if you already have those capabilities the added costs are marginal.
This the name - it is like dealing with decimals or integers with an accubus vs a calculator and scratch paper. Surprisingly despite a millenia gap the calculator is slower - the user needs to read symbols and enter them one at a time with a hunt and peckish sort of mechanism. Meanwhile a skilled user will have its operation essentially in muscle memory and only need to read the end state. The accubus is rare in comparison because it generally isn't worth it on its own vs the easier alternatives. Plus the paper has a less volatile storage format than bead position.
Command line operation isn't for everyone and in most cases and that is okay but it still has a valid purpose.
In some cases, this can feel a bit regressive. Stepping back to a bone-head simple pattern with fewer moving parts and less dependencies on complicated "others" I cannot see.
Part of why I hate WordPress with a passion. It's complex mess of plugins and themes, slapped together without the tight moderation and insistence on good style that we see in a place like the Python standard library.
I mean, you might as well install Python.
Do your task with PowerShell or DOS batch, and then this might cut it as computational survivalism on that box.
EDIT: Here, the solution to the author's problem without needing to pull a *nix tool chain in.
((Get-Content "<file name>").ToCharArray() `
| Where-Object { $_ -eq '<the character that you need>' } `
| Measure-Object).CountSecondly it looks like it depends on a fair amount of typical userland being available. It'd prefer something I could slap on top of busybox and be done. Something that only depended on the kernel.
But, maybe with some tweaking...
[1]: https://elv.sh
In more extreme cases, I've seen it as an argument against Microsoft software like Windows and Word.
In less extreme cases, I've adopted this attitude with respect to some tightly restricted engineering (CAD) software used back in grad school. I mean, it was really useful software, but I didn't feel that I could reasonably rely on having access to it whenever I'd need it.
Yes, the information contained might not be completely up-to-date or accurate, but at least it sets you up on the right path to search further.