https://my.cse.unsw.edu.au/thesis/thesis_topic_details.php?I...
552 karma · joined May 5, 2013
https://my.cse.unsw.edu.au/thesis/thesis_topic_details.php?I...
It's also quite likely that a pre-built docker image isn't going to satisfy everyone's deployment requirements.
Starting out with a philosophy that you'll make all the configuration and security choices for every user of your project isn't a great start, and that's how I read all these hard-coded assumptions.
When it comes to deploying, supporting and securing complex web application software, it is very rare that an out-of-the-box reference image from the project or vendor will have much relevance to your own environment's security and configuration standards.
Unless you're paying for support, and the vendor insists you'll have to suck it up and accept whatever madness their support agreement entails.
I'm all for convenient docker images, it just frustrates me to see the total and utter lack of imagination demonstrated by some projects that can't imagine their users ever locking down their database access, or running with SELinux, or separately patching the software/service components they depend on, etc.
And again, I don't expect the reference docker image for a given project to allow all of that flexibility and configurability, it's more the fact that these possibilities seem to be actively sabotaged with reckless abandon.
Indeed, and I use Docker myself to improve my own efficiency, and I've seen great stuff built with Docker that has been architected well.
However, it is painfully obvious that without the pressures that used to force most developers to keep their shit sane, there's now more workload for serious users that actually need to untangle this mess in order to support, secure and get stuff deployed.
I mean, when I first evaluated this app, it didn't even have a sane launcher script. Instead, ~15 lines of ruby config and a ruby script that I had to follow just to discover it was a weird, idiosyncratic way of executing "docker run".
Edit: I am not saying that Docker is somehow inherently flawed in this respect. It's solved a lot of problems for me and you can build great, well-architected stuff with it. But there is a trend to use it to wallpaper over poor development and deployment practices.
For everyone else, it's giant shit sandwich, and a constant reminder that the project couldn't arsed untangling their own weird idiosyncrasies from their distributed app, and that this crap wouldn't fly with a traditional distro package, or wouldn't have mattered with the more traditional tarball/REQUIREMENTS/INSTALL.txt that made way fewer assumptions about the end-user's environment.
I love docker, but it's letting people get away with murder.
It's a better, more isolated mess - but for anyone trying to enforce configuration policy into all of the running services in their environment, untangling gortesquely basic shit like not granting superuser privs on a database to your webapps - which would never fly in a traditional distro package - becomes even more work than the old source tarballs with INSTALL.txt.
There's giving up on traditional package management, and then there's throwing yourself off the cliff of decent release/config/dependency management that used to be a side-effect of traditional package management.
"Works for me on my computer" is now "Works for me with my Dockerfile [ and all the configuration assumptions I couldn't be bothered untangling from my app ]"
I can't tell you the number of times I've hoped to use a public docker image or just the Dockerfile, only to spend hours futzing with it because I was unhappy with the grotesquely insecure configuration or because I needed to work around a bunch of assumptions that are invalid once I've tuned the Dockerfile for my environment.
To paraphrase provocatively, 'machine learning is statistics
minus any checking of models and assumptions'.
~Brian Ripley, 2004 [1]
We should really check out some of the stuff bioinformatics people have been doing. That's not to say that cryptography has more to give to biology, quite the contrary. But I think the broader infosec community should try harder when it comes to doing stats properly. I think there is a lot more in common between bioinformatics and infosec analytics than most (on either side) realize. Huge amounts of messy, un/semi-structured data being the common theme.In my (admittedly limited) experience, I find that wetlab biologists tend to have a better understanding of the limitations of the experiments they're setting up, in terms of interpreting the significance, limitations and meaning of their results at least - than the computer science types playing in biology. But good stats is the one universal skill that all good scientists should master.
I spent a few years where part of my work was supporting evolutionary geneticists manage and figure out their data. Enormous amounts of data. So much raw data that the very large scientific institution I'm sure you've heard of couldn't accommodate it in their normal data repositories (at least not if every researcher started doing so). Having multi-million dollar research projects conclude and then archived to a handful of redundant sets of HDDs curated by teams and departments that now no longer exist is unnerving (supposedly things are better these days).
In any case, it's a fascinating area to work in for budding computer scientists. There are many problems to sink your teeth into, even basic stuff like repeatable analyses on research projects running for more than 10-15 years. "Yes, there are faster/more robust/more appropriate MCMV (Markov Chain Monte Carlo) solutions out there but we don't feel like re-running a decade's worth of work to ensure they're comparable". But I digress.
It amused me that each research team seemed to be either computer-science heavy (and ignorant about the quality or nature or "ground truth" of their data and the possibly embarrassing impacts on the signals they thought they were seeing), almost wilfully ignorant or at least dismissive of "old-school" biology methods and data...
... or were very wet-lab savvy and grounded in the raw biology side of things but limited their research questions according to the types of analyses they were comfortable doing themselves, using well-established methods, tools and software that would run on their biggest $20k PC under someone's desk. Although that was seriously changing by 2013 when I left (partly due to nice big web-UIs in front of HPC clusters, partly due to better mixing of team capabilities).
Now, it's easy to dismiss the latter group as just being not "up to date" or unskilled in "big data" (recent biology grad programmes are hopefully improving). But what was surprising was the obvious opportunities (to me) that highly CS-capable teams (inside and outside of the organisation I worked for) seemed to be deliberately staying away from. On more than one occasion I asked some of these people why they were ignoring the rich, machine-readable, highly-curated datasets from more traditional biology science (Eg. taxonomy/ID/phenotype/ecology info).
The answers are varied: they aren't even aware this data exists in machine-readable form; they seem unfamiliar with the very biology discipline they're playing in (they're more obsessed with algorithms); but a disturbing trend seemed to be a generic distrust of manually curated data. I guess because there's no repeatable algorithms to reproduce human curators :-)
Now, I'm not saying everyone should go out and use that data as a source of truth, but it should at least be a reference to sanity-check or compare to. And imagine if you actually did augment your sequence data with this stuff - one (admittedly naieve) thought was that there might be a fun way to look for at least some phenotype influences in a given genome, "for free", without having to start new wetlab work - i.e. just repurpose existing data.
When countering the aversion to manually curated data, I tried to point out that the very "garbage in" being dealt with - the genome sequences, spatial records, environmental layers etc. - all involve an element of human curation. I mean, that's how we end up with nearly 20% of non-human genomes containing human DNA [2] in the first place. And that's before we even consider whether stuff in the "good" 80% are even representative of the species group(s) of interest - you wouldn't believe how easy it is for people in the field (or the lab!) to get species identification completely and utterly wrong. Or scrape up DNA samples of some infection or parasite instead of the actual thing itself...
... and yet even with all these imperfections in the molecular data, we find that some of the old-school morphological (phenotype-driven) taxonomies are still 80%+ unchanged (hazy on the numbers now) when reconstructed with purely DNA-derived phylogenies. And that's on groups where we're still fighting over what the best DNA markers are to build those molecular phylogenies in the first place! A lot of people don't realize that the same DNA evidence can give different pictures of the evolutionary tree, depending on what genes you're picking on...
Apologies for the over-sized rant, it's been a while since I had to think of this stuff.
[1] http://stats.stackexchange.com/questions/6/the-two-cultures-...
But I haven't quite had time to figure it out myself yet; I've been interested in exploring Geode for a while.
Oversized egos are a problem at the best of times in open source, especially in the kernel. Which has traditionally had a culture of casual indifference toward the security priorities of PaX/grsec, and some would say toward security generally (I don't buy that - there's a lot of concern for integrity in the kernel, and that at least buys you some security).
Whilst it would have been great if Spender could have completely changed his personality, outlook and perspective on the world so that he could commandeer the linux kernel from the outside-in toward his vision of a safer kernel, is that really compatible with the wider project? And so, it's equally depressing there hasn't been more recruitment from the kernel side - the side that actually has resources - to figure out a path toward getting PaX stuff mainlined from the inside-out.
I guess my point is: grsec has one focus, one priority. But the Linux kernel is a project that has a much more vast, broader scope of competing priorities to untangle, and I just can't see that such an enormous, busy and byzantine project ecosystem truly will ever see a clear mandate to get something as disruptive and single-minded as grsec mainlined.
Not only would Spender have to become a different person, but the core kernel teams as well.
In any case, I've reduced my patreon donations and put what little that is towards grsecurity. It's certainly that important and useful to me.
That said, I didn't study Software Engineering at university. And so, in almost 10 years in various roles writing software, this is the first one in which I've allowed this title to apply (previously I've had: Systems Analyst, Software Developer, Bioinformatics Technologist, Instrumentation Engineer, etc).
I also didn't study systems engineering (which is a real discipline responsible for remarkable achievements such as taming the complexity of massive systems such as modern airliners), but not that Systems Engineering, I mean the wishy-washy, everyone's-an-engineer sense of Systems Engineer as typically used in software land - that's what I think a more accurate title for my role would be.
I work on the hardware for our widget, as well as the user-facing software. I set up our automated build system that builds everything including our ARM Linux OS image, including kernel and bootloader. I established the production and Q/A checklists. I actually do some of the assemblies. I determine what components we use, and I work with real engineers to validate and test and update customer spec sheets accordingly. I answer questions like, "what kind of power system would let us get XX hrs runtime at YY% duty cycle".
I have a complete picture in my head of the product, its capabilities and limitations. When it goes wrong, at any level, I'm the one that has to fix it. I've never worked with a Systems Engineer at a big tech company but it seems as if this is the sort of scope they work with. Except they probably don't have to write the userland code as well :)
That said, it really was an offline app. We could've done it as a native Qt or .NET app. Even so, the standard AngularJS document fragment paths supported bookmarking and so on.
But the 2002-2005 stuff had aged much worse. At some point there was fancy site generator that had used javascript for everything, and the javascript apparently only worked properly in IE6. So most of the navigation was busted in a modern browser, and needed special scrapers to parse out what should have been plain old <a href...> tags.
Now, I regularly think back to that crusty old HTML3 static site that had sat there for 15-20 years and think: I wonder if my AngularJS/D3.js/jqGrid/etc. single-page app will even load in a browser 20 years from now, let alone perform as originally intended...
Edit: i.e. NDT'd pieces are not immune from failure (but other factors such as age and in-service time/rotations/re-thread/re-collar runs etc. can provide additional context to NDT results and should correlate more or less with actual failure patterns).
In a single-use part, non-destructive testing mightn't be a complete picture of how the part will perform.
They were more experienced, convinced me I was being pedantic, and I myself wasn't coding on that project anyway. Months later, we had to do a panic refactor as real-world usage immediately made the app fall over.
These are the kinds of bugs I discovered as a kid writing crappy computer games for my friends and I: "floats for everyone! This is way easier!", followed later with: "floats are slow, and I don't understand half my bugs!".
Actually, on this project I found myself explaining several things I'd learnt from recreational games programming. Things that you apparently don't learn in CS (I did EE, so I wouldn't know), or a decade of doing the J2EE business middleware dance.
I just don't think it's healthy to embrace this as an alternative to proper maintenance.
Edit: I'm not singing redo's praises because I think the world should stop using make; fundamentally redo has potential scaling issues for very large codebases, but for the other 99.9% of projects if you want to explore an alternative, you could do much worse than djb redo.
Cross-compiled kernel, bootloader, custom packages, full distro rootfs to final sdcard image without a single Makefile. I owe that man a beer.
The most inspiring thing I saw was work on automated high-throughput materials discovery (I believe the examples at the time were polymers). Instrument operation, experiment design and results capture was done with OWL/RDF... And off it goes: repeatable results where the software drives experiment parameters until a goal or properties are reached. Semantic web tech has a lot of failed promises to answer for but those guys really seemed to make this stuff sing. Seeing what they're modeling/capturing, graph data models might not be such a bad fit, after seeing the contortions we went through in genetic studies trying to capture every facet of every piece of data and its provenance and the provenance of its methods and materials and specimens and specimen preparation and specimen identification and splitting and cloning and so on ad nauseum
Better stop ranting. I do software defined radio at dayjob and some silly osint graph-db driven thing I'm playing with at home to help security audit all packages, binaries and hopefully one day in-memory processes in linux/containers :)
You might scoff, but... btrfs send/receive is insanely fast and painless. To mitigate btrfs shenanigans, snapshots end up on non-btrfs filesystems too. I wrote a tool [2] which produces PGP signatures and sha512sums of snapshots to achieve reproducible integrity measurements regardless of FS.
Of course, in the time it took to polish up snazzer a bit for public release, many [3] other [4] cool [5] solutions [6] have materialized [7]... :)
[1] https://github.com/csirac2/snazzer
[2] https://github.com/csirac2/snazzer/blob/master/doc/snazzer-m...
[3] https://github.com/masc3d/btrfs-sxbackup
[4] https://github.com/digint/btrbk
[5] https://github.com/jimsalterjrs/sanoid/
So it seems a bit weird that somebody posts zed's adventure to HN and then suddenly he has to defend this experiment like it's a billion dollar company.
Perhaps I've missed the egregious thing you were referring to in that thread, but many of his responses to your (numerous!) replies asked for specific stats behind your assertions, and I think it's telling for our industry that you had none (I also have none, yet believe I'll never have to care), which for a person trying to write a fast server by deliberately doing things differently, you can see why he wouldn't be so quick to drop the question and defer to your unsolicited wisdom.
But perhaps I've misunderstood. The crazier antennas are generally down in HF frequencies, below ~30MHz. They're crazier because longer wavelengths put more physical demands on antenna optimization. Making an antenna which efficiently operates a 4MHz chunk of spectrum 144-148MHz means you can build something tuned at 146MHz and only suffer very small performance difference when moving <2% off either side of this range. That's a completely different story to HF freqs, Eg. operating 1.8-5.5MHz is <4MHz of spectrum but that <4MHz chunk represents a 300% increase in frequency moving from lower to upper end, a far cry from the 2% needed before.
So, to summarize, you can definitely and easily start out without any DIY gear, just usin turn-key stuff. Make friends with other amateurs working the higher/easier frequencies. Even in HF, you can still go a long just by draping bits of old wire around the place, there's still a bit of an art to this but not so difficult to learn :)