HNHacker News
TopNewBestAskShowJobs

csirac2

552 karma · joined May 5, 2013

I make software & systems: software defined radio, embedded Linux, build/test automation, data viz, web stuff. Electronics & infosec enthusiast.
submissionscomments
csirac2··on Unhackable Kernel?
NICTA/UNSW have apparently started working on a QubesOS port

https://my.cse.unsw.edu.au/thesis/thesis_topic_details.php?I...

http://sel4.systems/pipermail/devel/2015-March/000312.html

csirac2··on The sad state of web app deployment
I agree, and I do (if I end up actually keeping a piece of software running).

It's also quite likely that a pre-built docker image isn't going to satisfy everyone's deployment requirements.

Starting out with a philosophy that you'll make all the configuration and security choices for every user of your project isn't a great start, and that's how I read all these hard-coded assumptions.

csirac2··on The sad state of web app deployment
> But I still think that using docker is probably better than having the user do that manually

When it comes to deploying, supporting and securing complex web application software, it is very rare that an out-of-the-box reference image from the project or vendor will have much relevance to your own environment's security and configuration standards.

Unless you're paying for support, and the vendor insists you'll have to suck it up and accept whatever madness their support agreement entails.

I'm all for convenient docker images, it just frustrates me to see the total and utter lack of imagination demonstrated by some projects that can't imagine their users ever locking down their database access, or running with SELinux, or separately patching the software/service components they depend on, etc.

And again, I don't expect the reference docker image for a given project to allow all of that flexibility and configurability, it's more the fact that these possibilities seem to be actively sabotaged with reckless abandon.

csirac2··on The sad state of web app deployment
> But in reality of course it's not laziness, it's efficiency. They can get on with other tasks sooner than before.

Indeed, and I use Docker myself to improve my own efficiency, and I've seen great stuff built with Docker that has been architected well.

However, it is painfully obvious that without the pressures that used to force most developers to keep their shit sane, there's now more workload for serious users that actually need to untangle this mess in order to support, secure and get stuff deployed.

csirac2··on The sad state of web app deployment
Well yes, the fact they chose to require that the webapp have a superuser database account is not docker-specific, but the manner (and it should be said: ease) with which docker is used to neatly package an opaque ball of string creates a trend of ever more difficult-to-untangle software making it way harder than it should be to properly deploy apps conforming to your environment's security, config management standards.

I mean, when I first evaluated this app, it didn't even have a sane launcher script. Instead, ~15 lines of ruby config and a ruby script that I had to follow just to discover it was a weird, idiosyncratic way of executing "docker run".

Edit: I am not saying that Docker is somehow inherently flawed in this respect. It's solved a lot of problems for me and you can build great, well-architected stuff with it. But there is a trend to use it to wallpaper over poor development and deployment practices.

csirac2··on The sad state of web app deployment
If you're lucky enough to run in a production environment that has zero configuration policy or minimum security requirements, then yes, granting your webapp superuser privs to the database mightn't be a deal-breaker.

For everyone else, it's giant shit sandwich, and a constant reminder that the project couldn't arsed untangling their own weird idiosyncrasies from their distributed app, and that this crap wouldn't fly with a traditional distro package, or wouldn't have mattered with the more traditional tarball/REQUIREMENTS/INSTALL.txt that made way fewer assumptions about the end-user's environment.

I love docker, but it's letting people get away with murder.

csirac2··on The sad state of web app deployment
I love Docker; it's solved a lot of problems for me. But this article highlights the laziness that it can enable: there has to be a middle ground between traditional package management and all of the curation/QA that goes into getting stable releases out to multiple distributions, versus chucking shit over the wall into a git repo with a Dockerfile and calling it done.

It's a better, more isolated mess - but for anyone trying to enforce configuration policy into all of the running services in their environment, untangling gortesquely basic shit like not granting superuser privs on a database to your webapps - which would never fly in a traditional distro package - becomes even more work than the old source tarballs with INSTALL.txt.

csirac2··on The sad state of web app deployment
I've tried to install the exact software the author of this article was trying to install, and he/she/they didn't even encounter the insanity that turned me off.

There's giving up on traditional package management, and then there's throwing yourself off the cliff of decent release/config/dependency management that used to be a side-effect of traditional package management.

"Works for me on my computer" is now "Works for me with my Dockerfile [ and all the configuration assumptions I couldn't be bothered untangling from my app ]"

csirac2··on The sad state of web app deployment
As a happy Docker user with a lot of respect for what Docker has achieved, what concerns me is that Docker is clearly an enabler for people to totally abandon proper release and dependency management along with sane sysadmin friendly configuration.

I can't tell you the number of times I've hoped to use a public docker image or just the Dockerfile, only to spend hours futzing with it because I was unhappy with the grotesquely insecure configuration or because I needed to work around a bunch of assumptions that are invalid once I've tuned the Dockerfile for my environment.

csirac2··on P values are not as reliable as many scientists assume (2014)
Luckily some disciplines have a journal of negative results :-)

Eg. http://www.jnr-eeb.org/index.php/jnr

csirac2··on A Cryptologist Takes a Crack at Deciphering DNA’s Deep Secrets (2006)
This is a pretty old article, and well before this time markov/bayesian stuff was already in full swing for bioinformatics, AFAIK. Biology draws from many disciplines, and it's taken a lot of computer science to get to where it is now. But I think that the infosec community generally is pretty underwhelming in its application of statistical methods to big data. People seem to have limited imaginations when it comes to pattern matching, and too quickly jump to heuristics and "machine learning" without properly understanding the limitations of the data and techniques they're working with.

    To paraphrase provocatively, 'machine learning is statistics
    minus any checking of models and assumptions'.
    ~Brian Ripley, 2004 [1]
We should really check out some of the stuff bioinformatics people have been doing. That's not to say that cryptography has more to give to biology, quite the contrary. But I think the broader infosec community should try harder when it comes to doing stats properly. I think there is a lot more in common between bioinformatics and infosec analytics than most (on either side) realize. Huge amounts of messy, un/semi-structured data being the common theme.

In my (admittedly limited) experience, I find that wetlab biologists tend to have a better understanding of the limitations of the experiments they're setting up, in terms of interpreting the significance, limitations and meaning of their results at least - than the computer science types playing in biology. But good stats is the one universal skill that all good scientists should master.

I spent a few years where part of my work was supporting evolutionary geneticists manage and figure out their data. Enormous amounts of data. So much raw data that the very large scientific institution I'm sure you've heard of couldn't accommodate it in their normal data repositories (at least not if every researcher started doing so). Having multi-million dollar research projects conclude and then archived to a handful of redundant sets of HDDs curated by teams and departments that now no longer exist is unnerving (supposedly things are better these days).

In any case, it's a fascinating area to work in for budding computer scientists. There are many problems to sink your teeth into, even basic stuff like repeatable analyses on research projects running for more than 10-15 years. "Yes, there are faster/more robust/more appropriate MCMV (Markov Chain Monte Carlo) solutions out there but we don't feel like re-running a decade's worth of work to ensure they're comparable". But I digress.

It amused me that each research team seemed to be either computer-science heavy (and ignorant about the quality or nature or "ground truth" of their data and the possibly embarrassing impacts on the signals they thought they were seeing), almost wilfully ignorant or at least dismissive of "old-school" biology methods and data...

... or were very wet-lab savvy and grounded in the raw biology side of things but limited their research questions according to the types of analyses they were comfortable doing themselves, using well-established methods, tools and software that would run on their biggest $20k PC under someone's desk. Although that was seriously changing by 2013 when I left (partly due to nice big web-UIs in front of HPC clusters, partly due to better mixing of team capabilities).

Now, it's easy to dismiss the latter group as just being not "up to date" or unskilled in "big data" (recent biology grad programmes are hopefully improving). But what was surprising was the obvious opportunities (to me) that highly CS-capable teams (inside and outside of the organisation I worked for) seemed to be deliberately staying away from. On more than one occasion I asked some of these people why they were ignoring the rich, machine-readable, highly-curated datasets from more traditional biology science (Eg. taxonomy/ID/phenotype/ecology info).

The answers are varied: they aren't even aware this data exists in machine-readable form; they seem unfamiliar with the very biology discipline they're playing in (they're more obsessed with algorithms); but a disturbing trend seemed to be a generic distrust of manually curated data. I guess because there's no repeatable algorithms to reproduce human curators :-)

Now, I'm not saying everyone should go out and use that data as a source of truth, but it should at least be a reference to sanity-check or compare to. And imagine if you actually did augment your sequence data with this stuff - one (admittedly naieve) thought was that there might be a fun way to look for at least some phenotype influences in a given genome, "for free", without having to start new wetlab work - i.e. just repurpose existing data.

When countering the aversion to manually curated data, I tried to point out that the very "garbage in" being dealt with - the genome sequences, spatial records, environmental layers etc. - all involve an element of human curation. I mean, that's how we end up with nearly 20% of non-human genomes containing human DNA [2] in the first place. And that's before we even consider whether stuff in the "good" 80% are even representative of the species group(s) of interest - you wouldn't believe how easy it is for people in the field (or the lab!) to get species identification completely and utterly wrong. Or scrape up DNA samples of some infection or parasite instead of the actual thing itself...

... and yet even with all these imperfections in the molecular data, we find that some of the old-school morphological (phenotype-driven) taxonomies are still 80%+ unchanged (hazy on the numbers now) when reconstructed with purely DNA-derived phylogenies. And that's on groups where we're still fighting over what the best DNA markers are to build those molecular phylogenies in the first place! A lot of people don't realize that the same DNA evidence can give different pictures of the evolutionary tree, depending on what genes you're picking on...

Apologies for the over-sized rant, it's been a while since I had to think of this stuff.

[1] http://stats.stackexchange.com/questions/6/the-two-cultures-...

[2] http://www.nytimes.com/2011/02/17/science/17genome.html

csirac2··on Genode – Operating System Framework
I suspect you'd run unikernels under Geode. Instead of targeting Xen resources and Xen event/message channels and Xen security modules/FLASK, MirageOS would target Geode instead but perhaps have slightly higher-level resources and better-featured interfaces to lean on.

But I haven't quite had time to figure it out myself yet; I've been interested in exploring Geode for a while.

csirac2··on Important Notice Regarding Public Availability of Stable Patches
I can appreciate that it's been a bit like that, but isn't the truth really somewhere in between?

Oversized egos are a problem at the best of times in open source, especially in the kernel. Which has traditionally had a culture of casual indifference toward the security priorities of PaX/grsec, and some would say toward security generally (I don't buy that - there's a lot of concern for integrity in the kernel, and that at least buys you some security).

Whilst it would have been great if Spender could have completely changed his personality, outlook and perspective on the world so that he could commandeer the linux kernel from the outside-in toward his vision of a safer kernel, is that really compatible with the wider project? And so, it's equally depressing there hasn't been more recruitment from the kernel side - the side that actually has resources - to figure out a path toward getting PaX stuff mainlined from the inside-out.

I guess my point is: grsec has one focus, one priority. But the Linux kernel is a project that has a much more vast, broader scope of competing priorities to untangle, and I just can't see that such an enormous, busy and byzantine project ecosystem truly will ever see a clear mandate to get something as disruptive and single-minded as grsec mainlined.

Not only would Spender have to become a different person, but the core kernel teams as well.

In any case, I've reduced my patreon donations and put what little that is towards grsecurity. It's certainly that important and useful to me.

csirac2··on Ask HN: What role do you identify as in your occupation?
My contract says my title is software engineer, but titles don't mean much in a small company. Which I kind of like.

That said, I didn't study Software Engineering at university. And so, in almost 10 years in various roles writing software, this is the first one in which I've allowed this title to apply (previously I've had: Systems Analyst, Software Developer, Bioinformatics Technologist, Instrumentation Engineer, etc).

I also didn't study systems engineering (which is a real discipline responsible for remarkable achievements such as taming the complexity of massive systems such as modern airliners), but not that Systems Engineering, I mean the wishy-washy, everyone's-an-engineer sense of Systems Engineer as typically used in software land - that's what I think a more accurate title for my role would be.

I work on the hardware for our widget, as well as the user-facing software. I set up our automated build system that builds everything including our ARM Linux OS image, including kernel and bootloader. I established the production and Q/A checklists. I actually do some of the assemblies. I determine what components we use, and I work with real engineers to validate and test and update customer spec sheets accordingly. I answer questions like, "what kind of power system would let us get XX hrs runtime at YY% duty cycle".

I have a complete picture in my head of the product, its capabilities and limitations. When it goes wrong, at any level, I'm the one that has to fix it. I've never worked with a Systems Engineer at a big tech company but it seems as if this is the sort of scope they work with. Except they probably don't have to write the userland code as well :)

csirac2··on Web Design: The First 100 Years (2014)
URLs are a big deal, and I absolutely loathe systems which actively sabotage likability on the web.

That said, it really was an offline app. We could've done it as a native Qt or .NET app. Even so, the standard AngularJS document fragment paths supported bookmarking and so on.

csirac2··on Web Design: The First 100 Years (2014)
We used the Angular document fragment URLs, so it was possible to copy-paste and bookmark things.
csirac2··on Web Design: The First 100 Years (2014)
Around 2012 I worked with a team migrating some content from a very large static HTML site dating back to 1992. We scoffed at the awful ad-hoc nature of it all, just a pile of static hand-coded HTML pages.

But the 2002-2005 stuff had aged much worse. At some point there was fancy site generator that had used javascript for everything, and the javascript apparently only worked properly in IE6. So most of the navigation was busted in a modern browser, and needed special scrapers to parse out what should have been plain old <a href...> tags.

Now, I regularly think back to that crusty old HTML3 static site that had sat there for 15-20 years and think: I wonder if my AngularJS/D3.js/jqGrid/etc. single-page app will even load in a browser 20 years from now, let alone perform as originally intended...

csirac2··on SpaceX CRS-7 Failure Investigation Teleconference Thread
I've only been exposed to this in oil & gas, but there is a problem that non-destructive testing of drill stems does not test for all failure modes

Edit: i.e. NDT'd pieces are not immune from failure (but other factors such as age and in-service time/rotations/re-thread/re-collar runs etc. can provide additional context to NDT results and should correlate more or less with actual failure patterns).

In a single-use part, non-destructive testing mightn't be a complete picture of how the part will perform.

csirac2··on Show HN: Online IEEE 754 playground
I was working with someone who thought I was stupid for insisting on integer or fixed-point maths for what I insisted would be a particularly troublesome piece of functionality.

They were more experienced, convinced me I was being pedantic, and I myself wasn't coding on that project anyway. Months later, we had to do a panic refactor as real-world usage immediately made the app fall over.

These are the kinds of bugs I discovered as a kid writing crappy computer games for my friends and I: "floats for everyone! This is way easier!", followed later with: "floats are slow, and I don't understand half my bugs!".

Actually, on this project I found myself explaining several things I'd learnt from recreational games programming. Things that you apparently don't learn in CS (I did EE, so I wouldn't know), or a decade of doing the J2EE business middleware dance.

csirac2··on Docker: Not Even a Linker
Yeah, I totally get it - I realized I'm a hypocrite when I posted this; just a few weeks ago I put a custom build environment together with an obscure version of gcc because I'm dealing with some code that depends on some of those "bugs".

I just don't think it's healthy to embrace this as an alternative to proper maintenance.

csirac2··on Docker: Not Even a Linker
I'm a huge fan of docker, but I've also drawn a comparison to linking in the past: some people are using it to defer (not solve!) Dependency management and distro ecosystem complexities. Fossilizing dependencies is not the future; that's just like static linking. And just as in static linking, if you hide in the corner, avoiding your distro's dependency and patch and package management systems, this ultimately creates more avoidable pain.
csirac2··on Recursive Make Considered Harmful (1997) [pdf]
I spent a week pulling apart our old build scripts (which were dysfunctional) and putting it all back together again with something new. I'd have done almost the same amount of work if I had stuck with make. In fact I spent two days doing just that prior to my redo adventure: I gave redo a whole day, and found I'd achieved a lot very quickly despite never having used it before. So I completed my spring cleaning with redo.

Edit: I'm not singing redo's praises because I think the world should stop using make; fundamentally redo has potential scaling issues for very large codebases, but for the other 99.9% of projects if you want to explore an alternative, you could do much worse than djb redo.

csirac2··on Recursive Make Considered Harmful (1997) [pdf]
I recently spent a week replacing all Makefiles in a reasonably complex embedded arm project I maintain with apenwarr's redo [1] (revamping our cobbled-together CI).. Which allows recursive make structured .do files, without tbe downsides - global dependencies and build state is cached across invocations in a throw-away .redo directory. It's an absolute joy to be able to stamp things that aren't files (eg docker containers) and just use one syntax (script of your choice) to orchestrate builds.

Cross-compiled kernel, bootloader, custom packages, full distro rootfs to final sdcard image without a single Makefile. I owe that man a beer.

[1] https://github.com/apenwarr/redo

csirac2··on Ask HN: What are you working on?
That's awesome. Would love to know more. I was (only tangentially, really it was my manager) involved with (one of many) LIMS evaluations at CSIRO. I'm under the impression that "Generic" LIMS have awful track records in Australia at universities and research orgs like CSIRO: despite custom software projects also having poor cost/performance outcomes generally, historically it seems there aren't any great examples of LIMS aimed at diverse multidisciplinary research environments which have delivered better outcomes at less cost than even the most horribly expensive over-budget bespoke systems places like this churn out. There is a trail of train wrecks that is failed LIMS projects in this sector, largely from manufacturing/forensics sector thinking their LIMS fits all...

The most inspiring thing I saw was work on automated high-throughput materials discovery (I believe the examples at the time were polymers). Instrument operation, experiment design and results capture was done with OWL/RDF... And off it goes: repeatable results where the software drives experiment parameters until a goal or properties are reached. Semantic web tech has a lot of failed promises to answer for but those guys really seemed to make this stuff sing. Seeing what they're modeling/capturing, graph data models might not be such a bad fit, after seeing the contortions we went through in genetic studies trying to capture every facet of every piece of data and its provenance and the provenance of its methods and materials and specimens and specimen preparation and specimen identification and splitting and cloning and so on ad nauseum

Better stop ranting. I do software defined radio at dayjob and some silly osint graph-db driven thing I'm playing with at home to help security audit all packages, binaries and hopefully one day in-memory processes in linux/containers :)

csirac2··on Comparison of Attic vs. Bup vs. Obnam
As a heavy btrfs user backups have always been on my mind. I run a lab with a handful of busy VMs, all using btrfs. I was frustrated that there were no backup solutions (at the time) which leveraged btrfs, so I created snazzer [1] (one day soon it will support ZFS).

You might scoff, but... btrfs send/receive is insanely fast and painless. To mitigate btrfs shenanigans, snapshots end up on non-btrfs filesystems too. I wrote a tool [2] which produces PGP signatures and sha512sums of snapshots to achieve reproducible integrity measurements regardless of FS.

Of course, in the time it took to polish up snazzer a bit for public release, many [3] other [4] cool [5] solutions [6] have materialized [7]... :)

[1] https://github.com/csirac2/snazzer

[2] https://github.com/csirac2/snazzer/blob/master/doc/snazzer-m...

[3] https://github.com/masc3d/btrfs-sxbackup

[4] https://github.com/digint/btrbk

[5] https://github.com/jimsalterjrs/sanoid/

[6] https://github.com/lordsutch/btrfs-backup

[7] https://github.com/jf647/btrfs-snap

csirac2··on Dear Paul
To be fair I just trawled through that thread and found your posts most obsessive, and Zed seemingly sticking to technical discussion with (for him) not all that much profanity. Now, not that I disagree with your assertions (epoll/poll overhead << actual I/O processing) but I don't know why there was so much traffic on this. I don't know zed of course but I used to follow him before the term follow was owned by social media. Trying weird/apparently-pointless shit like measuring how polling scales is Zed's thing. Trying it for the goal of making the absolute fastest http server seems positively sane.

So it seems a bit weird that somebody posts zed's adventure to HN and then suddenly he has to defend this experiment like it's a billion dollar company.

Perhaps I've missed the egregious thing you were referring to in that thread, but many of his responses to your (numerous!) replies asked for specific stats behind your assertions, and I think it's telling for our industry that you had none (I also have none, yet believe I'll never have to care), which for a person trying to write a fast server by deliberately doing things differently, you can see why he wouldn't be so quick to drop the question and defer to your unsolicited wisdom.

csirac2··on How to Write Unmaintainable Code – Naming
It's been a while since I did much perl, but FWIW I seem to recall doing (undef, my $foo) = do_stuff(). Perhaps modern perls have tightened up allowing undef as the target of an assignment.
csirac2··on Wide-band WebSDR: Control a short-wave receiver over the web
You don't need an insane antenna setup to start out. A lot of Amateur radio people love optimizing their antennas for efficiency, skywave propagation, frequency agility etc. but it's not required for basic operation especially in higher frequencies. All radios, regardless of brand, have essentially the same laws of physics driving their antenna requirements (well, some have fancy in-built antenna tuning capability but we're talking small budgets here).

But perhaps I've misunderstood. The crazier antennas are generally down in HF frequencies, below ~30MHz. They're crazier because longer wavelengths put more physical demands on antenna optimization. Making an antenna which efficiently operates a 4MHz chunk of spectrum 144-148MHz means you can build something tuned at 146MHz and only suffer very small performance difference when moving <2% off either side of this range. That's a completely different story to HF freqs, Eg. operating 1.8-5.5MHz is <4MHz of spectrum but that <4MHz chunk represents a 300% increase in frequency moving from lower to upper end, a far cry from the 2% needed before.

So, to summarize, you can definitely and easily start out without any DIY gear, just usin turn-key stuff. Make friends with other amateurs working the higher/easier frequencies. Even in HF, you can still go a long just by draping bits of old wire around the place, there's still a bit of an art to this but not so difficult to learn :)

csirac2··on Wide-band WebSDR: Control a short-wave receiver over the web
I hate to be a buzz kill, but some of those Baofengs really need harmonic suppression filters installed. They're popular here in Australia, but I saw a friend put his on a spectrum analyzer and saw 2nd harmonic at ~-5dBc! I know they're only low power, so it's likely harmless, but it's stuff like this which makes it harder for amateur radio to retain the rights we have left on the spectrum, and raises the noise floor for everyone.
csirac2··on In Flight
The technical limitation is the basic geometry of GPS. Satellites are at ~20,000km, aircraft at ~10km max. And you want 0.0003km resolution in altitude. We barely even have 0.003km in the lat/long plane, and it's only that good because of fancy DSP algorithms which lean on special modeling of atmosphere and the satellites themselves - I.e. assumptions about a 3rd-party system, wholly inappropriate for safety critical stuff. In fact, cruising altitude (~"full span" indication) is roughly the magnitude of the specs for error in lat/lang back when GPS was first commissioned (IIRC).
← PreviousPage 3 of 9Next →