HNHacker News
TopNewBestAskShowJobs

foob

4,793 karma · joined February 14, 2011

I'm the CAIO at TenantBase (https://tenantbase.com) and a co-founder of Intoli (https://intoli.com). Feel free to get in touch at evan@sangaline.com if you would like to chat.

my public key: https://keybase.io/sangaline my proof: https://keybase.io/sangaline/sigs/uKMqGQ_auE6z2qTdrkpZihnoA-DcJ4BeHb35KLA1KZI

submissionscomments
foob··on One parameter is always enough [pdf]
Nobody evaluates a model based simply on how well it fits the data and its number of parameters; you also look at how well the model parameters are constrained and what the uncertainty bands on the fit are.

The existing practice isn't to blindly look at the number of parameters in a model without considering the actual fit. It's a strawman argument to pretend that that's the case. A model that statistically overfits the data is just as suspect as one that underfits it. If you try to publish a paper about a model with a 0.01 chi-squared value being fit to some data, then it's going to be rejected. It doesn't matter if it's one parameter or one hundred parameters, it's clearly encoding information about the actual dataset rather than being a general model. The model that they present in this paper would have a chi-squared value of essentially zero, and someone would be laughed out of the room if they tried to present it at a conference.

foob··on Building a YouTube MP3 Downloader with Exodus, FFmpeg, and AWS Lambda
I can definitely understand you cringing about the re-encoding, but I just wanted to point out that the reason that we took this approach was simply to give a semi-plausible real-world example of how tools like aws-serverless-express [1] and Exodus [2] can be used to build useful APIs with AWS Lambda [3] and AWS Gateway [4]. These articles are primarily meant to be educational tutorials that people can use as a reference when writing and deploying their own APIs. The whole "converting to MP3" thing is just an easily understandable premise which lays out a clear goal for the tutorials to build upon.

[1] - https://github.com/awslabs/aws-serverless-express

[2] - https://github.com/intoli/exodus

[3] - https://aws.amazon.com/lambda/

[4] - https://aws.amazon.com/api-gateway/

foob··on One parameter is always enough [pdf]
Their equation is cute, but this really isn't remotely surprising and the implications aren't as significant as they imply. The result relies heavily on the fact that their theta parameter has infinite precision. You can encode as much information as you want in a single real number with infinite precision. Think of it this way: a single precision float requires 4 bytes to store while a double requires 8. If all you need is single precision, then you can store two floats inside of one double variable with each occupying 4 of the 8 bytes. Now replace the double with an infinite precision number that takes infinite bytes to represent. Once you have an infinite number of bytes to work with, you can pack in as many floats of finite precision in there as you want. That's basically what they're doing here, they just have a simple closed form expression for decoding it.

The reason that the implications are a bit overblown is that their model is tremendously and chaotically dependent on the value of theta. The plots that they include in the paper require hundreds of thousands of digits of precision in the model parameter. Nobody evaluates a model based simply on how well it fits the data and its number of parameters; you also look at how well the model parameters are constrained and what the uncertainty bands on the fit are. With this model, theta would be completely unconstrained and their uncertainty bands would cover the entire range of the data that they're fitting. It simply doesn't matter how many parameters you have when that's the case, it means that your fit is useless.

foob··on Jeff Bezos announces Amazon is picking up 'The Expanse'
It's actually been reported that Netflix is significantly ramping up their production of sci-fi television shows [1]. 29% of their upcoming shows fall into the sci-fi category. Here's a quote from that article with some more context.

Responding to new data that shows science fiction and related programming powered the biggest viewer share of its genre content in the first quarter of 2018, Netflix is getting ready to aim even more of its seemingly limitless resources toward outer space, committing more money to more new sci-fi projects in a bid to give its still-expanding subscriber base more of what they already love.

[1] - http://www.syfy.com/syfywire/report-netflix-is-about-to-make...

foob··on Proxy, a new JavaScript ES6 feature
One possibility is adding syntactic sugar to objects. I posted elsewhere in this thread about how we used them in Remote Browser, but here's a simple example of adding support for negative indexing to an array. This basically creates a shorthand for array[-n] === array[array.length - n].

    class MyArray extends Array {
      constructor(...args) {
        super(...args);
        return new Proxy(this, {
          get: (target, name) => {
            if (name && name.match && name.match(/^-\d+$/)) {
              return Reflect.get(target, this.length + parseInt(name, 10));
            }
            return Reflect.get(target, name);
          },
          set: (target, name, value) => {
            if (name && name.match && name.match(/^-\d+$/)) {
              return Reflect.set(target, this.length + parseInt(name, 10), value);
            }
            return Reflect.set(target, name, value);
          },
        });
      }
    }
You could accomplish the same thing by creating explicit setters and getters, but proxies allow you to directly intercept and modify the behavior of property access, calling an object like a function, etc. This sort of flexibility allows you to create APIs that wouldn't be possible otherwise.
foob··on Proxy, a new JavaScript ES6 feature
Proxies are definitely fun. We use them extensively in our browser automation framework, which is largely based on the Web Extensions API [1], to implement syntactic sugar in the interface. The general goal behind Remote Browser [2] is to make it really easy to execute privileged code in a web browser, but to otherwise stay out of your way as much as possible. This is accomplished through the heavy use—some might even say abuse—of proxies.

For instance, take a look at this code example of opening a tab.

    await browser.evaluateInBackground(createProperties => (
      browser.tabs.create(createProperties)
    ), { url: 'https://intoli.com' });
The browser.evaluateInBackground() method is part of the Remote Browser API, and it simply evaluates code in a background script context in the browser. The browser.tabs.create() call, on the other hand, is part of the Web Extensions API. The vast majority of Remote Browser's power is provided through remote code execution and the Web Extensions API, so we really wanted to provide a more concise syntax for performing this sort of operation. Using proxies, we were able to make the browser object directly callable as a shorthand for background context evaluation. This code example is exactly equivalent to the previous one.

    await browser(createProperties => (
      browser.tabs.create(createProperties)
    ), { url: 'https://intoli.com' });
That's a bit more concise, but the real sugar is in the next step of abstraction. Instead of explicitly evaluating a function in the background script context, you can simply treat the browser object in the client context as though it were the Web Extensions API browser object in the background script context. This code example is again exactly equivalent, and again accomplished using proxies.

    await browser.tabs.create({ url: 'https://intoli.com' });
The tabs.create method isn't hardcoded into the Remote Browser library. Instead, proxies are used to evaluate any non-local API calls in the background script context. This means that you automatically get full access to the version of the Web Extensions API that your browser supports, whether you're using Chrome, Edge, Firefox, or Opera.

The one last piece of syntactic sugar is the shorthand for evaluating code inside of individual tabs. The syntax is very similar to how you evaluate functions in the background script context, you just need to first use square brackets to indicate the tab ID.

    await browser[tab.id](createProperties => (
      document.getElementById('some-link').innerHTML
    ), { url: 'https://intoli.com' });
This is also implemented using proxies. In fact, the same trap that handles the local Web Extensions API calls handles the tab access, and it differentiates between the two depending on whether or not the property name consists entirely of integers. In both cases, it returns another proxy that is able to handle either additional chaining or function evaluation.

Anyway, sorry that this ended up a lot more long-winded than I was expecting. If this has piqued your interest in Remote Browser though, we made an interactive tour of the project that you should check out [3]. You can actually launch and control browsers from inside of your own browser in the tour, and it explains a lot more about the project philosophy and what you can do with it

[1] - https://developer.mozilla.org/en-US/Add-ons/WebExtensions/AP...

[2] - http://github.com/intoli/remote-browser/

[3] - https://intoli.com/tour

foob··on Tarballs, the ultimate container image format
The packages that Exodus produces are actually quite similar to those introduced in this announcement. Both tools generate simple tarballs that can be extracted anywhere to relocate programs along with their dependencies, and both tools bootstrap the program execution using small statically compiled launchers written in C. They contrast guix pack against Snap, Flatpak, and Docker, but Exodus would probably make a more apt comparison in many ways.
foob··on Make your own sourdough
I'll have to checkout Ken Forkish's book, but, in the meantime, I feel the need to throw out a quick mention of Jennie Shapter's Bread Baker's Bible [1]. I'm not much of a bread eater, but it's one of my favorite cookbooks despite that fact. There is a fair bit of information on the science behind the recipes, and she doesn't take any shortcuts for convenience's sake. Her recipes for English muffins, pizza dough, and many other things have never left me wanting. Some of them take a lot of work to prepare, but they're absolutely worth it.

[1] - https://www.amazon.com/Bread-Bakers-Bible-Traditional-Recipe...

foob··on Machine Learning’s ‘Amazing’ Ability to Predict Chaos
I used to be part of a collaboration at a particle accelerator called RHIC. It was the most powerful accelerator of it's type before the LHC was built, and a lot of really great research was done there. One year, the budget got slashed and there just wasn't enough funding for it to operate. Simons donated a massive amount of his own money, and organized fundraising from other sources, both of which played central roles in the experiments being able to continue on with their research. There's now a road inside of Brookhaven National Lab called "Renaissance" in honor of him to celebrate his contributions. It's actually pretty sad that it was necessary for something like that to happen, but it certainly speaks volumes about Simons.
foob··on Remote-Browser – A browser automation framework based on the Web Extensions API
Thanks! You can use the extension in a CI setup by first installing the remote-browser extension, and then using it in your tests. You can check out remote-browser's own tests for an example of integrating the project with CircleCI [1].

[1] - https://github.com/intoli/remote-browser/blob/master/.circle...

foob··on Remote-Browser – A browser automation framework based on the Web Extensions API
The Remote Browser framework is designed to be an API where interactions occur at a much lower level than they do in a library like Selenium or Puppeteer. In many ways, it's more similar to the Chrome DevTools Protocol than it is to Puppeteer. It's more a tool for building libraries with user-friendly interfaces than it is a user-friendly interface in and of itself. Like you mention, simulating user interactions with the DOM is relatively easy. The idea here isn't that you would implement that sort of thing yourself, it's that higher level libraries will be built which incorporate this sort of functionality for you. We have several projects in the works here at Intoli that fall into this category, and we're looking forward to rolling them out in the near future.
foob··on From Express.js to AWS Lambda: Migrating Existing Node Apps to Serverless
It depends a little bit on whether there's a warm container or not, but I think that the overhead is much closer to ~100 milliseconds than to multiple seconds. The CircleCI artifacts API that I linked to takes less than one second to return an artifact, and there's a lot more going on there than a single request/response pair. The endpoint makes an external request to CircleCI's API, does some processing on the response, returns a 301 redirect, and then the redirect is followed by the client and another request is made to actually download the file from S3. You can test this directly by running

    time curl -L 'https://circleci.intoli.com/artifacts/intoli/exodus/coverage-report/total-coverage.json'

which outputs something like the following (the timing will vary a bit obviously).

    { "coverage": "92.43%" }

    real	0m0.798s
    user	0m0.046s
    sys	0m0.010s
foob··on From Express.js to AWS Lambda: Migrating Existing Node Apps to Serverless
I'm kind of surprised that the article doesn't mention aws-serverless-express [1], a node module provided by Amazon which makes it fairly trivial to expose an existing Express app as a Lambda function via AWS API Gateway. The Lambda handler module basically just consists of the following code.

    const awsServerlessExpress = require('aws-serverless-express');
    const app = require('./app');
    const server = awsServerlessExpress.createServer(app);

    exports.handler = (event, context) => (
      awsServerlessExpress.proxy(server, event, context)
    );
It's really convenient to be able to develop simple APIs using Express, and then to expose them via Lambda functions using aws-serverless-express. After you've done it once or twice, it's really easy to throw together single-purpose APIs and deploy them in a matter of minutes. For example, I was recently frustrated with the fact that CircleCI's API for accessing build artifacts is perpetually broken, and I wrote a little microservice to expose the same functionality [2]. It's definitely awesome to be able to deploy one-off projects like that without needing to worry about any sort of server maintenance.

[1] - https://github.com/awslabs/aws-serverless-express

[2] - https://intoli.com/blog/circleci-artifacts/

foob··on A Crude Personal Package Manager
While we're on the topic of homebrew package managers, I think some people here might be interested in checking out Exodus [1] [2]. A lot of the project philosophy actually overlaps quite a bit with that of qpkg. Both projects seem to be focused on simplicity, produce tarballs which can be extracted anywhere and run on any Linux machine, and avoid collisions without a centralized database. Unlike qpkg, however, Exodus isn't concerned with compilation and does automatically handle dependencies. It does so by using the system linker to discover required libraries, and then creating small wrapper executables which invoke the linker with the necessary library paths.

I could actually see it being used in conjunction with a script like qpkg, where the script handles compilation and then invokes Exodus to produces self-contained bundles with all of the necessary dependencies. Exodus is generally agnostic to how software it bundles was initially installed or compiled, which has lead to some interesting use cases. For example, we've had a number of people use Docker images to install software from Debian or Alpine package repositories, and then use Exodus to repackage single applications to run in microcontainers.

- [1] https://github.com/intoli/exodus

- [2] https://intoli.com/blog/exodus-2/

foob··on Show HN: HN Domain Leaderboard
Very cool, thanks for sharing! I did a somewhat similar analysis a while back [1], and I found that many of the top domains either had a YC affiliation or corresponded to extremely well-known companies or organizations. This made me interested in finding lesser known blogs that also produce high quality content. I tried to identify these by putting a limit on the number of unique users who had submitted content from each domain. My thinking here was that something like the GitHub blog would have submissions from many users, while smaller personal blogs would probably be mostly self-promoted. Using this approach, I was able to turn up some pretty interesting blogs that I had never heard of before.

I think it could really increase the usefulness of HN Domain Leaderboard if you added some additional filtering capabilities. Filtering based on the category would probably be pretty easy because you have that information there already, but perhaps also consider some measure of how broadly promoted each domain is. The time range option is already pretty cool, and I'll bet that a few more options would make it even more fun to play around with.

[1] - https://intoli.com/blog/pareto-optimal-blogs/

foob··on The Quantum Thermodynamics Revolution (2017)
> Yes, photons are said to go through both slits because there does not exist any information which would distinguish the paths. As soon as you arrange an experiment which provides such information, the chosen path becomes clear.

This isn't correct. Photons are said to go through both slits because they travel like waves and actually go through both slits. How would your interpretation account for the fact that a single photon at a time fired through a double slit still produces an interference pattern?

foob··on Exodus – Painless relocation of Linux binaries without containers
Thanks! For Python, `cx_Freeze` should be able to accomplish what you describe in a cross-platform way. I don't know of any similar PHP tools off the top of my head.

[1] - http://cx-freeze.readthedocs.io/en/latest/index.html

foob··on Exodus – Painless relocation of Linux binaries without containers
The short answer is that I used gifine [1]. The longer answer is that I researched this extensively and wrote a post called Terminal Recorders: A Comprehensive Guide that covers many of the available options [2]. If you're trying to pick a tool to use, you might find that post useful. Gifine was the right choice for me, but there are pros and cons for each tool.

- [1] https://github.com/leafo/gifine

- [2] https://intoli.com/blog/terminal-recorders/

foob··on Isso – a commenting server similar to Disqus
That sounds like it will be an extremely interesting post, I'm looking forward to it.

I can chime in with an anecdote about the first one on your list, RemarkBox [1]. We've been using it as the comment system on the Intoli Blog, and we couldn't be happier [2]. We first heard about RemarkBox from one of foxhop's posts here on Hacker News that we happened to see right around the time that we were thinking about adding comments to the blog anyway. We liked the idea of working with a smaller company instead of Disqus, so we got in touch and signed up to be what I suspect was one of the site's first paying customers.

The reason that I feel compelled to mention this here isn't that the comment system has worked well for us—although it has—but rather that it has been a real pleasure to communicate with foxhop and to be involved with RemarkBox at this relatively early stage. He's been extremely responsive to our feature requests, wishes us congratulations when one of our articles does well, lets us know if we forget to approve a comment, etc. Overall, it's just been a really pleasant experience. Definitely check them out if you're thinking about adding comments to your site!

[1] https://www.remarkbox.com/

[2] https://intoli.com/blog

foob··on CSS Paint API: New possibilities in Chrome 65
It's a joint W3C Technical Architecture Group and CSS Working Group initiative, you can find a lot of details in this post from the Opera developer blog [1].

[1] https://dev.opera.com/articles/houdini/

foob··on It is not possible to detect and block Chrome headless
That test actually wouldn't work:

    > navigator.webdriver
    true
    > Object.getOwnPropertyDescriptor(navigator, 'webdriver')
    undefined
As you say though, it's a cat and mouse game and you could always override the behavior of getOwnPropertyDescriptor() if it were used in a test.
foob··on Analyzing One Million Robots.txt Files
Just to clarify, this was intended as a tongue-in-cheek critique of tech companies that actively project superficial images designed to appeal to specific hiring demographics. I'm sorry if the meaning didn't come across as clearly as I had hoped for, but the statement was meant to be a condemnation of sexism and ageism rather than an endorsement.
foob··on Volvo is reportedly scaling back its self-driving car experiment
There's no question that machines have faster reaction times, but the other part of the equation is predicting how much a vehicle should slow down in any given situation in order to reduce braking time. The parent was talking about looking at a fairly complicated scene and making a decision about reducing their speed so that they could stop more quickly if it becomes necessary. We all know that there's a big difference between driving by a child playing aimlessly and somebody walking predictably on a sidewalk. It's much more difficult for an automated system to make these sort of predictions than it is for them to react quickly to emergency situations. My guess is that automated driving systems will have to err on the side of caution and that this will result in seemingly (to the passenger) unnecessary slow-downs. I don't personally think that this is a huge problem, but it's fair to say that reaction time is only one one of several significant factors in avoiding accidents.
foob··on Is it time to change the undergraduate physics curriculum?
It's certainly also true for astronomy, quantum gravity, and most of the condensed matter physicists that I know. Which subfields don't involve at least Monte Carlo simulations or writing code to analyze data and/or control experimental apparatus?
foob··on Is it time to change the undergraduate physics curriculum?
I'm a former physicist and I was expecting this article to make a different point (particularly because it's on HN). Physics education is heavily centered around--well... physics--but the majority of what most physicists do on a daily basis is really software engineering, and most physicists are woefully unprepared for that. The majority of them are talented enough to figure a lot of it out as they go, but best practices like testing and code reviews were basically unheard of when I was in the field. I wasn't in a small laboratory experiment either, I'm talking about a large-scale collaboration that cost hundreds of millions of dollars and involved thousands of people. A simple bug in someone's code could literally have a major impact on the field. Another comment asked whether universities are too focused on preparing people for academia rather than a job, but I honestly think that they're not even really focused properly on preparing them for academia.

Teaching physics is a hard problem because there's a long and rich history of building upon previous advancements. You can't teach somebody quantum chromodynamics without first teaching them quantum electrodynamics, quantum mechanics, relativity, thermodynamics, classical mechanics, and all of the mathematical methods that come along with these. It's hard to legitimately get anywhere close to the cutting edge in graduate school, let alone as an undergraduate. I hope that this has changed somewhat in the last decade, but most of the teaching methods in these subjects have historically been unchanged since the 70s. I strongly believe that it would be better preparation, for both academia and the real world, to integrate much more of an emphasis on simulation, visualization, and numerical methods into these requisite courses. These would be far more constructive for people developing skills that they simply can't get from doing problems on paper.

foob··on Extensions in Firefox 58
The timing on this is quite coincidental, I just published a short article this morning about adding support to Selenium for installing WebExtensions to Firefox profiles [1]. I write and use extensions frequently for web scraping and browser automation, and I've really been thrilled about what Mozilla has done with the WebExtensions API. It was a quite miserable process to write cross-browser extensions in the past, but that basically changed overnight when Mozilla first announced the WebExtensions API. I think that this is ultimately going to have an immensely positive impact on the extension ecosystem.

[1] - https://intoli.com/blog/firefox-extensions-with-selenium/

foob··on [dead]
This is the killer quote:

> The agreement—signed by Amy Dacey, the former CEO of the DNC, and Robby Mook with a copy to Marc Elias—specified that in exchange for raising money and investing in the DNC, Hillary would control the party’s finances, strategy, and all the money raised. Her campaign had the right of refusal of who would be the party communications director, and it would make final decisions on all the other staff. The DNC also was required to consult with the campaign about all other staffing, budgeting, data, analytics, and mailings.

It was already known, since Politico broke the story, that the Hillary Victory Fund was "essentially... money laundering," but I don't believe that there was any public information that indisputably showed that the Clinton campaign was controlling the DNC throughout the primaries. It was widely suspected, obviously, but this new information from Donna Brazile seems to show that there was actually a formal agreement in place.

foob··on The Source Code for NYC's Forensic DNA Statistical Analysis Tool
I haven't had a chance to dig into the analysis part of the code yet, but here's an amusing variable assignment that jumped out at me in `FST.Common/ComparisonData.cs`.

    compareMethodIDSerializationBackupDoNotMessWithThisVariableBecauseItsAnUglyHackAroundACrappyLimitiationOfTheFCLsCrappyXMLSerializer = comparisonID;
foob··on Why physicists still use Fortran (2015)
One of the key points of the article is that there is a lot of legacy code written in Fortran. As a former high energy physicist, I have an anecdote here that some people might find interesting.

There was a library written in Fortran called CERNLIB which included a broad variety of miscellaneous numerical algorithm implementations (e.g. minimization, integration, special functions, random number generation) [1]. I couldn't tell you exactly when the library was first released, but my best guess would be the early 80s. It can't possibly be later than 1986 when PAW was initally released [2]. The field has since transitioned from the Fortran based PAW to the C++ based ROOT since then, but many high energy collaborations still rely on CERNLIB for their own analysis frameworks (keep in mind that many of these experiments had been in planning and development stages for over a decade before they actually turned on).

The thing about this that I find interesting is that compiling CERNLIB has become a lost art and that this fact has had far reaching consequences. The last available binaries were compiled with GCC 4.3 in 2006 and packages are only available for Scientific Linux 4 [3]. This crucial dependency has led to collaborations using extremely outdated Linux distributions and GCC versions in their computing facilities. The majority of analysis code is written in C++, but not even C++11 additions can be used because everything is frozen on GCC 4.3. Nobody can even run the analysis environment on their local machines without resorting to the use of virtual machines running SL4. It was really a nightmare to deal with.

[1] - https://en.wikipedia.org/wiki/CERN_Program_Library

[2] - https://en.wikipedia.org/wiki/Physics_Analysis_Workstation

[3] - http://cernlib.web.cern.ch/cernlib/version.html

foob··on Designing the Wayback Machine Loading Animation
Thanks for letting me know! That was accidental and it should be fixed now.
← PreviousPage 4 of 9Next →