HNHacker News
TopNewBestAskShowJobs

davidfischer

420 karma · joined August 2, 2010

Advertising, security, and privacy @ Read the Docs / EthicalAds

To contact me about Read the Docs or Ads, use my first name @ our domain names (either readthedocs.org or ethicalads.io).

submissionscomments
davidfischer··on Malvertising on Google Ads
I work on ads, but not for Google and FWIW, I've only been able to reproduce a few of these malvertising reports. However, I wouldn't be surprised if there were additional targeting parameters on these campaigns. Rather than targeting just anybody searching for VLC, Blender, or Audacity, these malvertisers want to target folks more likely to click a "download now" malvertisement. Maybe only target older users, non-developers, Windows users only, or a number of other facets that probably have a higher rate of installing malware. I have no knowledge if these folks are doing this, but that's what I'd do if I were a scummy advertiser shilling malware. If they can avoid wasting their ad budget on sophisticated users, I'm sure they will.
davidfischer··on Privacy Teardown: Search Engines
Off the shelf privacy and tracking analyzers run against large and small search engines to see what they are doing with your data.
davidfischer··on Show HN: EthicalAds – Privacy-first ad network for developers
We're already on the EasyList (a big blocklist) so it should already be blocked.
davidfischer··on Show HN: EthicalAds – Privacy-first ad network for developers
Can you give a few more details of what you mean?

It's true that Do Not Track (DNT) is not a true standard in terms of implementation and intent and different folks mean different things when setting the DNT flag. However, when building DNT for EthicalAds and for Read the Docs, we followed the EFF's implementation guide for DNT[1]. This means a number things including:

* We do not store personally identifying information when users are merely browsing a site with our ads.

* We rotate our logs in less than 10 days

* We do not set cookies on ad requests. We also don't use some non-cookie alternative. Obviously for publishers or advertisers, logging into our backend requires a login cookie.

[1]: https://github.com/EFForg/dnt-guide

davidfischer··on How to advertise to developers: deep dive into paid developer marketing
Ads don't need to track folks to be targeted and effective =).
davidfischer··on How to advertise to developers: deep dive into paid developer marketing
I work on EthicalAds, one of the ad networks mentioned in the post. We focus on targeting relevant developer-focused ads to content rather than to individuals.

The post is great. My only nitpick is that the author mentions that they think retargeting will go away. The scope of retargeting might be reduced especially as third-party cookies stop working but I don't think retargeting will completely go away. Big ad companies with whom visitors have a direct relationship with (Google, FB) will always be able to do it. They know the pages on their sites and those linked from their site that visitors visit and they will use it to target ads. They don't need third-party cookies since they have first-party cookies.

Networks that just do ads and don't have relationships with visitors except for advertising will find it harder and harder. At EthicalAds, we don't do retargeting at all because of our privacy focus, but many ad networks will find retargeting more challenging (less precise, they can always fingerprint users) to do. Doing away with it is a good thing from a privacy perspective but it probably will further entrench the big players.

davidfischer··on Invasive ad targeting is bad for journalism and other high-quality publishers
While there's still a ways to go, NYT specifically is improving here: https://open.nytimes.com/to-serve-better-ads-we-built-our-ow...
davidfischer··on Invasive ad targeting is bad for journalism and other high-quality publishers
In some ways, "Patreon for good journalism" is Substack's pitch. It has its good and bad aspects but no doubt it's a very different monetization model than the ad-driven model.
davidfischer··on Invasive ad targeting is bad for journalism and other high-quality publishers
SEO content farms definitely aren't going away. In some specific cases, they aren't even that bad. However, they are enabled by the fact that with modern ad tracking you can target the same people there as on a premium site. This causes quality to decrease. Just to give an example, this[1] is how this dynamic affected Recode:

> "I asked him if that meant he’d be placing ads on our fledgling site. He said yes, he’d do that for a little while. And then, after the cookies ... helped him to track our desirable audience [he'd] begin removing the ads and placing them on cheaper sites"

[1]: https://www.theverge.com/2017/1/18/14304276/walt-mossberg-on...

davidfischer··on Handling 100 Requests per Second with Python and Django
We may have to eventually do that. That also might make more sense for handling requests across continents. Every 5 minutes still seems like not frequent enough but presumably the same approach could be used to write every 5-10s.

What we have considered is not doing any synchronous writes and just queuing it up and handling it async. There's definitely some questions about whether our current approach will scale 10x but it should be fine for the next 2-3x.

Just to shed a bit more light, we break things up into figuring out which ad to show and then handling when that ad is actually seen. The second part is mostly async already but the first part is the harder part. You want to choose the best ad for the content, the geographic targeting needs to match, and the ad campaign has to have budget left on the hour/day/total. If we only checked the budget every 5 minutes, that would be a problem. At some point, multiple servers need to know that there's still budget on a campaign and this requires some amount of synchronization.

davidfischer··on Handling 100 Requests per Second with Python and Django
Maybe I didn't run waitress through as full a test as I should have. In my (admittedly shallow) tests, it didn't offer any significant performance benefits over Gunicorn and I stuck with Gunicorn for the simple reason that the rest of our infra used it. Would you be willing to share a bit more about your configuration that made waitress significantly better for you?
davidfischer··on Handling 100 Requests per Second with Python and Django
I'm not an expert on async views at all but my hunch is it wouldn't do much. I do think async has some pretty interesting applications though on Read the Docs itself in our small proxy that serves docs which are just static files in s3. Currently it's a stripped down Django setup running in separate process from the main RTD app and it mostly sets headers picked up by nginx/sendfile.
davidfischer··on Handling 100 Requests per Second with Python and Django
I did check out ClickHouse but I haven't gotten a chance to load more real data to give it the full test. It's definitely on the todo.
davidfischer··on Handling 100 Requests per Second with Python and Django
Postgres is mostly bored (CPU sub 20% utilized) with 60 write requests per second which is normal peak traffic. Even traffic spikes to 100-120 don't usually change that much.

As our ad network grows though, everything has to grow pretty linearly (writes, reads, requests, etc.) and I hope our approach will continue to work well even with double or triple these numbers.

davidfischer··on Handling 100 Requests per Second with Python and Django
We've run some tests with PyPy on Read the Docs itself but not for ads. For Sphinx documentation builds, builds took around ~50% of the time. It was especially pronounced on builds with large numbers of doc files (hundreds) and therefore complicated side navigation.
davidfischer··on Ask HN: Ethical Advertising Companies
I work at Read the Docs and we're behind EthicalAds (ethicalads.io). Read the Docs has been running ads on our own sites for years and that's how we're funded. We do our own ad sales, we don't run advertiser supplied scripts, and we host the ad resources like images ourselves. You can read a bit about it here[1] or here[2].

EthicalAds is our ad network for sites not hosted by us and it's about a month old. It uses the same ad serving setup as Read the Docs and it follows the same principles.

Based on my affiliation, I'm going to be biased. However, if you have any questions, you can reach me at my first name @readthedocs.org.

[1] https://docs.readthedocs.io/en/stable/advertising/ [2] https://blog.readthedocs.com/archive/tag/advertising/

davidfischer··on AWS NLB now supports multiple TLS certificates via SNI
The 25 cert limit is pretty annoying. It also applies to Application Load Balancers (ALBs) on AWS as well.

At my company, we use Cloudflare's SSL for SaaS offering. It isn't exactly a load balancer per se, but it did allow us to have many users each with their own domains. We decided for our use case that it was better than rolling our own certificate management system for a couple thousand certs and figuring out how to hot load them onto web servers. https://developers.cloudflare.com/ssl/ssl-for-saas/

davidfischer··on Why the US Has No High-Speed Rail
I wrote that "there are some areas where high speed rail makes sense in the US" and I think that answers your question of "why the coasts don't sport high-speed rail". Specifically I cited the lower 48 density, not the US as a whole (including Alaska would be irrelevant), and I broke out California and the Northeast to show that they are more favorable to rail. I also mentioned that the Amtrak lines in Southern California and the Northeast make sense. Parts of the coasts do support high speed rail and as the video mentions, the coastal areas (California, Northeast, Florida, and even densely populated areas of Texas if you count the Gulf Coast) are building it albeit slower than many people would like.
davidfischer··on Why the US Has No High-Speed Rail
> If you count about 140 mph as high speed train, Finland (pop. density 17/km^2) and Sweden (22/km^2) seem to be doing just fine.

Citing the population density of Finland and Sweden as a whole is pretty disingenuous. Those two countries have densely populated urban centers in the south with high speed rail yet vast sparsely populated areas in the north with nothing. High speed rail gets built where it makes sense and no amount of hand waving is going to make building a high speed rail line from LA to Chicago or New York viable. With that said, I will agree that politics and culture are certainly factors.

davidfischer··on Why the US Has No High-Speed Rail
It's a little surprising that population density isn't mentioned as a major factor in this video. While it doesn't explain the overall cost or cost overruns, it is a major factor in whether rail makes economic sense.

Here's some key statistics:

* US lower 48 states: 40 people per square km (km2)

* California: 92 people/km2

* DC-Boston (this is hand-wavey a bit): 200+ people/km2

* France: 270 people/km2

* Japan: 330 people/km2

* China: 130 people/km2

* Eastern China: 250++ people/km2

Most of the Chinese high speed rail is in the Eastern part of the country where roughly 400M people live. That's more people than the entire US in a space about double the size of California. Likewise Japan's high-speed rail links Osaka and Tokyo which are among the most densely populated areas in the world. A high speed rail trip between SF and LA could be free and you still couldn't have a remotely full train leave every 10 minutes like you do in Japan.

The Amtrak lines that make the most sense are the shorter length ones that connect cities. The video sort of mentions this and that is in line with the Brookings Institution[1] findings. There's no coincidence that most of these lines operate in some of the most densely populated areas in the US[2] (So. California, Northeast). The goal of transit isn't strictly to turn a profit and highways don't turn a profit either. It has other goals like replacing car trips and alleviating congestion but all of these require people to actually ride. To get riders, trains need to operate where people are.

Long story short, there are some areas where high speed rail makes sense in the US but the US simply doesn't need as much high speed rail as China does.

[1] https://www.brookings.edu/wp-content/uploads/2016/06/passeng...

[2] https://en.wikipedia.org/wiki/List_of_states_and_territories...

davidfischer··on Python Developer Survey 2018 Results
Read the Docs hosts a lot of Python module docs and we ran community ads (meaning free) promoting this survey. There was no mention in our ads of JetBrains - only the PSF.

The ads wouldn't have appeared on the Requests docs but it could have on many others.

davidfischer··on When hiring senior engineers, you’re not buying, you’re selling
A couple things to add:

- The #2 source of candidates was almost universally meetups in my customer development interviews. Meeting people in person works! You don't have to be a referral although it does help.

- It's pretty easy at a small company to figure out who the hiring manager is and contact them via LinkedIn/Meetup/etc. At a big company, this can be hard or impossible but you are much more likely to know somebody at a big company who can put you in touch. Absolute worst case, reach out to the big company recruiter on LinkedIn asking for some details on a specific role and they might put you in touch directly or they'll get your questions answered some way. The recruiter 100% will talk to you. They are paid good money to find qualified people like you.

- While I said earlier it isn't a numbers game and you should only apply to roles/companies where you're a fit, to some degree it is a numbers game. Sometimes jobs are posted to a company website and the role isn't really available. Maybe they already have a candidate in mind. If you apply to ~3-4 jobs and don't hear anything, that isn't super uncommon. However, most people feel obligated to actually give you a direct answer if you contacted them directly.

davidfischer··on When hiring senior engineers, you’re not buying, you’re selling
Sorry to bum you out. That's not my goal.

As one becomes more senior the hiring manager expects a candidate to do their due diligence and ask some questions beforehand rather than hopping right into the application process. For a referral, that has already happened somewhat as the candidate already presumably talked to the the referrer about working at that company. You don't have to be a referral but a hiring manager expects you to do this due diligence. You don't necessarily need to go through a friend and you don't have to work with your friends. However, having an acquaintance at the company even in a different team is a massive help.

Think of it from the company's perspective for a second. They don't want somebody who just "wants a job". Money is always a factor but if somebody takes a job solely for money then they'll jump ship the moment somebody offers more. Companies want somebody interested in their space who wants to work there. You mentioned that you purposefully pursued some roles that looked like they'd be good fit for you. That's smart! Make sure the hiring manager knows it. By just clicking apply on the website, that may not be communicated clearly.

Here's two concrete pieces of advice:

- Before going in the front door and applying for roles, reach out to the hiring manager and ask them some more details about the role and what working for the company/team is like. If you know somebody at the company who can put you in touch (even if it isn't the same team), that's better, but even if you don't that's ok. They will talk to you and if they don't you probably don't want to work there anyway. LinkedIn is useful for this.

- If you're applying for a local job, try to meet the hiring manager or somebody else at the company in person if possible. Don't stalk them but to give you an example, I know that a CEO for a local security startup runs the local OWASP meetup and I told a person interested in working in security to go to the meetup and talk to him if she was serious about working there. I've been a hiring manager before and I looked way more closely at any candidate I talked to in person. Don't sound desperate but show them that you're interested and serious enough to go out of your way to talk to them. I had a couple candidates who I didn't know reach out to me via Meetup.com for roles posted to the company website and I had no problem meeting them for coffee.

On a related tangent, I help organize a local meetup and I get asked the question "how do I land my first dev job" a lot. I've helped half a dozen people land their 1st or 2nd tech job. Assuming they already have tech talent, I essentially offer them the above advice. It works at that level but it works for senior roles too! As a hiring manager, you want to talk to qualified people about your open role.

Good luck!

Edit: formatting only

davidfischer··on When hiring senior engineers, you’re not buying, you’re selling
While the definition of senior engineer is a bit squishy, I agree completely that as some one gets more senior they increasingly choose the company rather than the other way around.

I'm soft-launching a product to help companies recruit and I've done about a dozen customer development interviews with hiring managers (5 person companies through FAANGs) in the last two weeks. One thing I heard constantly was that referrals are their best source of candidates and that for roles with 7+ years of relevant experience they actually don't believe good candidates will come in by just randomly applying on their website/LinkedIn/etc. I had an engineering manager tell me "a senior engineer actively searching for a job is a negative signal of candidate quality" as opposed to going through somebody in their network or connecting with somebody at the company first. I had another manager tell me for the senior roles he's hiring for he only considers candidates explicitly sourced by his internal recruiter who searches LinkedIn for the specific skillset or people who are referrals for the specific job. This isn't some CTO level position; this is just a role for a ~10 years of experience individual contributor. Applications from the website go straight to the trash can. The implication was people with those skills aren't looking for jobs.

davidfischer··on Why Doesn't Google Comply with Do Not Track?
- Medium (https://medium.com/.well-known/dnt-policy.txt)

- EFF (https://www.eff.org/.well-known/dnt-policy.txt)

- Read the Docs (https://readthedocs.org/.well-known/dnt-policy.txt)

Sites that implement DNT should respond to either /.well-known/dnt-policy.txt or /.well-known/dnt/ or both.

Responding to /.well-known/dnt/ means a site has implemented the W3's Tracking Preference Expression (https://www.w3.org/TR/tracking-dnt/). This doesn't necessarily mean much as there's no agreed standard of what complying with DNT means. However, it typically implies that the site does something different for users with DNT enabled vs. disabled.

Responding to /.well-known/dnt-policy.txt is typically stricter and means a site adhere's to the EFF's guidelines (https://github.com/EFForg/dnt-guide) for DNT. This has rules for how long data is retained, which data can be retained, specifications around anonymizing data, and security precautions.

Disclaimer: I worked on Read the Docs' DNT implementation.

davidfischer··on Why Doesn't Google Comply with Do Not Track?
> I'm sorry, but this is crap.

I agree with your sentiment that on by default should be the standard. However, it isn't. The standard is specific and it says it should be off by default.

The standard was designed to be completely toothless and the advertising industry still tried to water it down. There's no enforcement mechanism and nothing that says what it means to comply.

If you want something with teeth, try the EFF's pseudo-standard for DNT: https://github.com/EFForg/dnt-guide

davidfischer··on Why Doesn't Google Comply with Do Not Track?
This is because on by default is against the standard. The only real standard around DNT is the Tracking Preference Expression[1] which says how to declare you comply with DNT without saying what compliance means. It isn't exactly a preference to not be tracked if the decision was made by your browser for you. Roy Fielding committed a change to Apache to ignore DNT on IE[2] as a result.

[1] https://www.w3.org/TR/tracking-dnt/ [2] https://arstechnica.com/information-technology/2012/09/apach...

davidfischer··on Ads just work, no matter what you think
While I can't speak for Hacker News as a whole and I'm going to talk about clicks and not purchases strictly, I am willing to talk about developers more generally who frequently say things like "advertising doesn't work on me". Collectively, they are mistaken and it does work on them.

I work on advertising at Read the Docs and when we were first added to the main ad blocker lists our revenue went down. We were 100% pay per click at that time and our revenue went down by the same percentage as the number of ad impressions went down. Looking at it statistically, developers who run ad blockers still click on ads at the same rate as those that don't as long as they actually see the ad.

davidfischer··on California teacher pension debt overwhelms school budgets
This was huge news in San Diego in 2012 when it was revealed. The worst of these types of deals are now outlawed as of 2013.

http://articles.latimes.com/2013/oct/02/local/la-me-ln-schoo...

davidfischer··on Cookie policy notifications have ruined user experience on the web
For the purpose of the GDPR, an IP is PII. However, one can get rough location using an anonymized version of an IP address (say one that has zeroed out the last octet or two).

From Recital 26 (https://gdpr-info.eu/recitals/no-26/):

> The principles of data protection should therefore not apply to anonymous information, namely information which does not relate to an identified or identifiable natural person or to personal data rendered anonymous in such a manner that the data subject is not or no longer identifiable.

← PreviousPage 2 of 3Next →