357 karma · joined March 26, 2018
Things have changed a lot since then: OS kernels became faster by eliminating a lot of unnecessary (?) cross-process overhead; browser makers made a number of potentially problematic decisions ("let's allow Javascript to create CPU threads — what could possibly go wrong?"); Linux kernel developers made few potentially problematic decisions ("let's allow unprivileged processes to invoke arbitrary BPF bytecode — that worked for Java, so what could possibly go wrong?")
A lot of small security lapses added up until it became viable to use CPU flaws to actually target ordinary users. To add insult to injury, certain corporations started spreading myth, that well-known insecure practices — such as knowingly running local software from questionable authors — are "safe enough" for general population. Topic web page even talks about running untrusted Android software, as if Android had some kind of impenetrable security boundary around untrusted apps.
Exactly: the integration requires cooperation from both websites, and there is zero reasons, why child website would want to cooperate. Unlike, say, using Google Analytics, letting user leave you via <portal> tag is against website's best interests.
Google is completely open about their intentions. Citation from https://github.com/WICG/portals/blob/master/key-scenarios.md:
> After being activated, a portal opened by the news aggregator will now receive the input events. This means that it will have to co-operate with the news aggregator in order to maintain the desired user experience.
Why exactly does it _have_ to cooperate? And _how_ can it cooperate? Will it be enough to include a script from parent domain? Will analytics be part of "desired user experience"? Will parent domain ban a site from it's index for trying to "break out"?
It is essentially impossible to build relationships of mutual trust to the point when one website can freely embed another... unless one party has overwhelming power advantage. Who is the target auditory of <portal>? What "news aggregator" has enough leverage to make pages switch back to it as told to? If such "aggregator" existed, I would not trust it to make web browsers or dictate, how they should work.
That depends on one's definition of "ok". I don't think that modern web is ok.
The difference is made crystal-clear in the draft [1]:
> Every browsing context has a portal state, which may be "none" (the default), "portal" or "orphaned"... "orphaned": top-level browsing contexts which have run activate but have not (yet) been adopted
In other words, the original google.com document will be kept active in background (in "orphaned" state) and the child document can continue to interact with it by using Javascript after "adopting" it (at which point roles reverse and google.com becomes child document itself).
In theory the specification allows child document to ignore parent and let browser close it, completing transition. It also allows to perform graceful switching between child and parent, keeping each in control as long as it remains top document.
In practice, child portals will run Javascript, written by Google, and will be subject to Google's complete discretion. The distinction between first-party and third-party scripts will be erased, effectively letting Google run analytics on third-party domain, and send results back from it's own domain via Javascript proxy.
It is a branch of traditional medicine, approximately 100000 years old and counting.
"Western medicine" ignores phage therapy, because growing a bunch of mud in a bowl does not bring anyone 10000% cuts. Cultivating phages still requires labor and equipment, and can be quite profitable, just not 10000% profitable, — because pretty much anyone can do it.
> How is that not how garbage collectors work?... I'm assuming we're talking about a standard, generational mark-and-sweep gc
GCs do not scan "all memory", but small fraction of memory. In case of generational GC the scanned fraction of memory is (usually) limited to single generation. Even without generational approach scanning heap itself is frequently avoided in favor of scanning separate data structure with highly compressed representation of object set.
GCs do not generally iterate over memory just because they can. They either reclaim space for new allocations, move things around to reduce fragmentation or fire periodically in response to increased allocation rate. If your program does not make allocations, it may never incur a GC at all.
The grandparent comment makes it sound like garbage collection is a simple effort, conducted solely by distinct GC code ("giant for loop"). This is often not the case: for example, JVM may generate additional memory barriers in any code, that uses references (exact nature and purpose of memory barriers depends on GC being used [1]). Augmenting the code with those barriers allows GC to operate more efficiently and quickly: achieve smaller pauses, scan less memory, collect memory for some threads without disturbing others.
It boggles my mind too. Imagine, what would have happened, if those small Javascript snippets, used mainly to add cute visual effects to pages, could check if some image from different site is already in browser cache by performing cross-site HTTP requests... That would allow completely new dimension of spying on web users!
Fortunately, browser developers are some of the most competent people in the word. They would never give web pages too much power by letting them start CPU threads, use OpenGL, allocate arbitrary amount of memory or read your battery level to set exorbitant taxi tariffs for people in a pinch. Browsers are well-designed and highly secure, because they are being updated with security fixes every day, sometimes even multiple times a day.
The infamous "a:visited" tracking didn't simply track your visits from Google — it tracked all your visits across entire Internet. Browser vendors are bunch of lazy hacks, who can't even implement per-site link history (just like they failed to implement per-site cookies). All "a:visited" states are source from single SQLite database, that stores your full web history. THAT is the "CSS Tracking", because it can tell a page about visits from completely different domains. Instead of separating your web history per-domain those <censored> have crippled :visited selector in several undocumented ways.
You know, that HTTP allows websites to "track" you each time you visit them, right? The horror!
This specific page uses :active, not :hover, so it is really no different from a web form, that performs web request each time you press a submit button. It just does not reload a page.
Cloudflare can check if nameserver and the actual server are run by different parties, and if so omit subnet information from EDNS response. It is not hard to implement — Google and OpenDNS used to require manual whitelisting to receive EDNS subnet responses (not sure if they still do).
Cloudflare's CDN leaks user's full online identity to Google via reCaptcha, especially when you use Tor. Maybe they should ask Google to be satisfied with client's subnet too?
Most recursive DNS severs on Internet can be categorized in two groups: local DNS servers, offered by Internet providers to their users, and enormous "generic" DNS like Google's 8.8.8.8. When someone makes a DNS request to those servers, they will in turn forward it to DNS servers of web page you are requesting. Content Delivery Networks use DNS to determine, which server should serve your request: if your DNS request arrived from Africa, CDN's DNS server will return IP in Africa. Of course, _users_ don't send DNS requests to CDN's server — recursive DNS servers do. In the past almost everyone used DNS, offered by their Internet provider, — CDN's had to use GeoIP or even static lists of providers to determine origin of that request. When world-wide DNS servers like Google's 8.8.8.8 started to gain popularity, that approach was broken, so EDNS was developed.
Cloudflare is a CDN. They are selling their CDN services for money. At the same time they are encouraging end users to use free DNS server, that does not support EDNS on purpose (they admit so on their website). In effect they are creating a situation, when competing CDNs are at disadvantage and can't determine, what country user comes from. Cloudflare itself does not suffer from that disadvantage, because they control both 1.1.1.1 and DNS, used by their clients' websites.
Why would you? What is the point of "escaping"? It is like asking, "how would Go world escape the grip of Google?" — no one cares to, because nobody asks Oracle for permission to make software for JVM.
Java EE has been in maintenance mode for a long time, but Spring and Dropwizard are alive and well despite that. Incidentally, Nginx Unit has recently implemented Servlet spec — a cornerstone of JEE! — but they are stuck at it's initial version, without asynchronous Servlet support. Maybe Oracle should give Nginx a couple of decades to catch up before drafting next JEE version...
I never understood the purpose of Valhalla. They should just introduce a 128-bit scalar type and call it a day. Nobody actually cares about automagically storing value types in HashMap — all interested parties (HFT, computing) have already written dedicated collection libraries for working with primitives, and those libraries have little in common with object-based ones.
> LLVM and webassembly are eating JVM's cake
pffft, hahahaha, no they don't.
JVM and LLVM don't even share same market niche — one is a feature-complete runtime, another is a "make-your-own-language" build kit. Webassembly? What webassembly?
There is a possibility that Mozilla implemented their backwards code-signing model on purpose — for example, it allows them to oust unwanted extensions without explicitly recalling their certificates. But personally I think that they just didn't give the matter enough thought.
"We accidentally uploaded all your HTTP requests to our servers, but we will definitely fix that in next addon version!~"
I don't believe, that everyone holds that opinion. Even if they did, the world does not revolve around Red Hat's team, — there are still kernel mail lists and other venues for discussion. But if proposed improvements aren't well thought-out, would anyone there back them up?
In my opinion, async-signal safety in itself is much bigger problem than robust registration of signals. The later is mostly solved by chaining signal handlers, while former is mostly unsolved (and keeps getting worse). Proliferation of new libraries and async-signal unsafe conventions. People keep using printf() in signal handlers. Occurrences of fork() in multi-threaded apps. Still no async-signal safe malloc() (some Googlers tried, but the idea didn't get much traction). And then you come and propose new interface for registering signals handlers, and say that "It’s okay for two functions can be async-signal-unsafe". If your proposed API is async-signal unsafe, how would it deal with signals arriving during dispatch of signal handler list?
Thanks for interesting read. That said, I can see why glibc people didn't appreciate the proposal.
The first part of article doesn't mention async-signal safety at all. Second part papers over async-signal safety, as if it were a non-issue. There are some dangerous-sounding paragraphs too:
> It’s occasionally useful to longjmp out of a signal handler. It’s reasonable to want to return non-locally from a shared signal handler too --- that is, to resume program execution after SIGNAL_CONTINUE_EXECUTION in a different state from the state the program had when we entered the shared signal handler. Since the signal system probably wants to maintain some kind of state to track its progress through its shared signal handler list, a plain longjmp out of a shared signal handler will likely leave the system in an unspecified state.
Are you sure, that we should worry about "signal system maintaining some kind of state"? Not about rest of application being in completely unspecified state?!!
The article proposes a primitive system for setting signal priorities, but stops at a half-baked solution. There is a mention of banning longjmp, but individual handlers still can bail via SIGNAL_CONTINUE_EXECUTION. What if I want my handler to always run regardless of registration order?
The proposal does not offer a way to retrieve a list of already installed handlers, which makes that part of it even worse than existing Posix signal API.
The proposed API does not address challenges of using signals in multi-threading programs.
The article mentions, that signal handlers can't be reliably unloaded, but proposed API does not address it.
Overall the proposed interface brings little to the table, does not work well alongside with existing sigaction() API and creates false illusion, that signal handlers are safe and ok to use. I imagine, that if it had more technical "meat" — more like robust mutexes or FD_CLOEXEC — it would have seen a lot more constructive discussion and less hostility from glibc maintainers.
Alternatively, there is no conspiracy, and every major government simply ignores dangers of aviation jet fumes. Sort of how everyone was sure, that invisible radiation is near-harmless and that radium dials are safe to use, until suddenly they weren't.
As soon as Mozilla started relying on Google's money, they were doomed. Now they are staffed by lots of well-paid US developers, and have to continue taking Google's deals to pay the wages.
If you want to hinder determined (but inept) adversary, impose reverse time limit: make your captcha a bit complex and deny answers, that arrive too fast. Legit users will spend a bit of time to solve captcha. Machine-learning-driven bots will blaze it. In addition to measuring speed of filling captchas you can measure amount of user time spent on other actions on your site — in process making your bot detector increasingly similar to Google's reCAPTCHA.
In general look for behaviors, distinguishing legitimate users from malicious. Hint: having Google account might or might not indicate a legitimate user, but it is probably more efficient to ask users for it directly than in roundabout way by using reCAPTCHA.
If they aren't legally obliged to purchase "broadcasting right" from "members", they can always stop doing so.
If they ARE legally obliged (or forced by corporate mob, depending on your preferences) to engage in the rent-seeking, their budget founding should be permanently set to zero.
If my (non-Apple) laptop is anything to go by, the keyboards with that crappy low-rise design have average life expectancy of one year. The 2018 modification will not get clogged with dust until the end of 2019, at which point people will be complaining about it too.
Each time your app opens and reads a file, the application on other side of ContentProvider can monitor user activity down to number of bytes read. Each time you open a directory, that action can be noted. If anything, that sounds like a privacy nightmare, and will undoubtedly be exploited.
Google's external storage ContentProvider is terrible: slow, buggy and lacks a number of basic features. It does not work as replacement of existing file managers, since it does not properly support searching and it's performance is abysmal, compared to using filesystem directly.
AFAIK, there is still no proper support for selecting a file by extension: https://stackoverflow.com/questions/45573171 — if your file extension isn't in Android MIME database, this is it for you. The mime-filters for file extensions are completely broken: https://stackoverflow.com/a/31028507/1643723. It is possible to work around that in file manager application (by using queryIntentActivities with different, modified Uri), but built-in Storage Access Framework file browser does no such thing.
MediaStorage used to have tons of crippling bugs. Wrongly reported file sizes, delayed file addition, files, filled with zeroes... Some were fixed in preparation for Android Q release, but I doubt, that all of them did. MediaScanner is abhorrent by design.
The options for flexible content generation are still bad. Prior to Android P you had to use pipes, which resulted in non-seekable files. ProxyFileDescriptorCallback remedied some issues of that approach, but it does not support FUSE interruption events — https://issuetracker.google.com/issues/38444582: if you open a buggy FUSE descriptor, your app gets STUCK (and closing the opened descriptor from another thread won't not unblock you!) In effect, using files from other applications can now hang your app (it could before too, but only during ContentProvider#open() stage, now read() is broken too). The bug is trivial to fix, but there is still no fix 3 Android versions later.
Google Drive is a great example, how NOT to mesh OS development and app developer interests. Google Drive does not properly implement Google's own DocumentProvider API — in effect, forcing third-party app developers to use it's proprietary web API. That's great for Google (vendor lock-in and all), but not exactly great for third-party apps, because the DocumentProvider API is effectively not cared about (see ProxyFileDescriptorCallback fiasco above). If one extrapolates this behavior to future Android development... imagine, that INTERNET permission is abolished, and you have to use Google's proprietary API for every single thing. Want to download a file? Use Google Downloads (tm) API in Google Services (the system DownloadsProvider will of course be broken as usual). Want to send analytics? Google Services! Crash reporting? Firebase. Communications? Cloud messaging. p2p file exchange? Too bad for you, — that kind of advanced functionality will likely require Chrome OS!
The extrapolation above may sound overly pessimistic until you realize, that Android developers already made great many steps in that direction. Half of OS functionality is locked behind "system"-level APIs, that aren't accessible to ordinary apps, ever. But Google Services can of course continue to use them! As Mark said in his article, “the writing was on the wall”, although a lot of people (himself included) don't want to fully accept it.