1,745 karma · joined June 11, 2017
I said their reporting on EUROPEAN politics ("politics here") copies points from the far left (e.g. all mass migration is unquestionably good, parties against it must be right wing populists if not racists, etc).
I think their reporting on campus politics, identity politics is also far left, but other than that their stance on Iraq war etc is more Hillary-left than traditional 'left'. It's pointless semantics though, outrage mode is already engaged in this thread and it will probably soon turn into a dumpster fire.
Someone cannot just be wrong or inaccurate, they must be the enemy ('right rant'), and culture wars demand I first clarify I am on 'the right side of the issues' before saying anything. The more objective people think they are, the blinder towards their own bias. Of course I am biased too, but what people engage with in my post is the 'far left' comment on European politics instead of the actual point.
My point was that the NYT engages in culture war because it sells. I can agree with many issues on the NYT but still observe that and be annoyed by it, but that does not matter in tribalistic discourse.
https://www.vox.com/2016/4/21/11451378/smug-american-liberal...
Nothing could be further from the truth. I think the most recent Sarah Jeong controversy and virtually all reporting on migration, feminism, campus politics etc shows this. Mind you this is from a European perspective where I see almost all reporting about politics here as copying off talking points from the far left.
This is the true genius of their marketing though: They are actually as polarized as any other source in the culture war, but market themselves to an audience that likes to think of themselves as rational, objective, sensible.
So surely there must be differentiation on whether any given job or emergency can affect your standard of living significantly.
I would categorically reject any collaborations with FB as an academic in ML.
They complain that Maps at new prices would be more than the cost of their infrastructure, when actually their entire startup revolves around this data, and 5k to be able to use an amazing piece of high tech software that is ahead of the competition is..peanuts.
No matter what other business interests and strategies Google follows, there is no right to using such valuable tech for less than the monthly salary of an engineer. I find this incredible entitled.
No mention of all the ongoing work in learning from demonstrations, or more generally incorporating any off-policy knowledge. Vague speculations about the philosophy of model free learning. Not really worth the read (as someone working in RL).
Is declining to accept a liquidation preference at seed level a red flag for any serious investor? What about subsequent rounds?
Thinking you can interact with customers of your employer like this in public even on a private account is naive.
Trying to turn it into an oppressor/sexism shitshow is being a toxic employee and a liability.
Hard to see a scandal..
This results in for instance there being proportionally many more Romanians at Oxbridge Cs/Math/Physics than you'd expect by population.
The interview process is relatively standardized, and if my results were to differ starkly from what more experienced interviewers do they would be disregarded, and the director of studies for that college would simply not invite me back to help.
I would add that asking PhD students to do this is not the worst thing because via supervisions and other teaching efforts, we have a good picture of what undergrads here need to be able to do. An interview is a like a short supervision.
The point is that if you go to a target school doing Math olympiads throughout your school life, the admissions exam and interview is a walk in the park. The applicants who didn't have any of this preparation can still do well but will fare relatively worse against that group, and I think this is very obvious in the ultimate intake.
They don't try that hard.
Being vaguely involved with the Oxbridge undergrad admissions process in CS/Math i can tell you there is very little trying. A fuss is being made about coming from a disadvantaged background but in practice sadly the people running it only care about one thing: how well you can grind out an answer to a math olympiad style question in 15 minutes. Yes, extra-curriculars and well-roundedness don't matter which I think is a good thing because I believe in focusing on being great at one thing.
What it comes down to nonetheless is preparation and school support, e.g. via training for math competitions. Saying the interviews are about 'evaluating the thinking process' of the applicant is a fantasy when most applicants come from schools where they have been trained to do them for years. Oxbridge are not forthcoming about this but ultimately they take people who are already well groomed Math olympiad winners, not raw potential.
It's probably still better than opaquely selecting for race and like-ability and if this means many math undergrads are Asian, why should that be a problem? It's still unfair to disadvantaged children and this sucks, but at least the criteria are clear.
Ps: on your question how they recruit rowers: They let them study land economy, that's the joke at least.
I don't know what the poster meant by suggesting 'Mathematician-MD', but it reads weirdly to me for that reason. It's highlighting an attribute of a person that is entirely unrelated to his career or this article. Why if not to denigrate him? The title should be changed to neutrally reflect his position.
https://www.lesswrong.com/posts/yCWPkLi8wJvewPbEp/the-noncen...
He is merely visiting USC so it strikes me as weird that they would claim this PR so quickly.
Also Mathematician-MD somehow makes it sound like the MD means he is a lesser mathematician or not a full mathematician. Fokas is a well respected Professor at one of the top applied Maths departments in the world. A better and less biased title would be 'Math Professor' or 'Cambridge math professor' claims..
Determining the scale needed, fiddling with the state/action/reward model, massively parallel hyper-parameter tuning.
I may be overestimating but I would reckon with hyper-parameter tuning and all that was easily in the 7-8 figure range for retail cost.
This is slightly frustrating in an academic environment when people tout results for just a few days of training (even with much smaller resources, say 16 gpus and 512 CPUs) when the cost of getting there is just not practical, especially for timing reasons. E.g. if an experiment runs 5 days, it doesn't matter that it doesnt use large scale resources, because realistically you need 100s of runs to evaluate a new technique and get it to the point of publishing the result, so you can only do that on a reasonable time scale if you actually have at least 10x the resources needed to run it.
Sorry, slightly off topic, but it's becoming a more and more salient point from the point of academic RL users.
The point being that that the bells and whistles of PPO and other relatively complaticated algorithms (e.g. Q-PROP), namely the specific clipped objective, subsampling, and a (in my experience) very difficult to tune baseline using the same objective, do not significantly improve over gradient descent.
And I think Ben Recht's arguments [0] expands on that a bit in terms of what we are actually doing with policy gradient (not using a likelihood ratio model like in PPO) but still conceptually similar enough for the argument to hold.
So I think it comes down to two questions: How much do 'modern' policy gradient models improve on REINFORCE, and how much better is REINFORCE really than random search? The answer thus far seemed to be: not that much better, and I am trying to get a sense of if this was a wrong intuition.
Well I guess my question regarding the expensiveness comes down to wondering about the sample efficiency, i.e. are there not many games that share large similar state trajectories that can be re-used? Are you using any off-policy corrections, e.g. IMPALA style?
Or is that just a source off noise that is too difficult to deal with and/or the state space is so large and diverse that that many samples are really needed? Maybe my intuition is just way off, it just feels like a very very large sample size.
Reminds me slightly of the first version of the non-hierarchical TensorFlow device placement work which needed a fair bit of samples, and a large sample efficiency improvement in the subsequent hierarchical placer. So I recognise there is large value in knowing the limits of a non-hierarchical model now and subsequent models should rapidly improve sample efficiency by doing similar task decomposition?