HNHacker News
TopNewBestAskShowJobs

aglionby

244 karma · joined October 30, 2017

PhD student in NLP, interested in interpretable reasoning strategies using knowledge graphs.

hello@ | www.guyaglionby.com

https://pinboard.in/u:guyaglionby

submissionscomments
aglionby··on Tell HN: Airbnb just stole me 5 minutes of my time adding dices
I don't know - with some edge detection I think you could distinguish between the faces of each die, and some heuristics would probably get you the upwards-facing one. Another thing that may be useful is that in each image the only face that is fully visible for all dice is the upward one (who knows if that is just for this example though).
aglionby··on BibTeX Tidy
This is great! Especially nice to be able to remove entire fields.

Relatedly, here are a couple of tools to ensure that references are complete (e.g. updating arXiv papers to their published versions, mostly for computer science papers):

- https://github.com/yuchenlin/rebiber (CLI, web interface)

- https://www.cl.cam.ac.uk/~ga384/bibfix.html (only *ACL papers, web interface with diff, disclaimer: mine)

aglionby··on Emoji Under the Hood
Great post, entertainingly written.

Back in 2015, Instagram did a blog post on similar challenges they came across implementing emoji hashtags [1]. Spoiler alert: they programmatically constructed a huge regex to detect them.

[1] https://instagram-engineering.com/emojineering-part-ii-imple...

aglionby··on Amazon buys 11 Boeing 767s to expand its cargo fleet
I'm guessing they won't come unless you chase it. Nothing happened to my refund for May flights until I called, after which the money came within a week.
aglionby··on Chess’s Cheating Crisis
It isn't fair (in the statistical sense) to make a comparison on the basis of rank between an over- and under-represented group. Interesting read here https://en.chessbase.com/post/what-gender-gap-in-chess
aglionby··on The growing short case on Facebook and Google
2002 was the year of the Nokia 7650, their first phone with a camera - 0.3MP which it displayed on its 176x208 screen. It cost €740 in 2020 money (€550 back then).

In 2020 you can buy a second generation iPhone SE for €489, with a 1334x750 screen (at >6x pixel density) and 12MP camera.

These are not roughly equivalent pieces of hardware, and the software they run differs immensely. MMS is not Instagram.

aglionby··on Pinboard Turns Eleven
> system where I actually refer back to things because I seem to google for the same things over and over again

Yep, this often frustratingly turns nothing up for me. Hence, Pinboard :)

aglionby··on Pinboard Turns Eleven
I have two use cases. The first is to keep track of interesting articles I find and plausibly want to refer back to in the future. A 3rd party browser extension and mobile app make saving very easy, and then I tag each item with a high-level category. This is also pretty painless, and brings a lot of value (otherwise you just have an unsorted collection of links - not helpful). An example is my 'long reads' tag https://pinboard.in/u:guyaglionby/t:long-read/. The 'unread' feature is also useful here - I've got >10 long reads banked for when I'm looking for things to do.

The second is as a kind of mechanism to give myself permission to close a bunch of tabs every time they accumulate. Each is _obviously_ open for a good reason and I may want to read it at some point, so sticking it on pinboard is a nice way of shoving them elsewhere. I don't save everything - curation is important (in the same way as with tagging). Lots of what remains are things that may be useful for me in the future but are not immediately, like design guides https://pinboard.in/u:guyaglionby/t:design/. Some of these things I leave as 'unread'; others that feel more like reference material I mark as 'read' immediately so as not to have them in my to-read queue.

aglionby··on The Bitter Lesson (2019)
My background is in NLP - I suspect we'll see similar in language processing models as we've seen in vision models. Consider this[1] article ("NLP's ImageNet moment has arrived"), comparing AlexNet in 2012 to the first GPT model 6 years later: we're just a few years behind.

True, GPT-2 and -3, RoBERTa, T5 etc. are all increasingly data- and compute-hungry. That's the 'tick' your second article mentions.

We simultaneously have people doing research in the 'tock' - reducing the compute needed. ICLR 2020 was full of alternative training schema that required less compute for similar performance (e.g. ELECTRA[2]). Model distillation is another interesting idea that reduces the amount of inference-time compute needed.

[1] https://thegradient.pub/nlp-imagenet/

[2] https://openreview.net/pdf?id=r1xMH1BtvB

aglionby··on The Sci-Hub Effect: Sci-Hub downloads lead to more article citations
There's some similar work out that analyses the impact on conference paper acceptance of having deanonymised arXiv versions of papers available before review. They look at ICLR papers for the last 2 years.

I've not read it in a lot of detail but it looks like there's a positive correlation between releasing papers and having them accepted. Not sure how they've controlled for confounders (you only release papers you're confident in the quality of on arXiv?) https://arxiv.org/pdf/2007.00177.pdf

aglionby··on Elevator.js – A “back to top” button that behaves like a real elevator
Home and end are most useful for me when I'm writing or programming, so both of my hands are on the keyboard anyway. I think I actually find these key combos more useful than a dedicated home or end button, as I'd have to move my hands a lot further for those.
aglionby··on Restaurants rebel against delivery apps as cities crack down on fees
Dominos in the UK has similar incentives but put in place less rigidly - any medium/large pizza will run you £18-£20, but there are loads of 40%-50% discount above £x (x >= 35) or 2-4-1 deals. This makes it more difficult to justify a solo order (unless you get 2 pizzas, which, fair enough) and incentivises ordering with friends or not at all.
aglionby··on Why Common Sense Is Not So Common in NLP
This problem of inferring something's truthfulness based on the number of times it's written down is nicely explored in the paper "Reporting Bias and Knowledge Acquisition".

https://openreview.net/pdf?id=AzxEzvpdE3Wcy

aglionby··on Facebook agreed to censor posts after Vietnam slowed traffic – sources
> It's an amazing channel for businesses to reach their customers.

Do you have any insight on how consumers find this? I can imagine it's nice to hear (rarely) from a few places that you care about, but taken too far I think I'd find the mixing of messages from friends and companies pretty annoying.

aglionby··on Shirt Without Stripes
Depends on your audience, but I imagine many find the answer boxes on Google search pretty useful. Getting the population of a city without having to click any links is probably good for your perceived value. For this you need some NLP tech to extract intent from the query and match it to the right entity in their knowledge graph (in addition to something to help you build the graph in the first place).

Google have a blog post from October last year with some more complex examples of where more sophisticated NLP helps https://www.blog.google/products/search/search-language-unde...

aglionby··on Ask HN: What's an unsolved problem in your field?
Relatedly, I'm curious about different ways that CS education can be framed. Speaking to teachers, one of the main reasons why many kids (and, crucially, parents) aren't interested in CS is the misconception that you can't be creative with programming. When kids find out that this is not the case, engagement rises. (There's also the point that CS seems to invariably be folded into programming, but I'm not sure this is the biggest problem.)

I wonder what the impact on uptake would be if the focus was shifted towards CS as a venue for building things and being creative and away from lines of monowidth code and indecipherable errors. More to the point, I wonder how this might be done.

aglionby··on YouTube accidentally permanently terminated my account
Isn't taking this as an absolutist view kind of myopic? Sure, people don't pay anything to watch and each person's eyeballs aren't worth that much. But pretty much the entire value of the service is locked up in the crowd of these people, led by the minority who actually create things. Annoy enough of the creators (or lose the trust of those who see what happens to their peers) and they'll start to leave, taking your advert-viewing crowd with them. The network effect makes this really difficult and slow to begin with, but we're already seeing attempts by people like Wendover Productions with things like Nebula and CuriosityStream. I'm curious to see how this plays out over the next few years (and if it's similar to anything that's happened historically?).
aglionby··on Ask HN: What is your blog and why should I read it?
I guess it depends on what you're writing about: tech might benefit more from a date than philosophy. I'm sure he's written multiple times about this, but patio11 has a thread you might be referring to here https://twitter.com/patio11/status/1234141833661440001 .
aglionby··on Rough.js – Create graphics with a hand-drawn, sketchy, appearance
Good example of where this has been used https://www.jwilber.me/permutationtest/
aglionby··on Open access to ACM Digital Library during coronavirus pandemic
It's true (at least in my experience) that the paywall is transparently dealt with when connecting from a university network, but most sites make it pretty easy to access their content via your institutional login even if you're not.
aglionby··on I had to build a web scraper to buy groceries
Don't have any personal experience, but in the UK I think most supermarkets are taking similar steps, at least when it comes to plexiglass shields and limiting capacity in shops. I was surprised at the extent of the distancing measures in some cases -- looks like Tesco defines a route around (some of? I got this from an ad [1], so, pinch of salt) their shops and has 2m intervals marked throughout.

[1] https://twitter.com/Tesco/status/1243670478524383232

aglionby··on Ask HN: What projects are you working on now?
I once took a look at traffic coming from my (Android) phone and IIRC the pressure was being sent to some Google server at a regular interval. I wonder if they use it for weather forecasting.
aglionby··on Ask HN: What projects are you working on now?
I feel the resemblance to Theme Park World - thanks for reminding me of it!
aglionby··on Ask HN: What projects are you working on now?
This sounds fantastic and is something I've broadly thought about for a little while. A little while back someone else linked a concept of what they had done with the idea -- initially a paragraph was visible, and some of the words had a coloured box around them which could be clicked to expand on that term, all the while maintaining a flowing prose. So, I think a slightly different goal to what you describe (it was technical writing I think?), but similar enough. Unfortunately I didn't save it and have previously (frustratingly) spent at least an hour looking for it to no avail, so I'm curious to see what you come up with!
aglionby··on Snap: Build Your Own Blocks
Though both Scratch and Snap are open source, I can't see any documentation for building on top of either of them. If you're interested in building something in this space, Blockly [1] is essentially the same and has some great docs for working with it (no affiliation).

[1] https://developers.google.com/blockly

aglionby··on What's so hard about PDF text extraction?
I spent some time extracting abstracts from NLP papers (ACL conferences) and it was mostly straightforward. Using pdfquery to extract PDF -> XML gave each character as an element, and they were mostly ordered sensibly and grouped into paragraphs.

However... this didn't work in some cases, mainly with formatted text but sometimes with PDFs that looked like they were compiled in some nonstandard way. As a result I ended up chucking the XML structure entirely and recompiling the text from character-level coordinates. Formatted text was also an issue, with slightly offset y coordinates from regular characters on the same line.

I'm not sure I could take this experience and say that extracting _all text_ would be straightforward. Hopefully for most documents the XML is nicely structured, but I imagine there are many more opportunities for inconsistencies in how the PDF is generated when thinking about diagrams, tables etc. rather than just abstracts.

Considered writing up a blog post about my experiences with the above but imagined that it was far too niche. Code's here [1] if it's of interest.

[1] https://gist.github.com/GuyAglionby/4b55d00803710f2e2e9877fd...

aglionby··on Why the Gov.uk Design System team changed the input type for numbers
The Government Digital Service does all kinds of cool work around accessibility for both users using assistive technology and users who are less technically able. Interesting talk here from 2014 https://www.youtube.com/watch?v=CUkMCQR4TpY
aglionby··on ALBERT: A Lite BERT for Self-Supervised Learning of Language Representations
Paper here: https://arxiv.org/abs/1909.11942

Really impressed at the speed with which Hugging Face ported this to their transformers library -- Google released the model and source code Oct 21 [1] and it was available in the library just 8 days later [2].

[1] https://github.com/google-research/google-research/commit/b5...

[2] https://github.com/huggingface/transformers/commit/c0c208833...

aglionby··on UK Met Office Climate Dashboard
Right, so what is plotted is the deviation from the mean reading in the time range given? Thanks for the links, those are insightful.
aglionby··on UK Met Office Climate Dashboard
Interesting graphs, but it's a little confusing that all but one of them use difference of the metric (i.e. relative values) instead of absolute values -- what and when are the differences relative to?
Page 1 of 4Next →