244 karma · joined October 30, 2017
hello@ | www.guyaglionby.com
https://pinboard.in/u:guyaglionby
Relatedly, here are a couple of tools to ensure that references are complete (e.g. updating arXiv papers to their published versions, mostly for computer science papers):
- https://github.com/yuchenlin/rebiber (CLI, web interface)
- https://www.cl.cam.ac.uk/~ga384/bibfix.html (only *ACL papers, web interface with diff, disclaimer: mine)
Back in 2015, Instagram did a blog post on similar challenges they came across implementing emoji hashtags [1]. Spoiler alert: they programmatically constructed a huge regex to detect them.
[1] https://instagram-engineering.com/emojineering-part-ii-imple...
In 2020 you can buy a second generation iPhone SE for €489, with a 1334x750 screen (at >6x pixel density) and 12MP camera.
These are not roughly equivalent pieces of hardware, and the software they run differs immensely. MMS is not Instagram.
Yep, this often frustratingly turns nothing up for me. Hence, Pinboard :)
The second is as a kind of mechanism to give myself permission to close a bunch of tabs every time they accumulate. Each is _obviously_ open for a good reason and I may want to read it at some point, so sticking it on pinboard is a nice way of shoving them elsewhere. I don't save everything - curation is important (in the same way as with tagging). Lots of what remains are things that may be useful for me in the future but are not immediately, like design guides https://pinboard.in/u:guyaglionby/t:design/. Some of these things I leave as 'unread'; others that feel more like reference material I mark as 'read' immediately so as not to have them in my to-read queue.
True, GPT-2 and -3, RoBERTa, T5 etc. are all increasingly data- and compute-hungry. That's the 'tick' your second article mentions.
We simultaneously have people doing research in the 'tock' - reducing the compute needed. ICLR 2020 was full of alternative training schema that required less compute for similar performance (e.g. ELECTRA[2]). Model distillation is another interesting idea that reduces the amount of inference-time compute needed.
I've not read it in a lot of detail but it looks like there's a positive correlation between releasing papers and having them accepted. Not sure how they've controlled for confounders (you only release papers you're confident in the quality of on arXiv?) https://arxiv.org/pdf/2007.00177.pdf
Do you have any insight on how consumers find this? I can imagine it's nice to hear (rarely) from a few places that you care about, but taken too far I think I'd find the mixing of messages from friends and companies pretty annoying.
Google have a blog post from October last year with some more complex examples of where more sophisticated NLP helps https://www.blog.google/products/search/search-language-unde...
I wonder what the impact on uptake would be if the focus was shifted towards CS as a venue for building things and being creative and away from lines of monowidth code and indecipherable errors. More to the point, I wonder how this might be done.
However... this didn't work in some cases, mainly with formatted text but sometimes with PDFs that looked like they were compiled in some nonstandard way. As a result I ended up chucking the XML structure entirely and recompiling the text from character-level coordinates. Formatted text was also an issue, with slightly offset y coordinates from regular characters on the same line.
I'm not sure I could take this experience and say that extracting _all text_ would be straightforward. Hopefully for most documents the XML is nicely structured, but I imagine there are many more opportunities for inconsistencies in how the PDF is generated when thinking about diagrams, tables etc. rather than just abstracts.
Considered writing up a blog post about my experiences with the above but imagined that it was far too niche. Code's here [1] if it's of interest.
[1] https://gist.github.com/GuyAglionby/4b55d00803710f2e2e9877fd...
Really impressed at the speed with which Hugging Face ported this to their transformers library -- Google released the model and source code Oct 21 [1] and it was available in the library just 8 days later [2].
[1] https://github.com/google-research/google-research/commit/b5...
[2] https://github.com/huggingface/transformers/commit/c0c208833...