HNHacker News
TopNewBestAskShowJobs

stephanheijl

344 karma · joined June 6, 2015

stephanheijl.com
submissionscomments
stephanheijl··on Show HN: ESM C, setting a new state of the art for protein language models
Love to see this on HN, very interesting research. I'm looking forward to the paper release and I appreciate the way that you are licensing these models. ESM 2-650M is still a solid baseline but seeing ESM C 6B outperforming it by these kinds of strides looks encouraging for the future possibilities of protein language models. Would be very interested to find out how well it performs on other benchmarks (ie ProteinGym zero-shot).
stephanheijl··on Honeycrisp apples went from marvel to mediocre
I would normally only buy apples around september/october in the Netherlands, trying to get them fresh from local orchards when possible. Elstar is amazing when plucked right from the tree, but becomes mealy before the end of the year IMO.

My new go to is the Magic Star variety, which has been sold as "Sprank" for the last 3 years at least in the Netherlands. These apples keep amazingly well; they ran out of stock around the summer in the last two years, but I found them delicious year round. I hope that this cultivar does not befall the same fate as the Honeycrisp, which I had the pleasure of tasting 5 years ago.

stephanheijl··on LoRA: Low-Rank Adaptation of Large Language Models
To be more exact, LoRA adds two matrices `A` and `B` to any layers that contain trainable weights. The original weights (`W_0`) have the shape `d × k`. These are frozen. Matrix `A` has dimensions `d × <rank>` (`rank` is configurable) and matrix `B` has the shape `<rank> × k`. A and B are then multiplied and added to `W_0` to get altered weights. The benefit here is that the extra matrices are small compared to `W_0`, which means less parameters need to be optimized, so less activations need to be stored in memory.
stephanheijl··on Why I’m Still on Strike: Portraits from the HarperCollins Picket Line
> Much too often, we are overworked and underpaid. We are in what people call a “passion industry,” one that ultimately capitalizes on our love of stories to excuse low wages and a “you better be grateful to this opportunity” attitude. We wish it was different. We’re not quite sure how to make it different.

The author identifies the situation completely accurately and is also able to see that they have no leverage whatsoever to make Harper Collins change their behavior. The fact is that doing a job that you are absolutely passionate about is something to be grateful for, and also something that is in high demand. Combined with the fact that the exact skill set required for this kind of work is currently possessed by a surplus of workers means that any position left by a striker will be rapidly fulfilled by a passionate scab.

I have no doubt that publishers will also readily take advantage of AI to make the pool of available jobs even shallower, which will decrease the viability of this movement even more.

stephanheijl··on Thanks to DALL-E, the race to make artificial protein drugs is on
Designing novel viral proteins might become trivial, but actually doing the lab work to produce them would still be a tough exercise. On the other hand, exactly by the mechanism that would give rise to such a novel pathogen, actors that do have access to large manufacturing capabilities would be able to create novel drugs rapidly or even prevantatively.
stephanheijl··on Ask HN: Where do you host images for your blog or landing pages?
I just add them to an img folder in my Github repository and it is then served as part of the Github page. Just using `src=img/picture.jpg` works fine.
stephanheijl··on A Web UI for Stable Diffusion
This is the dockerized version of this repo: https://github.com/AbdBarho/stable-diffusion-webui-docker
stephanheijl··on A Dyson sphere around a black hole
I am not claiming we should not have dams, or refrain from using dams to provide hydroelectricity. My point of contention is the assertion that hydro energy has less catastrophic failure modes than nuclear (presumably fission). Clearly the failure modes for hydro dams are at least on par with those of nuclear installations, given that we have historic evidence that they can inflict tens of thousands of casualties.

Hypothetical scenarios can be constructed for both methods of power generation (What if Braidwood plant melts down, somehow killing the 5 million people living within a 50 mile radius? What if the Three Gorges Dam busts, inundating an area inhabited by 600 million people?[1]) Either way, it is far from obvious to me that hydro has the superior safety profile, especially when their fatality rates are on the same order of magnitude even when the Banqiao incident is removed[2]*.

[1] https://web.archive.org/web/20210620174812/https://www.japan... p/opinion/2020/09/01/commentary/world-commentary/big-china-disaster/

[2] https://ourworldindata.org/safest-sources-of-energy

* Sovacool et al. (2016), the source of the data in [2] includes a hydro fatality rate per tWh 2.4x larger than nuclear when Banqiao is included.

stephanheijl··on A Dyson sphere around a black hole
> has less catastrophic failure modes

Given incidence of dam bursts like the 1975 Banqiao dam failure[1], with an estimate death toll of 26,000 to 240,000 people and flooding of over 12,000 square kilometers, I'm inclined to disagree with this assessment.

[1]https://en.wikipedia.org/wiki/1975_Banqiao_Dam_failure

stephanheijl··on AlphaFold Protein Structure Database
I'm impressed and grateful that DeepMind released this resource, this will save a lot of compute from labs trying to replicate an entire exome for themselves. While some structures look great, there are still some misses here. Important structures like BRCA1 (a well-studied breast cancer associated protein) are just structures for the BRCT and RING domains surrounded by a low-confidence string of amino acids, likely shaped to be globular: https://alphafold.ebi.ac.uk/entry/P38398

Maybe I was wrong for expecting the impossible here, but I was excited to see this specific structure and it appears that there is still work to do. Nevertheless, kudos to Deepmind on their amazing achievement and contributions to the field!

stephanheijl··on Just Be Rich
Based on the Global Wealth Book 2018 and 2019 from Credit Suisse, comparing Table 3-1 in both books, it appears that the number of adults in "under 10.000" range increased significantly in 2019. This is likely not a real change, but rather a change of methodology or data sources. As far as I can tell however, this is not detailed in the text. I have directed an email to Credit Suisse on this, as it's a rather interesting piece of data.
stephanheijl··on Just Be Rich
You linked an income inequality graph, which is distinct from wealth. If you look at the list on Wikipedia [1] and sort by Wealth Gini (2019) you will find that the Netherlands are #1.

[1] https://en.wikipedia.org/wiki/List_of_countries_by_wealth_eq...

stephanheijl··on A CO2 capture solvent with exceptionally low total costs of capture
I've been looking at the use of algae with regards to sequester img CO2 from the atmosphere. This seems to have some remarkable advantages: single molecules which make mass production easier, presumably less finicky operating procedure and way more straightforward to pump into disused oil wells. From the abstract it does seem to need CO2 being supplied to it as opposed to drawing it from the atmosphere actively. I assume this could be used in exhausts of some kind? Definite benefit is the fact that the CO2 is captured immediately as opposed to over a years long timeline, like trees.
stephanheijl··on Study finds people have short-lived immunity to seasonal coronaviruses
Important notes: No data is available on actual illness according to the article, which means that while there could be an immune response, that does not mean that symptoms were observed. This study also does not regard the impact of T-cells, which were determined to mediate the immune response to COVID [1]. The paper acknowledges this shortcoming.

[1] https://www.nature.com/articles/s41577-020-00436-4

stephanheijl··on Apple becomes first U.S. company to reach a $2T market cap
A large contributor to this would be the increase in investors pumping money in "safe" and easy ETFs. Instead of taking the time to investigate the market and looking into novel ventures, people want to ride the market into wealth.

If there ever was any social responsibility in investing, it would have been providing fluidity into new ventures and making markets efficient by making educated investments. I find this trend worrisome and I think it might be sending incorrect market signals.

stephanheijl··on Apple becomes first U.S. company to reach a $2T market cap
Something similar was commented below. These figures are not comparable, as GDP is "collected" every year. Revenue for Apple sits at 260B$, which is closer to the GDP of Vietnam.
stephanheijl··on Apple becomes first U.S. company to reach a $2T market cap
That's a pretty good bet, as Amazon is already sitting on a 1.6T market cap. Tesla is at a "meager" 342B$.
stephanheijl··on Apple becomes first U.S. company to reach a $2T market cap
The US military budget is an expenditure that occurs every year. Apple's market cap is the expected value of all of Apple's outstanding shares. These are really not comparable.
stephanheijl··on The hard part of the economics of Covid-19
That's not even close to true, schools, restaurants, pubs and sporting establishments are shut down. Events with 100 or more people attending are not allowed and the government is actively working to help companies deal with less productivity due to people working from home or in staggered shifts. They also recommend social distancing. This is not "Hey, we're going to let this run its course, take our lumps now, and try to get back to business as usual ASAP." as the grandparent mentioned.
stephanheijl··on The American Brain
A society’s mind is its marketplace of ideas, and the freer, more open, and more active that marketplace is, the sharper and clearer the giant mind is and the faster the pace of societal growth.

This is a good takeaway; even though the fringes of any discussion are generally filled with bad ideas, enabling people to have the discussion is paramount in finding the ideas that turn out to be "diamonds". It also permits us to gauge the quality of ideas that are not good allows us to consider why they are not. Sunlight is the best disinfectant.

stephanheijl··on Popularity of technology on Stack Overflow and Hacker News: Causality Analysis
Thanks for the answer, I missed that paragraph. It seems quite disconcerting to me that the average question in the Java tag has been rated negatively the last 4 years. I'm curious as to the specific reason for this downward trend, which isn't reflected in most of the other languages.
stephanheijl··on Popularity of technology on Stack Overflow and Hacker News: Causality Analysis
The graphs shown are cumulative, which is defined by the README to be the sum of all the points up to that date (the most common definition). However, some graphs show a dropping number of points, like the one for Java, even in the non-standardized plot. (https://raw.githubusercontent.com/dgwozdz/HN_SO_analysis/mas...) Would this indicate some sort of error in the data collection, or did I miscomprehend the "cumulative" label?
stephanheijl··on Edge displays “123456” in PDF but prints “114447”
I have also encountered something similar whilst attempting to print a ticket to a major amusement park in Europe using Edge. The page the tickets were on secured by a login mechanism, and attempting to print the tickets resulted a page with an error. I had to save the PDF to the computer and print from there to get the proper output. It definitely seems like Edge re-renders or even re-requests the PDF before printing.
stephanheijl··on Deep Text Correcter
Looks like a cool project, I would love to see this as a browser plugin of some sort. As for the corpus, I suspect that using articles from Wikipedia would be appropriate. Especially large articles are routinely checked and cleaned up. It has the added benefit of being available in multiple languages.

(https://en.wikipedia.org/wiki/Wikipedia:Database_download)

EDIT: I see this has already been suggested, along with a large amount of other source in another comment by daveytea.

stephanheijl··on The science of Westworld
> I think it's unlikely that the human mind tackles Go in the same way.

This is very certainly true, which is what makes AlphaGo interesting to watch and study. The human mind, even one that has trained on Go for years on end, will still work with abstractions and ideas that do not relate to the game. AlphaGo and other computers lack this attribute, as any and all abstractions they may have learned relate entirely to the game.

Any ideas about the "human perception" of Go they may have gleaned from games that are included in the initial training dataset, I suspect have long been supplanted by novel notions gathered during the phase where the Neural Nets played against themselves. These phases are documented in the AlphaGo blog from Deepmind[1].

I suspect that we may reach "human level intelligence", but that this intelligence will not arise in the same way. That is to say, computers will at some point match us in most tests of intelligence, but the solutions they devise will be completely novel.

[1] https://blog.google/topics/machine-learning/alphago-machine-...

stephanheijl··on Sweden reveals results from pilot of 30-hour work week
Then again, the hours that were cut also wouldn't have to be paid to the people not working them. In a perfect situation this would have led to higher employment numbers and happier employees. However, the Bloomberg source[1] also mentions that pay was not cut, which led to the overhead costs. You are right that this consequence could've been seen from miles away.

https://www.bloomberg.com/news/articles/2017-01-03/swedish-s...

stephanheijl··on Researchers create gel that regrows tooth enamel (2015)
Is it just me, or does it feel like we've been hearing about advances in dental treatment for decades, without them actually having any effect on the practice? From my personal experience as someone living in the Netherlands, where standards of health care are pretty high, I would expect these advancements to make at least _some_ impact on the field. Instead my (admittedly anecdotal) experience as a patient has been basically the same over the years.

2014 - "No more fillings as dentists reveal new tooth decay treatment" https://www.theguardian.com/society/2014/jun/16/fillings-den...

2011 - "An end to the dentist's drill: New painless cavity filler could be on the market in two years" http://www.dailymail.co.uk/health/article-2077816/Scared-den...

2004 - "No drilling, no filling in painless dentistry" http://www.telegraph.co.uk/news/uknews/1477044/No-drilling-n...

1998 - "Dental lasers – are they the safest way to fill your cavity?" http://judyforeman.com/columns/dental-lasers-are-they-safest...

stephanheijl··on A Mysterious Virus That Could Cause Obesity
Without regarding any other research on this topic, this particular paper has some interesting notes that may be considered. The sample size (52) is somewhat small, and the humans in this sample are already obese. (Even) the N-AGPT group (no antibodies found) had a 30.7 BMI on average. This makes it difficult to conclude that the virus was the cause of obesity based on this paper alone, as the entire sample set was obese by definition (BMI >30). At the very best one could infer that the level of obesity was affected by the virus to some extent.
stephanheijl··on Amazon LightSail: Simple Virtual Private Servers on AWS
I could see the appeal of some of the lower range hardware plans, but especially the higher tiers seem way off in terms of pricing. Can someone clarify why a 2 core/8GB machine costs 80$, where other IaaS providers charge far less for such rigs? (DigitalOcean gets you 4 cores for 80$/m, TransIP gets you 4 cores, 8GB and 300GB SSD for <55$/m...)
stephanheijl··on Show HN: Interactive deep convnet visualization for Keras and Tensorflow
Each of the channels (filters, in this case) is visualized separately. The user selects a layer and each of the image maps shown represents the activations of that channel.

You can verify this by looking at the `nb_filters` key in the JSON description of the layer on the left and counting the amount of image maps on the right.

Page 1 of 2Next →