HNHacker News
TopNewBestAskShowJobs

aljungberg

266 karma · joined February 24, 2013

submissionscomments
aljungberg··on Eye scans detect signs of Parkinson’s disease up to seven years before diagnosis
That's awesome. It's not easy to get seniors to spend that much time in the gym, so well done if you contributed to motivate her! Physical exercise is indeed one of the few Parkinson's treatments that there's little doubt about.

In terms of reversing the damage, look into photobiomodulation (PBM) aka red light therapy. In "Improvements in clinical signs of Parkinson's disease", Liebert et al in 2021, she shows improvement in all symptoms including cognition, which is one of preciously few such results I've found in my extensive search of the literature[1]. Caveats are that this was a proof-of-concept study with n=6 only and that the red light helmet is somewhat expensive if you want to try it. There's a Canadian company that makes one for above $2000 (modern, very sci-fi thing), and a more hackerish version from an Australian company for ~$700 (pairs of diodes on aluminum bands you have to finish assembling yourself).

[1]: At least which are actionable for the public. There are tons of trials but good luck getting in early unless you can donate a new library to the university! Gene therapy, stem cell therapy, GDNF, drugs that target alpha-synuclein are all promising but not yet accessible. PBM is something you can do today and since mitochondrial dysfunction is a leading hypothesis in the pathogenesis of PD, the treatment fits.

aljungberg··on Llama: Add grammar-based sampling
We already do tree searches: see beam search and “best of” search. Arguable if it is a “clever” tree search but it’s not entirely unguided either since you prune your tree based on factors like perplexity which is a measure of how probable/plausible the model rates a branch as it stands so far.

In beam search you might keep the top n branches at each token generation step. Best of is in a sense the same but you take many steps using regular sampling at a time before pruning.

aljungberg··on OpenLLaMA: An Open Reproduction of LLaMA
To an extent, but memory bandwidth soon becomes a bottleneck there too. The hidden state and the KV cache are large so it becomes a matter of how fast you can move data in and out of your L2 cache. If you don’t have a unified memory pool it gets even worse.
aljungberg··on The LLama Effect: Leak Sparked a Series of Open Source Alternatives to ChatGPT
Your triton code is great, nice work. Wouldn’t feel too bad about spending your time that way!

As it happens I was also thinking it might be worthwhile to dive into the Triton sources but for another reason: half2 arithmetic. That’s one thing that the Triton branch lost that the (faster) CUDA kernels had and I think it made a difference. In theory with compatible hardware you can retire twice as many ops per second when processing float16 data which we are in this case.

Can’t see anyone having tried to get half2 to work with Triton though.

aljungberg··on Ask HN: Worth it to buy 4x Nvidia Tesla K40 for AI?
For some workloads, it’s almost all about the VRAM. In those cases I’ve been wondering if getting a high memory M1 or M2 Mac could be a good lab machine thanks to unified memory. It’ll run more quietly, use significantly less power, no worries about overloading your electric circuit. On a 128 GB RAM Mac Studio you could theoretically run or even train models that otherwise would require multiple $6k A6000 GPUs in custom machine builds taking oodles of power at the plug. It’d be slow but slow beats not possible. And if you need a new development machine anyhow, you can justify some of that beefy Mac Studio’s cost as part of your required spend anyhow. PyTorch has supported “mps” as a target device for some time now.
aljungberg··on Data consistency is overrated
Within a closed system, consistency (and verifying it) is a fail fast mechanism. For example, it’s better to crash on a constraint failure when attaching a doodad to a non-existent user account than to figure out where all these orphan doodads came from next year.
aljungberg··on Crypto exchange Kraken to shut staking service, pay $30M fine in SEC settlement
Whatever that number is, it will be equal or less than the number already discussed. The network “hires” contractors to provide the services you mentioned and it pays a known figure for that. Not much else to it really. Since all we are discussing is whether the network is profitable or not in this thread we don’t need to dig into more specific analysis of the service providers’ internal costs (and indeed that would be difficult since they are globally distributed with different attendant costs and efficiencies). Just to note they are unlikely to themselves be making a loss is sufficient.
aljungberg··on Crypto exchange Kraken to shut staking service, pay $30M fine in SEC settlement
It seems like you’re making a semantic argument to equate the Ethereum network with its validators. That seems confusing. Here are some examples of how “A runs B” does not imply “A == B”:

“Employees” are part of a company, they run the company, don’t they? Yet the cost of having employees is not “therefore revenue” from the standpoint of the company.

“Drivers” are a part of Uber, they deliver the service, don’t they? Yet the money paid to drivers reduce Uber’s profits.

I think where your argument runs into trouble is “from the standpoint of the network”. If you want to equate the network and its validators, to say they are the same thing, then your sentence becomes, “The money [the validators] get paid [by the validators] is therefore revenue, from the standpoint of [the validators]”. That’s non-sensical. You can’t give yourself money and say it’s revenue. Either these two things are in fact not the same thing and we can analyse the cashflow of “Ethereum the network” separately from “the validation service providers”, in which case Ethereum is paying out less than it’s taking in, so it is profitable. Or they are the same thing, in which case the “profit”, to the extent you can say a virtual entity like a network can have such a thing, is even higher.

This is because whatever costs the validators bear are less than the ETH they receive is worth. This is true if we assume validators are rational actors (they wouldn’t validate if they were losing money doing so). And even if we take away the assumption that they are profit motivated (maybe they’re all doing it as charity work for some higher purpose), the cost of running an Ethereum validator is tiny, so we end up in the same place: outgoings are smaller than receipts when considering the whole.

(The fact that Ethereum the network “burns” its receipts and then “mints” its outgoings to the validators does not affect this calculation since it’d work out the same if Ethereum paid validators from fees directly.)

aljungberg··on Showering at the South Pole
I wanted to provide a clever insight here saying they should bring some more solar panels, but unless my back of the napkin calculations are totally off that would work poorly. To begin with, there’s very little sunlight at the south pole and efficiency of panels is 10% of normal. Admittedly I’m only glancing through this paper, but an installation to generate an average continuous 2.5kW would have a total cost $250k [1].

The good news is that if you did do that, 2.5kW would allow you to heat a very significant amount of water per hour, even from ice to shower temperature. Like a US gallon in 4 minutes from 0º to 35ºC.

1: https://ir.canterbury.ac.nz/bitstream/handle/10092/14220/Jam...

aljungberg··on ChatRWKV, like ChatGPT but powered by the RWKV (RNN-based, open) language model
It does say on there they are training it on the Pile training data. And they have this bit comparing inference with GPT2-XL:

RWKV-3 1.5B on A40 (tf32) = always 0.015 sec/token, tested using simple pytorch code (no CUDA), GPU utilization 45%, VRAM 7823M

GPT2-XL 1.3B on A40 (tf32) = 0.032 sec/token (for ctxlen 1000), tested using HF, GPU utilization 45% too (interesting), VRAM 9655M

So it looks about twice as fast for inference while using only about 80% as much VRAM. Obviously at such a small size, just 1.5B, you can run it even on consumer GPUs but you could do that with GPT2 as well. If it remains 80% of VRAM usage when scaled up, we’re still talking 282GB once it’s the size of BLOOM w/ 176B parameters. So yeah still 8x A100 40GB cards I guess. Not going to be the Stable Diffusion of LLMs.

aljungberg··on ChatRWKV, like ChatGPT but powered by the RWKV (RNN-based, open) language model
THe RWKV model seems really cool. If you could get transformer-like performance with an RNN, the “hard coded” context length problem might go away. (That said, RNNs famously have infinite context in theory and very short context in reality.)

Is there a primer for what RWKV does differently? According to the Github page it seems the key is multiple channels of state with different decaying rates, giving I assume, a combination of short and long term memory. But isn’t that what LSTMs were supposed to do too?

aljungberg··on Ask HN: Working in a VR Headset?
I used the Quest 2. It wasn’t just the hardware though, something about the software too. The “main” display was a reasonably sharp and almost retina like. But the other displays were unable to keep up with that level of quality. Not enough bandwidth? Video encoding or decoding CPU bound? Not sure what it was. This was using Immersed
aljungberg··on Ask HN: Working in a VR Headset?
I have tried it. Working in VR is much better than I expected it would be, with the right equipment. Still, I quit after a while mainly because of the quality of text rendering. It is so much better than it used to be, but still not good enough, at least not for me.

I also found it surprisingly frustrating that I couldn’t see the keyboard. Every time I took my hands off it I had to do this blind search for it and get my fingers back to the right starting position. Symbols I don’t know how to touch type because they are rarely used was also more frustrating than I would have thought.

aljungberg··on Crypto trading firm Alameda Research might be insolvent
If those loans are no-recourse loans with this FTT token as collateral, then should the token crash the liability just "disappears". The collateral will be sold to cover the loan. If the collateral is now worthless that was the risk the lender agreed to take on when issuing a no-recourse loan.

If they are Defi loans for example, they're pretty much automatically no-recourse loans.

aljungberg··on Musk’s inner circle worked through weekend to cement Twitter layoff plans
Hasn’t it been a popular topic on here of how bloated Twitter’s staff seems to be given how little the platform has apparently changed over the years? Not taking a stance on that, I don’t know, but to that cohort of commenters this probably seems like victory for common sense.
aljungberg··on Using AI to compress audio files for quick and easy sharing
Yes, I believe it had been done with Google’s Wavenet for example where normally it’s trained to generates speech conditioned on textual input. For fun they trained it on piano music and gave it no conditioning and it improvised piano pieces.

Also see OpenAI’s jukebox.

aljungberg··on How the U.K. Became One of the Poorest Countries in Western Europe
Which goes to my point: the title is explicitly saying the UK is a "poor country" which is what makes it silly. You want to write an article about poor income distribution, that's fine, but your title should say that.

On the same note, the gini index is a kind of inequality measure. If your definition for "poor country" is one with uneven income distribution then the US is one of the poorest countries in the world. Like I don't know, I just don't feel like this definition of "poor" works.

aljungberg··on How the U.K. Became One of the Poorest Countries in Western Europe
Sure, the sixth largest economy in the world is “poor” because too much wealth is generated by financial services in the capital and not mom and pops in small towns or whatever. That’s like saying California is poor because most of the wealth isn’t generated by McDonalds employees. You picked some romantic ideal for how money “should” be made and voila now you can complain when it isn’t so.

The article actually wants to make a point about falling living standards, but the author couldn’t resist a clickbait title so nonsensical it overshadows the rest. There’s probably some interesting analysis that could be done about UK PPS, absolute, mean and median incomes, incomes vs average cost of living etc. This is not that article.

aljungberg··on Evidence-based conclusions about ADHD
Vyvanse is available in the UK although under the name Elvanse.
aljungberg··on A method to promote sleep in crying infants using the transport response [pdf]
Meaning it failed to detect a drop in oxygenation? Unless it fails 100% of the time, it would still be useful since “alerts on some non-zero percentage of life threatening conditions” is better than the alternative which is no alerts at all.

We had one and I think it made one or two false positives where basically it fell off. I can’t say anything about false negatives since none of our children stopped breathing. (It had two alerts, one for losing skin contact and one for low oxygen. The first one we got all the time when forgetting to pause the base station before removing the sock, so that bit certainly worked.)

aljungberg··on Ask HN: What's the best source code you've read?
Sounds like nbdev which indeed exports code and documentation from a single notebook, plus unit tests for external execution.
aljungberg··on Covid: Summary of lab-origin hypothesis
I disagree that people were spreading it in bad terms, if by that you mean the general public. Just look at what the officials actually said. Here’s March 2020:

> “You can increase your risk of getting it by wearing a mask if you are not a health care provider,” Surgeon General Jerome Adams said.

England’s chief medical officer, same month:

> Prof Whitty said: “In terms of wearing a mask, our advice is clear: that wearing a mask if you don’t have an infection reduces the risk almost not at all. So we do not advise that.”

Dr Fauci:

> “Right now, in the United States, people should not be walking around with masks,” said Dr. Anthony Fauci, an immunologist and a public face of the White House Coronavirus Task Force, on CBS’ “60 Minutes” earlier this month. He, like the others, suggested that masks could put users at risk by causing them to touch their face more often.

The WHO advised against it (to your point they did couch their language so much that it wasn’t even clear what they were really advising [1]).

And these are just the ones I could find quickly right now. From memory, the message was even stronger than this and even proliferated in this very forum. There was a time when you kind of had to duck and speak quietly if you wanted to bring up the idea that maybe this anti-mask thing wasn’t settled fact.

1: https://blogs.bmj.com/bmj/2020/03/11/whos-confusing-guidance...

aljungberg··on Covid: Summary of lab-origin hypothesis
It’s fine to make mistakes, especially in a fluid, developing situation. And it is precisely therefore we shouldn’t dress up our hypothesises as fact.

The actors here presented mask dictates as gospel. They were either wrong when anti-mask or wrong when pro-mask, but somehow conveyed absolute confidence in both cases. This is very harmful for the public discourse and for the reputation of the authorities during a time when reputation was paramount.

It’s okay not to know everything, just be honest about it.

aljungberg··on Covid: Summary of lab-origin hypothesis
The fact that we were strongly told no and then yes with equal conviction is sufficient to demonstrate the point, regardless of whether masks work. Whether intentionally or not, media and governments can mislead in an in effect coordinated fashion.

(Note there is a straw-man counterargument where one might say “so they changed their mind as evidence evolved, that’s allowed”. But the information was presented as established and factual in both cases. We have always been at war with Eastasia.)

aljungberg··on Faithful Reasoning Using Large Language Models
Is it only me or does the bear training data in Figure 13 not make sense? “Q: Is the bear loud? A: The bear is soft.”

And what are we to make of the reasoning step that “All round things are loud. The bear is sound. Therefore the bear is soft.” The bear is sound? What does that even mean?

Is that a typo that should read, “the bear is round”? If so, why isn’t the conclusion that the bear is loud rather than soft? From the context we do not know that all round things are soft, only that all soft things are round.

aljungberg··on DALL·E: Introducing Outpainting
GPT-3 was fine-tuned after release to be better at following instructions. I don’t think that’s been done for BLOOM.

BLOOM incorporates some new ideas like ALiBi which might make it better in a more general sense. They haven’t released official evaluation numbers yet though so we’ll have to see.

aljungberg··on OpenAI API pricing update FAQ
GPT-3 has been fine-tuned after release to better interpret prompts (see InstructGPT). Perhaps Bloom is more like the original GPT-3; a little more raw and requiring better prompt engineering?

In my small amount of testing of Bloom so far it seems capable of advanced behaviour but it can indeed be trickier to coax that out. Playing with temperature and sampling matters for sure.

aljungberg··on Do breastfed children have higher IQs? The answer is annoyingly hard to uncover
On the contrary, it was not easy at all but that’s a story for another day.

My argument is merely that our opening position, until proven otherwise, should be that the child benefits. (It can, unfortunately, simultaneously be hard for the mother and beneficial for the child. Many aspects of parenting are.)

And again I do think it’s important to note that huge numbers of children grow up well on formula. So if breastfeeding doesn’t work out for whatever reason there’s no reason to lose sleep over it. We can suspect it is beneficial and yet make a rational choice to forego it based on other factors.

aljungberg··on Do breastfed children have higher IQs? The answer is annoyingly hard to uncover
Yes evidently many children grow up fine on formula. So given that the effect of breastfeeding is capped, it’s reasonable to choose to optimise for other factors such as practicality, father-child bonding time and so on.

As a counterpoint I’ll say that to the extent that the transfer of antibodies prevents or reduces infant illness, that is a win in the dimensions you describe on its own. A sick infant is both wildly impractical and hard to bond with.

For example, in this interventional cohort study with breastfeeding promotion [1], ”the percent of children having pneumonia and gastroenteritis declined 32.2% and 14.6%, respectively, after the intervention.”

[1] https://www.nec.navajo-nsn.gov/Portals/0/NN%20Research/Biolo...

aljungberg··on Do breastfed children have higher IQs? The answer is annoyingly hard to uncover
I have to agree with the argument that the prior should be that breastfeeding is beneficial and we would need massive and strong evidence to the contrary to change that position. It is an expensive process so it must have an evolutionary advantage.

The analysis presented here seems convincing in the sense that we don’t have good evidence that breastfeeding increases IQ. But from what I remember there is good evidence that it improves the baby’s immune system among many other things.

So bottom line is in the absence of massive amounts of strong evidence against breastfeeding, by default you should favour it. And if in the end the IQ effect isn’t real at least you improved your baby’s health (and hence survival chances). If you have the option to breastfeed, go for it.

And as the article stated, if for whatever reason you cannot or will not you probably don’t need to feel too bad about it because the effects while positive are at the margin to the best of our current knowledge and formula is getting better every decade.

Page 1 of 3Next →