1,956 karma · joined June 20, 2020
Are you open to pull requests? If I have the time I'd love to contribute. I'm sure others would as well.
You should write up a short article on this, even something really simple like one of the examples in the README but with some commentary and examples of output and then post it to a few places like https://dev.to/ or maybe https://hashnode.com/ or even Medium (even though I'm not a big fan). There aren't many newer implementations of PyTorch in JS and I've been looking for one to learn from for some time so I'm sure there are a lot of other JS/TS developers out there that would be interested. Getting to the front page of HN certainly helps but having an article somewhere will help everyone after this week find it through a Google search.
Again, thanks so much for doing this work! It's really helpful to have everything spelled out in JS for those of us who haven't used Python much (I'm sure Python devs can relate when they think about JS projects).
Absolutely amazing product by the way! The 1000 free tokens is enough, the fact that people are complaining about running out too soon is good, it shows that they like the product and want to use it more. They do have a point about adding a rough word count, maybe just a subheading that says "on average, X spoken words".
*edit: I should specify that when I say LLMs are a collection of facts and heuristics I mean they are a collection of those things encoded as language which itself has been encoded as vectors of floats which in turn have modified the weights of the network to produce yet another encoding. I don't mean that the facts and heuristics are stored in a lookup table or as procedures.
Your idea of splitting the probabilities based on whether you're starting the sentence or finishing it is interesting but you might be able to benefit from an approach that creates a "window" of text you can use for lookup, using an LCS[3] algorithm could do that. There's probably a lot of optimization you could do based on the probabilities of different sequences, I think this was the fundamental thing I was exploring in my project.
Seeing this has inspired me further to consider working on that project again at some point.
[0] https://github.com/karpathy/nanoGPT
[1] https://en.wikipedia.org/wiki/Trie
[2] https://en.wikipedia.org/wiki/N-Triples
[3] https://en.wikipedia.org/wiki/Longest_common_subsequence
Natural resources exist to be utilized. Once again, they provide the necessities and also the comforts that all deserve. If we limit our energy use our ability to extract natural resources will suffer. The resources we can access grow in proportion to the amount of energy we make available. No where is this relationship more direct than in the production of fresh water via desalinization and that alone should be sufficient incentive to utilize more energy. It takes resources and energy to develop more resources and produce more energy, you can't stop it or reverse it, you need to keep moving forward.
Fiat currency by definition is infinite being created by decree.
The rest of the points are increasingly wobbly so I'll leave you with this exchange from the comments on his page:
---
Christopher Toth You say that GPT-3 training consumed 700,000 liters of water, as if that is a large amount. With five seconds of research, I found that the global average water footprint for beef is around 15,415 liters of water per kilogram of beef produced, so an average cow costs >4.6 million liters of water. For a single cow. I am disappointed in your inability to contextualize the numbers you use.
Gary Marcus dude the context is that it will be way more for gpt-4, gpt-5 etc, but maybe you were unable to read that far.
Christopher Toth Okay, so can you speculate as to how much more water? Three cows worth? Ten cows worth? A hundred cows worth of water to train GPT-5?
Turns out we kill 900,000 cows every day, so around four trillion liters of water are used for beef production for a single day.
Do you expect GPT-5 to use more than this?
Otherwise why ever would you mention it other than because it looks like a large number to the uninformed?
---
How many cows indeed.
[0] https://www.hackster.io/news/the-lisperati1000-is-a-cyberdec...
[1] https://shop.m5stack.com/products/m5stack-cardputer-kit-w-m5...
I'm trying to imagine the simplest case, say a button you could press and every time you pressed it something entirely random would happen, always guaranteeing surprise. It would have a great deal of novelty at first but after a while it would cease to hold your attention even though your prediction of what would happen would never be accurate. I'd bet that after a while you might even never bother to push it again. The only way you would be convinced to push it consistently would be a) if you were assigned a reward for pushing it e.g. money in which case it is a slot machine or b) if by pushing it you could somehow reduce your uncertainty about what would happen which would as a by product reduce your surprise.
Thinking about it this way, surprise is certainly a key element at first. It grabs your attention initially but it doesn't hold it. What keeps you focused on exploring the thing that surprised you initially is the learning process which involves reducing prediction error i.e. reducing surprise. So there is a tension between the two.
The combination probably makes for a good exploration strategy. Initial surprise, look for a learnable pattern and follow it until another surprise, maybe backtrack and try other familiar patterns until those are exhausted and then investigate each sequence that led to a surprise by recursing through these steps.
This would also explain the example where my curiosity was prompted by the unknown link but I was not motivated to explore further. The website wasn't interesting to me because it was too unfamiliar and I wasn't able to find any familiar routes to explore through it due to my lack of interest in that area of mathematics but our hypothetical mathematician with a fondness for integers would see lots of familiar patterns they could explore attached to which are likely some enticingly unknown and surprising links.
Thanks for the prompt to think about this more!
You have to store wealth somewhere otherwise you'll constantly be losing wealth. Ideally, one could simply bank their savings without it losing value while also investing in a variety of other assets rather than putting everything into stocks or real estate but that's not the world we live in at the moment.
If you believe people should be paid for their work then you should also believe that the money they are paid for their work should not be bled away from them.
There's a long history of national building projects in many countries that have created problems like these.
The real-estate market in the West may be bonkers but at least someone in their prime earning years can invest in stocks if they can't purchase a home. That has its own risks but with a little knowledge you can mitigate them and with patience you can wait them out and even benefit from them.
> In 1917, California passed a new hotel act that prevented the building of new hotels with small cubicle rooms.[12] In addition to banning or restricting SRO hotels, land use reformers also passed zoning rules that indirectly reduced SROs: banning mixed residential and commercial use in neighbourhoods, an approach which meant that any remaining SRO hotel's residents would find it hard to eat at a local cafe or walk to a nearby corner grocery to buy food.