HNHacker News
TopNewBestAskShowJobs

grph123dot

14 karma · joined March 23, 2023

submissionscomments
grph123dot··on AI-enhanced development makes me more ambitious with my projects
Yesterday I asked chatgpt to compute the probability that two words of five letters both began with the same three letters and it could not solve it. It gave me about 10 different results but all were wrong, it seems the number five mislead the algorithm. Another question was computing the variance of the sum of 60 dices but omitting the results where the sum was greater than 200, the reasoning was incorrect in all cases, finally I suggested creating a simulation to estimate the variance in R, but it made an error about the var command in R suggesting that it has an option for the degree of freedom.
grph123dot··on Ask HN: Math for Programmers?
Perhaps the concepts of vector embedding and the attention mechanism could be explained. Also you could use some prompts for LLM to generate new content for your intended audience, mostly to classify them, so that your reader could copy your prompt and ask for further clarification.

Example: In chapter 5 here is a prompt that resembles the content of this chapter, to solve some of the exercises or clarify some points, your prompt would be along the lines ...

Just some ideas.

grph123dot··on CodeAlpaca – Instruction following code generation model
The Alpaca lora result is incorrect since, for example, if all elements in array are the same the result in just an array with one element.
grph123dot··on LoRA: Low-Rank Adaptation of Large Language Models
>> that the change in the weights as you train also have low intrinsic rank

It seems that the initial matrix of weights has a low rank approximation A and this implies that the difference E = W - A is small, also it seems that PCA fails when E is sparse because PCA is designed to be optimum when the error is gaussian.

grph123dot··on LoRA: Low-Rank Adaptation of Large Language Models
Since we are in ELI5, it seems that the concept of low rank approximation is required to understand this method.

(1) https://en.wikipedia.org/wiki/Low-rank_approximation

Edited: By the way, it seems to me that there is an error in the wikipedia page because if the Low-rank approximation takes a larger rank then the bound of the error should decrease, and in this page the error increases.

grph123dot··on LoRA: Low-Rank Adaptation of Large Language Models
Your explanation is crystal clear. I suppose it works well in practice, but is there any reason it works that well?
grph123dot··on ChatGPT can now call Wolfram Alpha
I think that google could use a LLM with wxmaxima and some database, for example using maxima the LLM could show the source code used and perhaps it could adapt the code to new problems.