I have a basic grasp of algebra (including quadratic functions such as in your example).
So am I on the right track in extrapolating this to say that LLM parameters are the adjustable coefficients and constants in the chain of functions?
I have a basic grasp of algebra (including quadratic functions such as in your example).
So am I on the right track in extrapolating this to say that LLM parameters are the adjustable coefficients and constants in the chain of functions?
Here, (a_1, a_2, ..., a_n, b) are trainable, i.e. adjustable, parameters. (x_1, x_2, ..., x_n) are the inputs to the neuron. y is its output. Because linear functions are inflexible, neurons typically pass the output y to a non-linear function. The classical non-linear function is a sigmoid, but there are many choices. So you have z = sigmoid(y) = sigmoid(ax + b) = sigmoid(a_1 * x_1 + a_2 * x_2 + ... a_n * x_n + b). These non-linear functions usually have no trainable parameters.
Such a structure, a linear function composed with a non-linear function is the cornerstone of modern statistics, where it is called a generalized linear model (GLM). Hence, neural networks are just many GLMs or neurons grouped into layers, with layers connected to each other.