An LLM's judgements can be verified as consistent with past judgements, or other criteria. A judge's cannot.
The bias can be quantified in advance of making actual rulings.
An LLM's judgements can be verified as consistent with past judgements, or other criteria. A judge's cannot.
The bias can be quantified in advance of making actual rulings.
You will analyze billions of weights the LLM is made from? What will you look for? You will analyze the trillions of text training inputs that created the LLM weights? What will your analysis do? How to get this "perfect insight" that you claim?
Assuming reproducible settings are used, that will give you perfect insight into the behavior in the tested circumstances, but unless it lets you reconstruct the entire network, it won't tell you the impact that untested changes even if they are recombinations of tested elements behave. It is not perfect insight into the biases.