Yea that’s a good point! On the one hand, it’s cool to look at the JavaScript the LLM generates to test its hypotheses in the logs (you can click “expand” to see them). On the other hand you’re right that especially given that it’s (presumably) an autoregressive LLM, if it writes the choice before the pattern there’s no way for the pattern to influence the choice.
You can muck with the prompts by clicking “Settings”! I think there’s a lot of room for improvement with all aspects of the experiment.
One prompt I tried made the LLM take tons of turns, but it ended up with a weird fixation with assuming everything was some sort of approximately geometric sequence and would always conclude none of the options matched. Before I gave it the eval_js tool it still did a pretty good job making choices, but the explanations were worse.