Antrhopic's Haiku model can routinely pass —about 300 times so far--the extra class US ham exam based solely on what's contained in its model. (The LLM was not permitted, programatically I might add, to access the interent and therefore the question pool .) I'm studying this to reinforce my grad school business statistics class for use in writing agentic evals. Yesterday, I reasoned that the order the questions were asked might influence the results since the single LLM—essentially a single chat session—might get discouraged if it missed the first question and then project onto a 'poor test taker' part of its model space. I found out I was half right. The question order does seem to matter, but instead of influencing the entire exam, it influenced the results on the first question asked. I ran three 90 exam test fllets, one each with a first question that Haiku always gave the correct answer, the wrong answer, or a mixed result (sometimes correct, sometimes wrong...