Antrhopic's Haiku model can routinely pass —about 300 times so far--the extra class US ham exam based solely on what's contained in its model. (The LLM was not permitted, programatically I might add, to access the interent and therefore the question pool .) I'm studying this to reinforce my grad school business statistics class for use in writing agentic evals. Yesterday, I reasoned that the order the questions were asked might influence the results since the single LLM—essentially a single chat session—might get discouraged if it missed the first question and then project onto a 'poor test taker' part of its model space. I found out I was half right. The question order does seem to matter, but instead of influencing the entire exam, it influenced the results on the first question asked. I ran three 90 exam test fllets, one each with a first question that Haiku always gave the correct answer, the wrong answer, or a mixed result (sometimes correct, sometimes wrong...
The answer is very much, yes! I'm teaching myself eval statistics using ham exams taken by Claude LLMs as my demonstration vehicle. The picture below shows the disttribution of correct vs incorrect questions Claude Haiku and Sonnet on the extra class exam. The tiny text to the left of the rows is the test subelement, E0A, E1A, E1C, etc... Expect to see much more on this in the future. For a complete look at the data with statistics I can't promise are meaningful yet, see the initial report .