Skip to main content

Posts

Showing posts with the label LLM eval

Lab Book 2026-07-24: Encountering and Hopefully Fixing Subagent Overflow

 I'm still seeing Sonnet fan work out beyond its model capablities using subagents. You can see in the window below that the Gas Town polecat has had to stop and compact its context before proceeding. The good news is that when these subagent work fans don't turn into storms, I'm seeing more thorough reseearch findings returned. The bad news is that with Sonnet's limited context, those findings occasionally overwhelm the polecat that started it all.  I found a preprint from MIT for using what the authors call recurive language models or RLMs. The general idea is basically what I'm seeing with subagents—split the task, split the context—but they have clever ideas about how to manage the output stages to avoid the exact overflows shown above. I'm still harboring my hunch that this all started with the advent of Sonnet 5. Anthropic's release notes include "Sonnet 5 is much more agentic than its predecessors. Testers described how it finishes complex tas...

Lab Book 2026-07-20 Claude, Many Agents, and Subagents

 Just noting a feeling, nothing quantitative, but I think I've noticed that if more Claude agents are working at once, they each tend to create fewer subagents. If that's true, and it's a big if, I wonder if there's something at Anthropic that trottles or budgets overall effort to an account. Just to be clear, I've also watched five polecats descend into a complete and utter subagent storm like this one. True or not, I need the storms to either stop or I need to gain the ability to control them, so I'm working more on subagent storm control today. As promised, I did not watch my agents. Intead, I showed the mayor how to watch and mange polecats creating subagents. Here's the start of my conversattion At that stage, I hadn't considered the obvious fact in the figure above that subagents are occasionally beyond recursive. (As an aside, the kid here tells me this is actually a plot point in one of Matrh Well's MrubderBot books.) The next thing I asked...

Lab-Book 2026-07-19 How to Get Your Thinking Time back from LLMs

 I've seen a few posts in the last week that talk about engineer's loosing their time to think because they're too busy watching what their LLM agents are doing. Anthropic's models, even under GasTown, have decide to spontaneously fork and find other ways to farm work out to subagents, occasionally causing agent storms on my project in the last week, so I can sympathize a bit more now that I did at the start of the week. There are a few solutions I'd like to propose for this issue. Here goes. How to Get Your Thinking Time Back First, if you're agents are doing marginally what you want, this is pretty simple. Stand up, back away from your desk, and leave . If it makes you feel better, setup a notification so the agent messages you when it's done. Perhaps, something like this If you don't trust your agents for the moment, still don't watch them. Come up with metrics to know if they worked or not and control your agents or modify their activities AUTOM...

Lab Book 2026-07-19: Is the Antrhopic Increased Token Deal Related to Fan Experiments?

 Working with Fable last week, I noticed that it can get very excited about fanning out tasks, (to the point of crashing my WSL session.) I experimented with prompting agents to split out tasks on their own, but haven't found a successful to do this yet. Consequnetly, I reverted back to my original prompt which does not explicitly call for task splits. Then, this morning, running on Sonnet 4.6, another task storm sprung up. I wonder if the new 50% higher token limit  till August is Anthropics way of buying themselves some head  room while they experiment with models creating subagents?  Are other  people seeing the same thing? It's particularly worrisome that Sonnet 4.6 has started a subagent storm becuase it really doesn't ahve the context to deal with the results.

LLM Lab Book 2026-07-12: Claude fable-5 agent forks

 I'm still tracking down what makes some agents find a Nikola Tesla research finding and why others do not. Today, that's led me into investigating Claude CLI's forking harness. A few notes from me. Forking looks pretty spectacular! The agent kicks off a subagent that automatically has a copy of the parent's context. The subagent doesn't add to the parent's context until it's done. So, it seems to make things cheaper, at least for my passenger manifest research. The agent that used forking made the Tesla association. The other two agents with the same inputs and the same model did not use forking and did not find the Tesla assocation. This is important. It seems that agents that don't fork lack the persistence to look for more than one "really good" finding. They make that one good finding, and then kind of take any results for the rest of the passengs as good enough. Each forked subagent is looking for its own "really good" finding ...

Variance in Research; Manual Agent Orchestration - LLM Lab Book 2026-07-02

 More instances of variance in research results, this time around Isidore Nobel. Also, saving tokens by keeping conversations short. Missing Isidore Nobel This was probably due to a misspelling on the British manifest of Nobel's last name as Noble. The U.S. manifest has it as Nobel.  Another midweek usage reset I would say this was caused by the introduction of fable yesterday, but my usage numbers didn't reset until late last night. Yesterday, my limits were at about 21% and set to renew on Saturday. What do ticket number clusters reveal in the sorta solved Hedy Lamarr mystery? First, an update. Hedy Lamarr aka Hildegard Mandl is present on the English version of the manifest. You might remember she is not present on the United States side. No answers yet as to why. However, the English manifest has ticket numbers and people traveling together seem to have similar ticket numbers, so this is a reminder to look for something there. Why we sometimes want substrate-mediated pro...

W. E. D. Stokes - Ruddering From LLMs Back Towards Ham Radio

 While doing research for a book I'm working on, The Gladych Files , I wondered into the weeds of statistical analysis of LLM AI agent performance which relates to my everyday sort of work in engineering. One of the things I really enjoy about The Gladych Files, however, is that it's never long before the project pulls me back towards ham radio. The statistical analysis project involved determining how often, and with what certainty AI agents could find out that Lucia Hobson was the daughter of Rear Admiral Richmond Pearson Hobson and then make the further link that Nikola Tesla was the best man at Rear Admiral Hobson's wedding. While estimating how difficult this was to do with plain old human operated web searches this morning, I came across W. E. D. Stokes! Stokes came into the picture as Lucia Hobson's husband. What I didn't know was that he was one of the founders of The Radio Club of America. His original interest in radio came from wanting to control a mod...

LLM Agent Research Protocol for Avoiding Stigmergy - a Lab Notebook Entry

 I'm working through a methodology to study the behavior of teams of agents via observation of real-world tasks. As usual with LLMs, the concept of repeatable results is squishy, especially as compared to non-LLM deterministic computing. My finding last week was that LLM agents, especially Claude (per Google's research), can exhibit stigmergic , (a fancy word for how insects, like ants, 'learn' where important locations are from other insects), learning and behavior. In short, agents given the exact same instructions, (prompts), can and often times will exihibit different behaviors if they can see the results of the work of other agents. If you want to study the variance in the behavior of an LLM agent over multiple runs, this stigmergic behavior has to be accounted for. Otherwise, we're not measuring the behavior of an LLM agent with a set of inputs and prompts. With stigmergic behavior, if we're not careful, we're observing the behavior of a community of ...

LLM Evals Lab Book: The Importance of Statistics and Also Stigmergy

 Recap During an analysis of a travel manifest, two agents, (referred to as polecats in Gastown terminology), were accidentally handed the same manifest page for input. The agents produced different results. One agent found an association between Lucia Hobson and Nikola Tesla, a very valuable association for the research project. The other agent did not. A set of eval experiments ensued to determine how often polecats missed the association. The initial answer was that they missed it quite frequently with only 3 out of 16 agents making the association. Models Used In the following, all agents are using Sonnet 4.6. Orchestration is handled with Gastown. New Findings On the fourth batch of five test case runs, four polecats made the Tesla association. The chances of this happening randomly were less than 3% in the absence of any other process changes. Here's the Fisher's test run by Gemini. Fisher's Exact Test (Recommended) This compares your two distinct groups (the past 16...