I'm still seeing Sonnet fan work out beyond its model capablities using subagents. You can see in the window below that the Gas Town polecat has had to stop and compact its context before proceeding.
The good news is that when these subagent work fans don't turn into storms, I'm seeing more thorough reseearch findings returned. The bad news is that with Sonnet's limited context, those findings occasionally overwhelm the polecat that started it all.
I found a preprint from MIT for using what the authors call recurive language models or RLMs. The general idea is basically what I'm seeing with subagents—split the task, split the context—but they have clever ideas about how to manage the output stages to avoid the exact overflows shown above.
I'm still harboring my hunch that this all started with the advent of Sonnet 5. Anthropic's release notes include
"Sonnet 5 is much more agentic than its predecessors. Testers described how it finishes complex tasks where previous Sonnet models would stop short, how it checks its own output without explicitly being asked..."
An eval of Sonnet 5 by Microsoft indicates that the new model is willing to do more work as well, but not always with a better outcome.
The subagent fan issues I'm seeing started soon after the release of Sonnet 5. I'm also seeing the "checks its own ouput" behavior more frequently now. That's actually not a feature for the current project I'm working on.
The final clue is that I'm seeeing Antrhopic advertise features to control workflow fans.
I'll keep you posted.
Comments
Post a Comment
Please leave your comments on this topic: