Building a Data Analysis Agent: Context Engineering Design Principles
Key point
Prevent unnecessary context accumulation and improve exploration efficiency through progressive disclosure and sub-agent utilization
Details
As the autonomy of LLM agents increases, context engineering—the process of curating and organizing information to ensure the model accurately understands goals and progress—has become crucial. Based on experience developing a data analysis agent, this article explores how context was designed around three elements: the system prompt, tool specifications, and message arrays.
Separating Roles Between System Prompts and Tool Specifications
The system prompt defines the agent's role, safety rules, judgment criteria, and usage principles common across multiple tools. In contrast, detailed usage conditions for individual tools, the meaning of input values, and result delivery methods are specifically described in the tool specifications. Following Anthropic's guidelines, tool descriptions must be highly detailed; in the case of Claude Code, the tool specifications reached approximately 15k tokens, about four times the length of the system prompt.
When designing tools, the concept of 'deep modules' can be applied to simplify interfaces by bundling related functions into a single tool. For example, the creation and modification of SQL notes can be integrated into a single upsert_note tool, or the process from SQL writing to execution and error correction can be handled by a sub-agent to reduce the complexity of tool selection.
Progressive Disclosure and Sub-Agent Utilization
We applied the Progressive Disclosure approach, where the agent finds necessary information in stages. Initially, a separate exploration agent was used, but this was changed to a model where the main agent pre-recognizes the full list of accessible tables in the system prompt and queries detailed schemas only when needed. This reduced unnecessary exploration turns and allowed for faster identification of missing data situations.
Additionally, to address the issue of error messages and failure records from SQL writing and execution accumulating in the main context, sub-agents were introduced. Sub-agents handle SQL generation, execution, and error correction in a separate context, passing only the final results to the main agent. This prevents contamination of the conversation history and helps maintain the accuracy of the main agent's judgments.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.