Agentic Design Patterns: A Book That Made Me Re-Understand "What Is an Agent, Really?"

链捕手Published on 2026-05-25Last updated on 2026-05-25

Abstract

"Agentic Design Patterns" is a 2025 book by Antonio Gullí, a Google engineering director, which offers a systematic framework for AI Agent development through 21 design patterns. A core contribution is the "Four Levels of Agency": Level 0 (bare LLMs) are not true agents. Level 1 agents actively decide when and how to use tools. Level 2 agents engage in strategic planning, context engineering (curating and filtering information), and self-reflection. Level 3 involves multi-agent collaboration with defined communication topologies. The book introduces **Context Engineering** as a superset of prompt engineering, managing four layers of information for the agent: system prompts, external data, implicit context (user history, environment), and feedback loops for automated optimization. A key pattern is **Reflection (Producer-Critic)**, where two distinct agents with different prompts collaborate iteratively—one produces output, the other critiques it—until quality is satisfactory or a max iteration limit is reached. For **Memory**, a three-layer model is proposed: Session (ephemeral conversation context), State (temporary task data), and Memory (persistent, long-term storage). Regarding **Multi-Agent Systems**, the book advises against unnecessary complexity, recommending simple topologies like Supervisor or Peer-to-Peer based on task needs. It emphasizes perfecting a single Level 2 agent before moving to multi-agent setups. The author concludes with three actionable takeawa...

Author: Yanhua

Antonio Gullí is an Engineering Director at Google. He wrote a 453-page book, breaking down AI Agent development into 21 design patterns.

But this is not a book review. My motivation for reading this book was specific: I've written about Harness Engineering, shared my experience with pitfalls in Clawdbot, and written "AI Agents Are Not Magic" about the seven turning points from burning tokens to becoming truly usable. After each piece, there remained an unanswered question: Is there an underlying, reusable logic behind all these things?

This book gave me an answer, and it went deeper than I expected.

What You're Writing Might Not Be an Agent At All

The most incisive judgment in the book is hidden in the prologue.

The "AI" most people are using is just Level 0: a bare LLM, with no tools, no memory, and no ability to act. You ask it which film won Best Picture at the 2025 Oscars, and it guesses. The book is blunt: Level 0 stuff is not an Agent.

Only the higher levels are true Agents:

  • Level 1: Tool User

    The Agent starts using tools: search, APIs, databases. But it's not just "able to call an interface"; it must decide *when* to call, *what* to call, and *how* to use the result. The book gives a concrete example: a user asks "What are some new TV shows?" The Agent realizes on its own that this information isn't in its training data, actively calls a search tool to find it, then synthesizes the results. The key step is "realizing on its own." It's not a human telling it "go search for this"; it judges *for itself* that a search is needed. This judgment ability is the threshold for Level 1.

  • Level 2: Strategic Thinker

    Adds two more things: Planning and Context Engineering. The book defines Context Engineering: it's not about dumping information, but about carefully selecting, trimming, and packaging context. A great example: a user wants to find a coffee shop between two locations. The Agent first calls a mapping tool to get a bunch of data, then judges that "only street names are needed for the next step," trims the map output into a short list, and feeds it to a local search tool. Every step is about reducing information noise.

    There's a sentence in the book I read several times: "To achieve the highest accuracy from AI, you must give it short, focused, and powerful context." Context Engineering is exactly about doing this.

    At this level, the Agent can also self-reflect. It reviews its own work after finishing, identifies issues, and makes corrections itself. I'll talk about this in more detail later.

  • Level 3: Multi-Agent Collaboration

    The book's stance is clear: stop trying to build one all-powerful super agent. The reliable approach is to build a team: a Project Manager Agent + a Researcher Agent + a Designer Agent + a Copywriter Agent. The example given is for a new product launch: a "Project Manager Agent" coordinates overall, assigning tasks to "Market Research Agent," "Product Design Agent," and "Marketing Agent." The key is communication: how Agents pass data, synchronize state, and handle conflicts. This chapter diagrams six communication topologies, from the simplest single Agent to the most flexible custom hybrids, with explanations for each scenario.

After reading these four levels, I suddenly understood why many people say "my Agent doesn't work well." The model isn't the problem; the problem is you're using it like a chatbot, and it might not even be at Level 1.

Context Engineering: The Book's Most Underrated Concept

I wrote an article about Harness Engineering, discussing how the design of the racetrack is more important than the engine's horsepower. After reading this book, I realized that Context Engineering is the mapping of Harness Engineering at the prompt level.

Traditional Prompt Engineering only cares about "how you ask." Context Engineering in the book cares about "what the Agent sees in front of it before it's asked." It includes four layers of information:

  1. First layer, the system prompt. Defines who the Agent is, its tone, its boundaries. Most people only write this layer.

  2. Second layer, external data. Documents retrieved via RAG, return values from tool calls, real-time API data. This is where most people get stuck: they know they need to feed data, but not how to do it without overwhelming the model.

  3. Third layer, implicit data. User identity, interaction history, environmental state. Things you don't explicitly state but the Agent should know. For example, if you tell the Agent, "Help me email John to confirm tomorrow's meeting," it should know what tomorrow's meeting in your calendar is and what your relationship with John is.

  4. Fourth layer, the feedback loop. After each output, the Agent automatically evaluates quality and adjusts the context strategy for next time. The book calls this "automated context optimization." Google's Vertex AI Prompt Optimizer is the engineering implementation of this idea.

When I read this part, I remembered my article "AI Agents Are Not Magic," which included the insight: "Your Agent needs rules, and a lot of them." Looking back now, those rules were essentially a manual version of Context Engineering; the book systematizes it.

Reflection: Two Agents Are Truly Better Than One

This is the pattern with the most practical value for me in the entire book.

The core of Reflection is simple: after finishing work, the Agent reviews itself, finds problems, and corrects them. But the implementation matters. The book states clearly: The Producer and the Critic must be two different Agents, with different system prompts. The same persona reviewing its own work will always have blind spots. If you let the same LLM write code and then review the code it just wrote, it will most likely say, "It's fine."

The book provides a complete code example.

  • The Producer's prompt is: "You are a Python developer. Write a function to calculate factorial, handle edge cases and exceptions."

  • The Critic's prompt is: "You are a nitpicking senior engineer. Review the code line by line, check for bugs, style, missed edge cases, and areas for improvement. If perfect, output CODE_IS_PERFECT, otherwise list all issues."

  • Then there's a for loop: Producer writes code → Critic reviews → Producer revises based on feedback → Critic reviews again → until Critic says CODE_IS_PERFECT or the maximum iteration count is reached.

It's that simple. But the book warns about a cost issue easily overlooked: each reflection loop is a new LLM call; the more iterations, the more expensive. Also, as the conversation history expands, the context window gets filled with earlier versions and criticisms, reducing the actual reasoning space available. So the best practice for Reflection is: set a reasonable maximum number of iterations (the book uses 3), stop once the Critic is satisfied, don't pursue perfection.

Its uses go far beyond writing code. Writing articles, making plans, summarizing documents, solving logic puzzles—the Producer-Critic model applies everywhere. The book lists seven application scenarios, with the same core logic: produce, review, revise.

Multi-Agent Isn't About Being More Complex

In the Multi-Agent Collaboration chapter, my favorite part is the six communication topology diagrams. Many people start with complex structures, but in reality, three are sufficient for most scenarios:

  1. Single Agent (Independent Execution): The task can be broken down into independent sub-problems, each handled by its own Agent. Simple, easy to maintain.

  2. Peer-to-Peer Network: Agents communicate directly with each other, with no central control node. Decentralized, good fault tolerance—if one Agent fails, it doesn't affect the whole. But coordination costs are high, and it can get chaotic.

  3. Supervisor (Centralized Orchestration): A Supervisor Agent manages a group of Worker Agents. Assigns tasks, collects results, resolves conflicts. Clear hierarchy, easy to manage. But the Supervisor is a single point of failure and a performance bottleneck.

The other three (Supervisor-as-Tool, Hierarchical, Custom Hybrid) are variations and combinations of the first three. The book is very practical: the topology you need depends on your task complexity. The more fragmented the task, the higher the communication cost. At a certain point, the Supervisor pattern becomes more efficient than the hierarchical one.

My takeaway is that many people building Multi-Agent systems spend 80% of their time on communication protocols, forgetting to ask a more fundamental question: does this task *really* need multiple Agents? The book is clear: a single Level 2 Agent with Reflection is often sufficient. Level 3 is for scenarios where a single Agent genuinely can't handle it.

The Three-Layer Memory Model: I Felt It Vaguely But Never Named It

I resonated most with the Memory chapter because when I wrote those two articles about Obsidian + Claude, I kept wondering: how should an Agent's memory be layered?

The book provides the answer:

  1. Session (Conversation Layer): The context window for the current conversation. This is the shortest memory; it's gone when the conversation ends. Long-context models simply enlarge this window, but it's still temporary, and each inference has to process the entire window, which is expensive and slow.

  2. State (State Layer): Temporary data during the current task. For example, "what is the ongoing task," "what step has been completed," "what intermediate data has been generated." Longer than Session, but cleaned up when the task ends. The book provides a complete example using Google ADK's State mechanism.

  3. Memory (Persistent Layer): Long-term memory across sessions and tasks. User preferences, learned experiences, important historical decisions, stored in databases or vector stores, retrieved semantically. The book emphasizes an important point: Memory isn't just about storing; you must design a full strategy for *what* to store, *when* to store it, and *how* to retrieve it. Store too much, noise increases; store too little, it's insufficient.

In my previous article about Clawdbot, I mentioned "state files" and "workspace documents," which were essentially handcrafting the State and Memory layers. The book has framed this.

Five Hypotheses, the Fifth Is the Most Outlandish

At the end of the book, it presents five hypotheses about the future of Agents. The first four are within reasonable speculation: General-purpose Agents evolve from writing code to managing projects; Deep Personalization proactively discovers your needs; Embodied Intelligence moves from screens into the physical world; Agents become independent economic entities.

The fifth one stunned me: Shape-Shifting Multi-Agent.

You only declare a goal, like "start a premium coffee e-commerce business." The system automatically decides: first create a "Market Research Agent" and a "Brand Agent." After running a round of data, it judges that the Brand Agent is no longer needed, splitting it into three new ones: "Logo Design Agent," "Website Builder Agent," "Supply Chain Agent." If the Website Builder Agent becomes a bottleneck, the system automatically replicates three parallel Agents to work on different pages simultaneously. Throughout the process, the system continuously auto-tunes each Agent's prompts and constantly restructures the team architecture.

The book calls this a "goal-driven, self-transforming multi-agent system." It's not executing a plan you wrote; it's generating the plan itself, adjusting the plan itself, and reorganizing the execution team itself.

This reminds me of Karpathy's AutoResearch: write a program.md, define goals, metrics, boundaries, and press "Launch." Humans are outside the loop. But this book pushes further: even how the Agent team is formed and restructured is left for the system to decide. Humans only declare "what they want."

Three Things You Can Do Immediately

After reading this book, I have three actionable items to implement immediately:

  • First, add a Critic to your current Agent. Whether you use Claude Code, CrewAI, or your own framework, add one step at the end of your existing workflow: have another Agent (with a different system prompt) review the previous step's output. Code generation + code review, article writing + fact-checking, plan creation + feasibility assessment. It's one more LLM call, but the quality improvement is often doubled. The book's Producer-Critic pattern is plug-and-play.

  • Second, start doing Context Engineering, not just Prompt Engineering. Go back and look at your instruction files for the Agent. If they are all rules about "how you should do things" but lack the context of "what environment you are currently facing," add it. Tell the Agent which project it's in, what decisions it made before, what the user's preferences are. The Context Engineering chapter in the book and your AGENTS.md are two expressions of the same thing.

  • Third, don't rush into Multi-Agent yet. Get your single Agent to Level 2 first: with tools, Reflection, and Memory. The book repeatedly emphasizes that a Level 2 single Agent with Producer-Critic and Context Engineering can cover the vast majority of practical scenarios. Level 3 is for truly cross-domain, multi-stage tasks requiring parallel division of labor. Most people's problem isn't having too few Agents; it's that they haven't even tuned one Agent properly.

This book is 453 pages, published by Springer in 2025. Code examples cover LangChain/LangGraph, Google ADK, CrewAI, and the OpenAI API. The foreword is written by Google Cloud AI VP, and there's a surprising and engaging recommendation preface from a Goldman Sachs CIO.

But my reason for recommending it isn't "comprehensive." It's because you'll realize something after reading: the pitfalls you've encountered with Agents in the past six months have been organized into patterns. You don't need to reinvent Reflection, guess how Memory should be layered, or experiment with which communication topology to use for Multi-Agent.

Someone has drawn the map for you. The rest is just walking.

Are you using AI Agents for development? What Level is your current Agent at?

Related Questions

QWhat are the four levels of AI Agent maturity described in the book 'Agentic Design Patterns'?

AThe book describes four levels: Level 0 (Bare LLM, not a real Agent), Level 1 (Tool User), Level 2 (Strategic Thinker with planning and Context Engineering), and Level 3 (Multi-Agent Collaboration).

QAccording to the article, what is the core difference between Prompt Engineering and Context Engineering?

APrompt Engineering focuses on 'how you ask,' while Context Engineering manages 'what is in front of the Agent before it asks.' It involves structuring four layers of information: system prompt, external data, implicit data, and feedback loops to provide the Agent with focused, actionable context.

QWhat is the 'Reflection' pattern, and what is a key practical implementation detail highlighted in the book?

AThe Reflection pattern involves having an Agent review and revise its own work. A key implementation detail is that the Producer (who creates) and the Critic (who reviews) must be two different Agents with different system prompts to avoid blind spots. The process involves iterative loops until the Critic approves or a maximum iteration limit (e.g., 3) is reached.

QWhat are the three main memory layers defined for AI Agents in the book's model?

AThe three memory layers are: 1) Session (the current conversation's context window), 2) State (temporary data for an ongoing task), and 3) Memory (the persistent, long-term storage for cross-session and cross-task information like user preferences and learned experiences).

QWhat are the three actionable recommendations the article author suggests after reading the book?

AThe three recommendations are: 1) Add a Critic Agent to your current workflow for review. 2) Start doing Context Engineering, not just Prompt Engineering, by providing environmental context. 3) Focus on perfecting a single Level 2 Agent with tools, reflection, and memory before rushing into Multi-Agent systems.

Related Reads

Senior Trader's Confession: How to Trade Market's False Expectations?

Veteran trader's case study: trading the market's "wrong expectations". This trade centered on a textbook "expectation error" after a weak CPI report. While the market initially priced in broad monetary easing (sending Nasdaq to 30,060), the crucial 30-year real yield hit a 20-year high. This signaled a fractured transmission mechanism: short-term rates eased, but long-term funding costs (vital for tech valuations) refused to fall. The trader executed five short positions on the Nasdaq (NQ) as it fell from 30,060 to 28,768. The core methodology: don't just trade the data, but analyze the market's implied causal chain and identify where it breaks. In this case, the chain was: Weak CPI → Policy Easing → Lower Long-Term Funding Costs → NQ Valuation Expansion. The break occurred between policy easing and long-term rates. The "veto variable" – long-term real yields – refused to confirm the bullish narrative. Trades were structured around "fast variables" (price) temporarily repairing while "slow variables" (funding conditions) remained broken. The article outlines a repeatable framework: 1) Map the market's implied causal chain. 2) Identify the veto variable. 3) Observe if it rejects the narrative. 4) Enter when price still follows the old script. 5) Choose the cleanest asset expression (e.g., short NQ, not broad S&P). 6) Define both invalidation and fulfillment exit conditions. The key insight: Alpha often comes not from an information edge, but from a "reaction function edge" – recognizing when the market is applying an outdated causal logic to new data. The critical question: What causal chain is the market's first reaction relying on, and is that chain still valid today?

marsbit24m ago

Senior Trader's Confession: How to Trade Market's False Expectations?

marsbit24m ago

Opinion: The Hedging Relationship Between U.S. Treasuries and Stocks Has Broken Down, and BTC, as a Risk Asset, Is Under Dual Pressure

For the past 20 years, U.S. investors relied on a free insurance policy: when stocks fell, bonds rose, cushioning portfolio losses. This reliable inverse correlation underpinned entire financial strategies. However, this mechanism broke down around 2020 and has not recovered. Currently, the two-month rolling correlation between the S&P 500 and 10-year Treasury yields is at -0.69, its lowest level since 1996, indicating stocks and bonds are moving in sync to an unprecedented degree, eliminating the traditional portfolio shock absorber. The失效 of this hedge is not simply due to lost confidence in U.S. debt. The key driver is the shift from growth-dominated to inflation-dominated market narratives. When growth fears prevail, stocks and bonds move inversely. Since 2022, persistent inflation volatility has been the dominant factor, causing both asset classes to suffer simultaneously from higher inflation expectations. Investors now seek safety without duration risk, favoring cash, dollars, and short-term Treasuries while selling long-duration bonds. Record U.S. deficits, rising net interest payments, and waning foreign demand (e.g., from Japan) are pressuring long-term yields, with the 30-year yield surpassing 5%. This environment places Bitcoin, as a risk asset on the far end of the risk curve, under dual pressure. Higher risk-free rates increase the opportunity cost of holding non-yielding assets like Bitcoin, while falling equities reduce overall risk appetite. Bitcoin's performance has become highly sensitive to macro conditions such as real yields, dollar strength, and financial conditions. While its long-term thesis as a fixed-supply asset outside the sovereign credit system is strengthened by these fiscal trends, the same conditions hurt it in the short term. The return of bonds as a effective hedge requires inflation volatility to subside, growth risks to retake dominance, and the Fed to have room to ease policy. Until then, Bitcoin trades in a market where the deepest asset class no longer absorbs shocks, removing the safety floor for all risk assets, especially those that pay nothing to wait.

marsbit24m ago

Opinion: The Hedging Relationship Between U.S. Treasuries and Stocks Has Broken Down, and BTC, as a Risk Asset, Is Under Dual Pressure

marsbit24m ago

Funding Weekly | Crypto.com Secures $400M Investment, CeFi and Stablecoin Sectors Continue to Attract Capital

Crypto Weekly Investment Recap: Funds Converge on CeFi, Stablecoins, and AI Last week's crypto and AI investment landscape saw significant capital concentration, with a few large deals dominating. **Crypto/Web3 Highlights (July 13-19):** Total investment exceeded **$812 million** across 17 deals. Key trends: * **Centralized Finance (CeFi) & Stablecoins** remained a major magnet, led by **Crypto.com**'s massive **$400 million** raise from Citadel Securities, valuing the exchange at $20B. * **Infrastructure & Tools** saw 7 deals, including **Cyclops** ($20M for stablecoin payments) and **ADI Chain** ($50M for stablecoin settlement infrastructure). * **DeFi** had 2 deals, such as AI-native DEX **Quote Trade** ($4M). * Other notable raises: API broker **Alpaca** ($135M), cross-border platform **Flex** ($70M), treasury firm **ORANGE JUICE** ($40M), and prediction market **Pascal** ($9M). * **Acquisitions:** Keyrock bought BlockFills' trading unit, SBI Holdings acquired Singapore exchange Coinhako, and MoonPay bought startup Glide. **AI & Robotics Highlights:** Investment momentum in AI remained very strong. * Nvidia-backed AI cloud service **Fireworks** raised a massive **$1.5 billion** at a $17.5B valuation. * In robotics, **Walden Robotics** (spun out from Toyota) secured **$300 million**, and Chinese humanoid firm ****逐际动力** **(Climax Dynamics) raised nearly **$200 million** in a Pre-IPO round. * Other significant AI raises included Indian programming platform **Emergent** ($130M) and drone company **Brinc** ($125M), backed by Sam Altman. Overall, the trend shows capital flowing heavily into established crypto financial services, stablecoin infrastructure, and large-scale AI/robotics commercialization.

marsbit1h ago

Funding Weekly | Crypto.com Secures $400M Investment, CeFi and Stablecoin Sectors Continue to Attract Capital

marsbit1h ago

The Gentlest Bear Market? BTC Bears Exit, ARK and Bitwise Collectively Bullish

**Title: The Mildest Bear Market? BTC Shorts Exit, ARK and Bitwise Collectively Bullish** Bitcoin continues to consolidate around $75,000 while Ethereum struggles near $1,900. Market data shows $116 million in liquidations over 24 hours, with $62.7 million from short positions, and the Fear & Greed Index remains at 35 (Fear). According to Polymarket, there's a 33% probability BTC falls below $50,000 this year. ARK Invest's Q2 2026 Bitcoin report notes technical weakness but identifies potential bottoming signals, including a record high in long-term holder supply, suggesting selling exhaustion. Bitwise's Juan Leon calls this the "structurally mildest bear market" on record, with a ~50% drawdown from highs, less severe than past cycles. He highlights institutional accumulation and a shift in investor dialogue from survival to entry points. Technical analysis from Bit suggests a potential C-wave low may have formed, with an ideal bottoming range between $50,000-$55,000. However, glassnode's CryptoVizArt warns that failure to break $66,000 could signal a local top, as new buyer accumulation is concentrated there. Analyst Darkfost identifies a critical support band between $59,000-$70,000, where 50% of BTC's circulating supply has changed hands. Notably, trader Doctor Profit announced closing all crypto short positions—including BTC shorts from $115k-$125k and over 100 altcoin shorts—for significant profit. He has begun a phased accumulation of Bitcoin spot starting at $64,000, reversing his previous $40k-$50k target, citing structural positives like regulatory clarity and institutional adoption. He argues the anticipated September/October bottom may arrive earlier than the herd expects.

Foresight News2h ago

The Gentlest Bear Market? BTC Bears Exit, ARK and Bitwise Collectively Bullish

Foresight News2h ago

Trading

Spot
活动图片