How to Add an MCP Server to Claude Code: Step-by-Step
How to add an MCP server to Claude Code with claude mcp add: hosted, local and OAuth servers, scopes, .mcp.json, status checks and fixes for common errors.
10 min read
AI agent architecture is the set of layers and patterns that turn a language model into a system that plans, uses tools and acts. Here's how it breaks down.
A demo agent can answer ten questions correctly, get shipped, and then fall over the first time someone asks it something that needs three lookups across two systems. The model didn't get worse. What broke is everything wrapped around the model: the loop that decides what happens next, the memory that should have kept the user's last request, the tool call that quietly returned a stale row. That wrapping is what people mean when they say AI agent architecture, and it's usually the difference between a system that survives contact with real users and one that doesn't.
AI agent architecture is the set of layers, wired together, that turn a language model into a system that takes in context, reasons about what to do, acts through tools, and remembers what happened across steps. It spans everything from a single well-crafted model call up to several specialized agents coordinating through an orchestrator, and the right choice depends on how much of the task's path can be predicted in advance.
PuppyGraph maps the working parts of an agent onto four layers: perception, reasoning, action and memory. Perception is everything that puts the world into the model's context: the user's message, retrieved documents, prior tool outputs, and any event feeds the agent watches. Reasoning is the planning step, where the model decides whether to answer directly, call a tool, break the task into smaller pieces, or ask a clarifying question. Action is the tool call itself, a database query, an HTTP request, or handing a subtask to another agent. Memory holds the working state between steps and, in fuller architectures, a longer-term record the agent can draw on later.
Celigo splits that memory layer further into episodic memory, which stores what happened in past interactions, and long-term knowledge stores the agent queries on demand. Anthropic's engineering team calls the base building block the "augmented LLM": a model given retrieval, tools and memory, with the Model Context Protocol as one standard way to wire tools into that loop instead of writing a custom integration for every data source. PuppyGraph frames the whole exercise as answering four questions: where context comes from, how the agent decides what to do next, what it can actually do, and how you keep it from going off the rails. Most production failures, in its telling, trace back to one of those four rather than to the model itself.
Not every task needs a full agent loop. Microsoft's Azure Architecture Center lays out a three-level spectrum: a direct model call, a single well-crafted prompt with no agent logic and no tool access, handles classification, summarization, translation and other single-step tasks in one pass. A single agent with tools reasons over available tools, knowledge sources and APIs, looping through multiple calls to refine a result, and is described as "often the right default for enterprise use cases" since it's simpler to debug than a multiagent setup while still supporting dynamic logic. Multiagent orchestration, where an orchestrator or peer-based protocol manages multiple specialized agents, is reserved for cross-functional problems, scenarios needing distinct security boundaries per agent, or tasks that benefit from parallel specialization.

Microsoft's guidance is to use the lowest complexity tier that reliably meets the requirement, since each step up adds coordination overhead, latency and cost, and to set an iteration limit on a single agent with tools to guard against infinite tool-call loops. Anthropic's Building Effective AI Agents: Architecture Patterns and Implementation Frameworks makes the same point from the other direction: start with single-purpose agents that do one thing well, and only add complexity as requirements demand it, because simpler systems are cheaper to run, easier to debug, and give clearer metrics tied to business outcomes.
Anthropic draws a specific line between the two: workflows are systems where LLMs and tools are orchestrated through predefined code paths, while agents are systems where LLMs dynamically direct their own process and tool usage. Both can use the same model and the same tools. The difference is who decides the order of operations.
Anthropic names five workflow patterns that cover most production cases without needing a full autonomous agent. Prompt chaining decomposes a task into a sequence of steps where each call processes the previous one's output, useful for a draft-then-review-then-polish pipeline. Routing classifies an input and sends it to a specialized follow-up prompt or model, such as directing easy questions to a cheaper model and hard ones to a stronger one. Parallelization either splits a task into independent subtasks that run at once (sectioning) or runs the same task multiple times for a vote. Orchestrator-workers has a central LLM break down a task and delegate pieces to worker LLMs before synthesizing their results, well suited to tasks like a coding change where the number of files that need edits can't be known in advance. Evaluator-optimizer has one model generate a response while a second evaluates it and sends it back for revision, which works when a human's feedback would demonstrably improve the output and the LLM can supply that same kind of feedback.

A true autonomous agent goes further than any of these workflows: it plans and operates independently, gathers "ground truth" from tool results or code execution at each step, and keeps going until it completes the task or hits a stopping condition such as a maximum number of iterations.
Once a task genuinely needs more than one agent, Microsoft's Azure Architecture Center names five ways to coordinate them. Sequential orchestration chains agents in a fixed pipeline where each one processes the previous agent's output, the shape behind a law firm's contract pipeline of template selection, clause customization, compliance review and risk assessment agents run in that order. Concurrent orchestration runs several agents on the same input in parallel and aggregates their results, which a financial services firm uses to get several types of analysis on the same stock at once. Group chat orchestration lets agents converse to reach a decision together. Handoff orchestration passes a task between agents based on which one is best suited to the current step. Magentic orchestration adds a coordinating agent that adapts the plan as the task unfolds instead of following a fixed sequence.
PuppyGraph's own survey of production patterns lands on a similarly sized but differently named set. Single-loop ReAct agents fit a narrow tool surface and a well-defined task. Plan-and-execute suits long-horizon work where mid-task drift is a risk. Orchestrator-worker fits open-ended research that benefits from parallel exploration. Hierarchical or role-based agents, organized as a small org of planner, researcher, critic and writer, handle complex deliverables that need multiple personas. Reflexive or self-critique agents, where a primary agent's output is checked by a critic agent, suit quality-sensitive work like code or legal analysis.

PuppyGraph's advice for choosing among them is to match the pattern to the failure mode you actually see: plan-and-execute if the agent loses the thread on complex tasks, add a critic agent if it gets the right answer with poor judgment, and move to orchestrator-worker if it's too slow because everything runs in sequence.
CIO grounds these patterns in deployed systems. A conversational assistant gives users a chat interface backed by persistent memory and citations, the pattern behind a law firm's internal research assistant and a wealth management firm's client-facing portfolio assistant. A triggered workflow runs silently once an email, ticket or file event fires it, the shape a commercial insurer uses to classify inbound broker submissions and extract risk fields without anyone opening the tool directly. An autonomous agent with sub-agents plans its own steps and calls in specialists as needed, how a consulting firm's research agent chooses which internal databases and licensed data sources to query for a given engagement. A multi-agent team splits work across specialized agents with a proposer-critic loop in the middle, which a global bank uses to draft marketing copy, check it against brand guidelines, and check it again against disclosure rules before a human sees only the flagged disagreements. A human-in-the-loop agent automates the mechanical part of a task while keeping a person's approval on the sensitive parts, and CIO cites Moody's finding that 42% of compliance professionals consider that human oversight mandatory rather than optional.
OWASP names Excessive Agency as a distinct risk category, describing agent-based systems as ones that "will typically make repeated calls to an LLM using output from previous invocations to ground and direct subsequent invocations." The vulnerability's root cause, per OWASP, is excessive functionality, excessive permissions, or excessive autonomy, any of which lets a hallucination or a prompt injection turn into a damaging action instead of a wrong answer.
Futurice lists the failure modes that actually take down deployed agents: indirect prompt injection arriving through email, web content or tool output; over-privileged agent identities and connectors that widen the blast radius of a single mistake; tool timeouts and partial failures; non-idempotent actions that duplicate a write on retry; silent behavior drift when a model, harness or prompt changes; runaway workflow costs from repeated reasoning and retries; a compromised MCP server or tool in the supply chain; and long-running state that goes stale or corrupts mid-task. Futurice's fix is to treat the model plus its runtime as a replaceable execution layer and keep identity, policy, tool allowlists, approval gates, end-to-end tracing and cost budgets as a separate layer the company owns regardless of which model or managed platform sits underneath. Comparing agent observability tools for that tracing layer means choosing the piece of architecture that catches drift and cost overruns before they reach a customer.
Harness engineering, the runtime layer that mediates between the model and everything outside it, moved this year from something teams built themselves toward something vendors ship. Futurice dates Anthropic's Managed Agents to April 2026, making the split between model and execution runtime explicit, and reports that OpenAI's Agents API, which shipped in September 2026, exposes the managed Codex harness so a cloud agent can run for days at a time. Microsoft, Google and AWS have packaged similar runtime, tool, memory, identity and observability primitives into their own managed agent platforms over the same period, according to Futurice's reporting.
The multi-agent pattern specifically is growing fast: CIO's July 2026 piece cites Databricks research finding that multi-agent system usage grew 327% in four months as enterprises moved past single chatbots. Futurice also points to concrete outcomes from a 2026 State of AI Agents report: cybersecurity firm eSentire compressed expert threat analysis from 5 hours to 7 minutes, L'Oréal's orchestrated conversational analytics reached 99.9% accuracy across 44,000 monthly users, and Novo Nordisk cut clinical study documentation production from 10 or more weeks to 10 minutes using retrieval-augmented generation with approved text and case-specific variables.
AI agent architecture is the structural design that determines how an AI agent perceives input, reasons about what to do, accesses tools, and takes action, according to Celigo. It's the engineering layer around a language model that turns a single request-response call into a system able to plan, use tools and adapt across multiple steps.
Anthropic defines a workflow as a system where LLMs and tools follow predefined code paths, while an agent dynamically directs its own process and tool usage. A workflow is more predictable and cheaper to debug; an agent trades that predictability for the ability to handle tasks where the number of steps can't be known in advance.
Not usually as a starting point. Microsoft's Azure Architecture Center recommends starting with a direct model call or a single agent with tools, and moving to multiagent orchestration only when a single agent can't reliably handle the task due to prompt complexity, tool overload, or a requirement for separate security boundaries per agent.
OWASP names Excessive Agency, an agent given more functionality, permissions or autonomy than its task requires, as a distinct risk category. Futurice's field checklist adds indirect prompt injection through tool output, over-privileged connectors, and a compromised MCP server or tool in the supply chain as the failure modes that actually take down production deployments.
Celigo splits agent memory into episodic memory, which stores the record of past interactions, and long-term knowledge stores the agent queries as needed, while PuppyGraph treats memory as the layer holding working state between the steps of a single task. Both are what let an agent stay coherent across a multi-step task instead of treating every model call as a fresh start.
Comparing agents that put these architectures into practice is easier with a directory built for it, and AI Agents Listing tracks agents across the patterns covered here.
One email a week. New agents, MCP servers and skills, and what is actually getting traction.
How to add an MCP server to Claude Code with claude mcp add: hosted, local and OAuth servers, scopes, .mcp.json, status checks and fixes for common errors.
10 min read
Agentic workflows let a model choose its next step at runtime. Five patterns from Anthropic and LangGraph, real GitHub and IBM examples, and when to skip them.
8 min read
A comparison of open source AI agents: LangChain, CrewAI, OpenHands, goose, AutoGPT, Dify and n8n, with real license terms and 2026 release dates.
7 min read