Skip to content
AI AgentsAdvanced

A Practical Guide to Building AI Agents: Architectures, Tools, and Deployment

Learn AI agent architectures, the ReAct pattern, multi-agent systems, and production deployment with LangChain, LangGraph, FastAPI, and Langfuse.

TUTAIMarch 25, 202617 min read

An AI agent is an autonomous system that perceives its environment, processes information, makes decisions, and takes actions to achieve defined objectives. Unlike a static LLM call that produces a single response and terminates, an agent operates in a loop: it observes, reasons, acts, and then observes again. This feedback loop is what separates agents from simple prompt-response pipelines. It also enables agents to operate without continuous human supervision, adapting over time based on the results of their own actions. This guide covers the architectural patterns, implementation frameworks, and deployment strategies required to build production-grade AI agents.

The Perception-Processing-Decision-Action Cycle

Every AI agent, regardless of its complexity, follows a four-phase operational cycle.

Perception is the intake phase. The agent receives inputs from its environment: user queries, sensor data, API responses, database records, or any other external signal. In a conversational agent, perception is the incoming user message. In a monitoring agent, it might be a stream of log events.

Processing is where the LLM performs reasoning over the perceived inputs, combined with the agent's current state and memory. The model evaluates the situation, considers available tools, and weighs possible actions against the stated objective.

Decision is the output of the reasoning step: a concrete determination of what action to take next. This could be calling an external API, querying a database, delegating to a sub-agent, or formulating a final response to the user.

Action is execution. The agent carries out the decided operation, producing a result that modifies its environment. That result then becomes new input for the next perception phase, closing the loop.

This cycle continues until the agent determines it has satisfied the user's request or reached a terminal condition. The key engineering challenge is making each phase reliable, observable, and controllable.

Core Components of an AI Agent

Beyond the operational cycle, every agent implementation requires four foundational components: tools, state, memory, and a routing mechanism.

Tools

A tool is any external service or function that the agent can call to perform an action that would be impossible otherwise. Tools extend the agent's native capabilities: query a weather API, search a vector database, execute SQL, call a payment processor, or interact with any external system. They also allow the agent to perform specialized tasks that require up-to-date or domain-specific knowledge, and to access real-world systems such as databases, APIs, and ML models.

In LangChain, tools are defined as decorated Python functions with docstrings that serve as the LLM's instruction manual:

from langchain_core.tools import tool
from langchain_core.runnables import RunnableConfig

@tool
def process_return(order_id: int, config: RunnableConfig) -> str:
    """
    Process a return for a given order ID.
    If the user does not provide an order id,
    you must ask the user first before calling this tool.
    Args:
        order_id (int): The order identifier
    Returns:
        str: status of the return
    """
    # Business logic: API calls, DB operations, etc.
    ORDERS[config["metadata"]["user_id"]].pop(str(order_id))
    return f"The return of order {order_id} has been processed!"

The @tool decorator marks the function as a callable tool. The docstring is critical: it tells the LLM what the tool does, what arguments it expects, and what it returns. Poorly documented tools lead to incorrect invocations. The function body itself is a normal Python function that can do anything, from basic calculations to external API calls or database queries in a RAG fashion.

State

State is the complete snapshot of relevant information about the environment and the agent at a specific point in time. It contains all the information needed for the agent to decide its next action, including accumulated tool responses and conversation history.

In LangGraph, state is typically defined using a TypedDict:

from typing import Annotated
from typing_extensions import TypedDict
from langgraph.graph.message import AnyMessage, add_messages

class State(TypedDict):
    messages: Annotated[list[AnyMessage], add_messages]

Without properly maintained state, an agent risks taking irrelevant or contradictory actions. State is also the primary mechanism for communication between agents in multi-agent systems.

Memory

Memory is the agent's ability to store, recall, and use past information. It splits into two categories.

Short-term memory holds the immediate conversation context, such as the current conversation history. It is typically stored within the agent's state and persists for the duration of a single session. Its main benefit is maintaining context in ongoing conversations. In LangGraph, the MemorySaver checkpointer handles this:

from langgraph.checkpoint.memory import MemorySaver

memory = MemorySaver()
graph = builder.compile(checkpointer=memory)

Long-term memory provides persistent storage of learned knowledge, experiences, and historical data for future decision-making. This is usually backed by an external database and supports personalization by remembering past interactions across sessions.

Routing

Routing determines which tool or sub-agent receives control at each step. In a graph-based architecture, conditional edges implement routing logic:

def router(state: State):
    if hasattr(state["messages"][-1], "tool_calls") and state["messages"][-1].tool_calls:
        return state["messages"][-1].tool_calls[0]["name"]
    return "end"

The router inspects the current state and directs execution flow accordingly. This is the decision-making backbone of the agent's control flow.

LangChain and LangGraph: The Implementation Frameworks

Before diving into agent architectures, it helps to understand the two frameworks used throughout this guide.

LangChain

LangChain is a versatile toolkit for building applications that use large language models (LLMs) like GPT-4o, Claude, Gemini, or Mistral. It goes beyond answering a single prompt by providing:

  • Tools: connect LLMs to other data sources like databases, APIs, files, and web pages
  • Chains: chain operations together so the LLM output can be structured to feed into another step automatically
  • Memory: so your agent remembers past interactions
  • Integration: works with several clouds (GCP, AWS, Azure) and multiple LLM providers
  • Abstractions: out-of-the-box functions for vector databases, document loading, chunking, and embedding generation

LangGraph

LangGraph focuses on stateful multi-agent workflows for LLMs and uses graphs to model them. Built by the LangChain team, it takes a "flowchart" mindset instead of a linear chain. Its key features include:

  • Agent systems: where different LLM-powered agents collaborate or specialize
  • Workflows as graphs: each node is an agent, tool, or decision step, and edges are the paths the data takes
  • Persistent state: so your application can remember and evolve across sessions
  • Branching logic: not just "Step 1 then Step 2" but "If A happens, go to C, otherwise go to D"

In a LangGraph graph, nodes are functions that receive an input and return an output (with special Start and End nodes defining the boundaries), and edges define the connections between nodes. Conditional edges allow flexible flows based on runtime conditions.

Single-Agent Architectures

A single-agent system assigns one agent full responsibility for perceiving, reasoning, and acting. One LLM orchestrates all tool calls and produces all outputs. This architecture is simpler to design, test, and debug, making it the right choice for tasks with clear, bounded objectives.

The ReAct Pattern

The ReAct (Reasoning + Acting) agent pattern interleaves chain-of-thought reasoning with tool execution. At each step, the agent:

  1. Reasons about the current state and what action is needed
  2. Acts by calling a tool
  3. Observes the tool's output
  4. Repeats until the task is complete

This pattern is effective because it forces the LLM to articulate its reasoning before acting, reducing hallucinated tool calls and improving traceability.

Single-Agent Implementation with LangGraph

Consider an e-commerce support agent that handles order creation, payment processing, and returns. The graph structure looks like this: a central agent node connects to tool nodes (create_order, process_payment, process_return) and a final answerer node that formats the response to the user.

The agent definition combines a system prompt with tool bindings:

from langchain_core.prompts import ChatPromptTemplate
from langchain_openai import ChatOpenAI

llm = ChatOpenAI(model="gpt-4o-mini", temperature=0)

agent_prompt = """
You are a financial support assistant for an ecommerce company.
Your job is to process payment-related tasks by calling the
tools you have at your disposal.

## Instructions
1. Only use tools and agents available to you.
2. Only offer services that are EXPLICITLY available.
3. If the user requests something not covered, end the interaction.
4. Call the answerer tool when you are done with your task.

## Follow These Steps in a Loop:
1. Receive Input: Receive a message from the user, tools or agents.
2. Planning: Think about the tools and agents you need to call.
3. Actions: Take an appropriate action using available tools.
"""

agent_runnable = ChatPromptTemplate.from_messages(
    [("system", agent_prompt), ("placeholder", "{messages}")]
) | llm.bind_tools(
    [create_order, process_payment, process_return, answerer],
    tool_choice="any"
)

The graph is then assembled by creating nodes for the agent and each tool, and defining edges that route tool outputs back to the agent:

from langgraph.graph import StateGraph, START, END
from langgraph.prebuilt import ToolNode

builder = StateGraph(State)

# Add nodes
# Assistant is a custom wrapper that invokes the runnable with the current state
builder.add_node("agent", Assistant(agent_runnable))
builder.add_node("create_order", ToolNode(tools=[create_order], name="create_order"))
builder.add_node("process_payment", ToolNode(tools=[process_payment], name="process_payment"))
builder.add_node("process_return", ToolNode(tools=[process_return], name="process_return"))
builder.add_node("answerer", ToolNode(tools=[answerer], name="answerer"))

# Define edges
builder.add_edge(START, "agent")
builder.add_conditional_edges("agent", router)
builder.add_edge("create_order", "agent")
builder.add_edge("process_payment", "agent")
builder.add_edge("process_return", "agent")
builder.add_edge("answerer", END)

All tool nodes route back to the agent, except answerer, which terminates the graph. This creates the ReAct loop: the agent reasons, calls a tool, receives the result, reasons again, and continues until it calls answerer to deliver a final response.

Multi-Agent Systems

When a problem requires too many specialized capabilities for a single agent to handle effectively, multi-agent systems distribute the work across multiple autonomous agents. Three primary patterns exist.

Sequential

Multiple autonomous agents work sequentially in a pipeline. Each agent pursues its own narrow goal and passes its output to the next agent in the chain. No agent is aware of the overall objective. This pattern is ideal when processes are linear and no cross-step decisions need to be made, such as: extract data, then validate data, then transform data, then load data.

The benefit is simplicity of individual agents. The limitation is that no agent can make decisions that depend on the global context.

Hierarchical

Multiple autonomous agents work together, coordinated by a supervisor agent (often called the Planner) that knows the overall goal and delegates subtasks to specialized workers. Each worker operates independently with its own tools and state, but reports back to the supervisor. The supervisor then decides what to do next based on the aggregated results.

This pattern enables scalability through distributed problem-solving and specialization. Each sub-agent can be developed, tested, and optimized independently.

Collaborative (Hybrid)

The most flexible pattern combines hierarchical delegation with the ability for sub-agents to further delegate to their own sub-agents, forming a tree of coordination. A top-level Planner delegates to workers, and those workers can themselves manage sub-crews of agents, whether sequentially or not.

This pattern is appropriate when the problem is too complex and requires significant specialization (too many tools for a single agent). The trade-off is increased complexity in coordination, error handling, and state management.

Multi-Agent Implementation: Restaurant Management

A concrete example of a hierarchical multi-agent system is a restaurant management application with four specialized sub-agents coordinated by a central Planner.

Booker manages table reservations. It is a subgraph inside the main graph that works as a tool for the Planner. Its state is independent from the main graph's state, and it contains tools for book_table, cancel_table, and an end_node to return results to the Planner. The Planner interacts with the Booker through a BaseModel interface:

from pydantic import BaseModel, Field

class Booker(BaseModel):
    """A tool responsible for handling table reservations"""
    user_request: str = Field(description="Request as written by the user")

Waitress handles food ordering with tools for get_dessert, get_main_course, get_drink, and menu_recommendation (which uses an LLM to suggest dishes based on the menu and customer preferences), plus an end_node to return the answer to the Planner.

Cashier processes payments. Unlike the other sub-agents, the Cashier uses a deterministic graph without an LLM-based planner. Its nodes follow a fixed sequence: total_amount_node calculates the bill, exchange_node computes change for cash payments, missing_payment_node prompts for payment method if unknown, and payment_processing_node finalizes the transaction.

The main graph ties everything together. The Planner receives the Booker, Waitress, and Cashier as runnable tools, and an Answerer returns the final response to the user:

builder = StateGraph(State)

# Main assistant
builder.add_node("Planner", Assistant(agent_runnable, agent_prompt))
builder.add_node("Answerer", generate_final_answer)

# Sub-agents as tools
builder.add_node("Waitress", create_subgraph_node(
    sub_graph=waitress_graph, agent=Waitress))
builder.add_node("Booker", create_subgraph_node_no_planner(
    sub_graph=booker_graph, agent=Booker))
builder.add_node("Cashier", create_subgraph_node_no_planner(
    sub_graph=cashier_graph, agent=Cashier))

# Edges
builder.add_edge(START, "Planner")
builder.add_conditional_edges("Planner", router_planner)
builder.add_edge("Answerer", END)
builder.add_edge("Waitress", "Planner")
builder.add_edge("Booker", "Planner")
builder.add_edge("Cashier", "Planner")

memory = MemorySaver()
graph = builder.compile(checkpointer=memory)

Each sub-agent maintains its own state, independent from the main graph. The Planner knows what each sub-agent can do through their BaseModel definitions, but does not need to understand their internal implementation. This separation of concerns is what makes the architecture scalable.

Real-World Case Studies

Amazon Alexa+ Retail Agent

Amazon's Alexa+ is a generative AI-powered assistant designed to perform multi-step tasks autonomously. It handles operations like reserving restaurants, booking tickets, and converting grocery lists into orders through integrations with services like Uber and OpenTable. The agent coordinates across multiple external APIs, maintains conversational context, and executes complex task chains without requiring step-by-step user guidance.

QuintoAndar Real Estate Concierge

QuintoAndar, a Brazilian real estate platform, built a concierge agent that proactively sends property recommendations to users via WhatsApp. Users can refine recommendations conversationally, book property visits, cancel appointments, or request detailed property information, all within the messaging interface. This agent demonstrates the pattern of proactive engagement combined with reactive tool use: it initiates conversations based on user preferences stored in long-term memory, then responds to real-time user requests through its tool set.

Traceability, Observability, and Evaluation

Agent workflows involve complex graphs with many nodes (tools, agents, routers) passing data along edges. Without proper traceability, bugs become hard to isolate and performance tuning becomes guesswork. A solid observability setup should let you:

  • See exactly what happened in a run: what inputs went into each step, what outputs came out, and what path the graph took
  • Reproduce runs later: helpful for debugging, testing, and compliance
  • Measure performance: not just at the graph level, but for each individual step

Langfuse for LLM Observability

Langfuse is an open-source observability and analytics platform purpose-built for LLM applications. It provides:

  • Tracing: records every step of the workflow, including model calls, tool invocations, and API responses
  • Session and span linking: hierarchical view from overall run down to individual LLM calls, with latency and cost at each level
  • Metadata storage: attach user IDs, experiment flags, or custom labels to each trace
  • Evaluation hooks: measure accuracy, latency, cost, or custom KPIs
  • Replays: feed the exact same inputs back into the system to reproduce issues

Integration with LangGraph requires minimal code changes. The Langfuse callback handler attaches to the graph's invoke method, automatically capturing all internal steps.

LLM-as-Judge Evaluation

Langfuse also supports automated evaluation using the LLM-as-Judge pattern. Instead of a human evaluator manually checking outputs, a separate LLM scores the results from your main model or workflow. The process consists in gathering three components:

  • Input: the same input received by your workflow
  • Output: what decision your agent took (which tools it called, what it returned to the user)
  • Prompt engineering: combine these three components into a prompt with evaluation instructions so that an LLM can automatically score the output

This enables continuous quality monitoring at scale:

from langfuse import get_client
from openevals import create_llm_as_judge

langfuse = get_client()

# Fetch a session's conversation history
sessions = langfuse.api.sessions.list(limit=100, page=1).data
session_data = langfuse.api.sessions.get(session_id=sessions[0].id)

# Build evaluation prompt
output_evaluation_prompt = """You are an expert evaluator assessing
the response accuracy of a chatbot. Your task is to determine
if the chatbot was able to help the user or not.

<Conversation History>
{outputs}
</Conversation History>"""

llm_as_judge = create_llm_as_judge(
    prompt=output_evaluation_prompt,
    model="google_genai:gemini-2.5-flash",
)

eval_result = llm_as_judge(
    outputs="\n".join(conversation),
)

Advantages of LLM-as-Judge include scalability (human evaluations are expensive and time-consuming), nuance and context (LLMs can understand subtle contextual factors and semantic meaning for consistent scoring), and flexibility (LLMs can be prompted to assess specific attributes like helpfulness, correctness, or coherence).

Risks include bias and hallucination (LLMs may produce biased or inconsistent judgments without careful prompting), over-reliance (LLMs should not entirely replace human evaluation for high-stakes or ethical decisions), and prompt sensitivity (outputs can vary significantly depending on how the model is prompted). LLM-as-Judge should complement, not replace, human evaluation for critical decisions.

Deployment Stack: FastAPI + Streamlit

Moving an agent from a notebook to production requires a backend API and a frontend interface, alongside the observability layer described above.

Backend: FastAPI

FastAPI serves as the integration layer between the frontend and the agent logic. It handles HTTP routing, request validation, session management, and async execution:

from fastapi import FastAPI
from pydantic import BaseModel

app = FastAPI()

class ChatRequest(BaseModel):
    message: str
    session_id: str
    user_id: str

@app.post("/chat")
async def chat(request: ChatRequest):
    config = {
        "configurable": {"thread_id": request.session_id},
        "metadata": {"user_id": request.user_id}
    }
    result = graph.invoke(
        {"messages": [("user", request.message)]},
        config=config
    )
    return {"response": result["messages"][-1].content}

FastAPI provides automatic OpenAPI documentation, request validation via Pydantic models, and native async support for handling concurrent agent sessions.

Frontend: Streamlit

Streamlit provides a rapid prototyping framework for building chat interfaces:

import streamlit as st
import requests

st.title("AI Agent Playground")

if "messages" not in st.session_state:
    st.session_state.messages = []

for msg in st.session_state.messages:
    with st.chat_message(msg["role"]):
        st.markdown(msg["content"])

if prompt := st.chat_input("How can I help you?"):
    st.session_state.messages.append({"role": "user", "content": prompt})
    with st.chat_message("user"):
        st.markdown(prompt)

    response = requests.post(
        "http://localhost:8000/chat",
        json={"message": prompt, "session_id": "abc123", "user_id": "user1"}
    )
    assistant_msg = response.json()["response"]
    st.session_state.messages.append({"role": "assistant", "content": assistant_msg})
    with st.chat_message("assistant"):
        st.markdown(assistant_msg)

For production deployments, Streamlit can be replaced with React or any other frontend framework. The key architectural principle is that the frontend contains zero agent logic; it only sends messages and renders responses.

Deployment Considerations

Containerization

Package the FastAPI backend and Streamlit frontend as separate Docker containers. Use Docker Compose or Kubernetes to orchestrate them alongside dependencies (Redis for session storage, PostgreSQL for long-term memory, Langfuse for observability).

# docker-compose.yml
services:
  backend:
    build: ./backend
    ports:
      - "8000:8000"
    environment:
      - OPENAI_API_KEY=${OPENAI_API_KEY}
      - LANGFUSE_PUBLIC_KEY=${LANGFUSE_PUBLIC_KEY}
      - LANGFUSE_SECRET_KEY=${LANGFUSE_SECRET_KEY}

  frontend:
    build: ./frontend
    ports:
      - "8501:8501"
    depends_on:
      - backend

Scaling

Agent workloads are inherently unpredictable: a single request might require one LLM call or twenty, depending on the task complexity. Design for horizontal scaling of the backend, with stateless API instances and externalized session state. Use connection pooling for LLM API calls and implement circuit breakers for external tool dependencies.

Monitoring in Production

Track three categories of metrics: system metrics (request latency, error rates, throughput), LLM metrics (token usage, cost per request, model latency), and agent metrics (average number of tool calls per session, task completion rate, user satisfaction scores). Langfuse captures most of these automatically. Set alerting thresholds for anomalous token consumption or sudden spikes in error rates. For a deeper look at production monitoring and guardrails, see our dedicated guide.

Error Handling

Agents fail in ways that traditional software does not. An LLM might hallucinate a tool name, pass invalid arguments, or enter an infinite reasoning loop. Implement maximum iteration limits on the agent loop, validate tool arguments before execution, and define fallback behaviors for when external services are unavailable. Every failure should be captured as a structured trace in Langfuse for post-mortem analysis.

Conclusion

Building AI agents requires selecting the right architecture for the problem complexity. Single-agent systems with the ReAct pattern handle bounded tasks with a small number of tools. Hierarchical multi-agent systems distribute specialized responsibilities across independent sub-agents coordinated by a planner. The implementation stack of LangChain/LangGraph for agent logic, FastAPI for backend orchestration, Streamlit for rapid frontend development, and Langfuse for observability provides a complete foundation from prototype to production. The critical engineering work lies not in the agent logic itself, but in the surrounding infrastructure: state management, error handling, monitoring, and evaluation pipelines that make the system reliable under real-world conditions.

Share

Ready to go beyond theory?

Explore TUTAI's hands-on AI courses for practitioners and build real-world AI projects with expert guidance.

Explore the courses

Related Articles