Skip to main content
← Back to Blog
2026-01-20 · 11 min readAI AgentsPythonUpdated Oct 03, 2026

Building AI Agents from Scratch: A Practical Guide

A hands-on Python guide to the four parts of every agent (model, tools, loop, memory), with a working tool-calling loop, guardrails, and a clear rule for when to reach for a framework or n8n instead.

Avnish Yadav
Avnish Yadav
Developer & Automation Builder
5 views
Building AI Agents from Scratch: A Practical Guide

The first version of this post was a five-line outline that promised a guide and never delivered one. This is the guide. We will build a small but real agent in plain Python: no framework, one SDK, about a hundred lines. Once you can see every moving part, frameworks such as LangChain and LangGraph stop feeling like magic, and you can make an honest decision about when to use one, when to write your own loop, and when a no-code tool like n8n is the better fit.

I use the Anthropic Python SDK and Claude Sonnet 5.5 (claude-sonnet-5-5) because the tool-use format is simple and well documented. The same ideas carry over to any model that supports tool calling. If you work in JavaScript, I already covered the Node.js version in How to Build an AI Agent in 10 Minutes; this post goes deeper on the loop, memory and guardrails, and stays in Python.

What an AI agent actually is

Strip away the marketing and an agent is four parts:

  1. A model that reads the conversation and decides what to do next.
  2. Tools: functions you write and describe to the model (read a file, query a database, call an API).
  3. A loop: your code sends the conversation to the model, runs any tool the model asks for, sends the result back, and repeats until the model gives a final answer.
  4. Memory and state: the message history for the current task, plus anything you persist between runs (a database, a JSON file, a vector store).

The model never runs your code. It returns a structured request ("call add_task with these arguments"), and your program decides whether to run it. That single fact is where all your control comes from, and it is why guardrails live in your loop, not in the prompt.

A chatbot is the model alone. A workflow is a fixed sequence of steps you wrote. An agent sits between them: you fix the tools and the limits, and the model chooses the order.

The example: a task-tracking agent

To keep things concrete, we will build an agent that manages a to-do list stored in tasks.json. You can ask it "What is still open?" or "Add a task to send the invoice by Friday, then show me everything due this week", and it will call the right tools in the right order.

Setup is two commands. The SDK reads your key from the ANTHROPIC_API_KEY environment variable.

pip install anthropic
export ANTHROPIC_API_KEY="<YOUR_ANTHROPIC_API_KEY>"

Step 1: write the tools as plain functions

Tools are ordinary Python functions. Keep them small, deterministic and strict about their inputs. They should raise a clear error rather than guess.

import json
from datetime import date
from pathlib import Path

STORE = Path("tasks.json")

def _load() -> list[dict]:
    return json.loads(STORE.read_text()) if STORE.exists() else []

def _save(tasks: list[dict]) -> None:
    STORE.write_text(json.dumps(tasks, indent=2))

def list_tasks(status: str = "open") -> list[dict]:
    if status not in ("open", "done", "all"):
        raise ValueError("status must be 'open', 'done' or 'all'")
    tasks = _load()
    return tasks if status == "all" else [t for t in tasks if t["status"] == status]

def add_task(title: str, due: str | None = None) -> dict:
    title = title.strip()
    if not 3 <= len(title) <= 120:
        raise ValueError("title must be 3-120 characters")
    if due is not None:
        date.fromisoformat(due)  # raises ValueError if not YYYY-MM-DD
    tasks = _load()
    task = {"id": len(tasks) + 1, "title": title, "due": due, "status": "open"}
    _save(tasks + [task])
    return task

Step 2: describe the tools to the model

The model only knows what you tell it. Each tool definition has a name, a description and an input_schema written in JSON Schema. Anthropic's docs are blunt about this: detailed descriptions are the most important factor in how well tool use works. Say what the tool does, when to use it, and what it will not do.

TOOLS = [
    {
        "name": "list_tasks",
        "description": (
            "List tasks from the user's to-do list. Use this before answering any "
            "question about existing tasks. Returns id, title, due date (YYYY-MM-DD "
            "or null) and status. Does not modify anything."
        ),
        "input_schema": {
            "type": "object",
            "properties": {
                "status": {"type": "string", "enum": ["open", "done", "all"]}
            },
        },
    },
    {
        "name": "add_task",
        "description": (
            "Add one new open task. Use only when the user clearly asks to add or "
            "create a task. 'due' must be an ISO date (YYYY-MM-DD); convert relative "
            "dates like 'Friday' using today's date from the system prompt."
        ),
        "input_schema": {
            "type": "object",
            "properties": {
                "title": {"type": "string", "description": "Short task title"},
                "due": {"type": "string", "description": "Due date, YYYY-MM-DD"},
            },
            "required": ["title"],
        },
    },
]

HANDLERS = {"list_tasks": list_tasks, "add_task": add_task}

Step 3: the agent loop

This is the heart of every agent. When Claude wants a tool, the response comes back with stop_reason set to "tool_use" and one or more tool_use blocks, each carrying an id, a name and an input. You run the tool and reply with a user message whose content is a list of tool_result blocks, each pointing back to its tool_use_id. Per the docs, the tool results must come first in that message, and you set is_error: true when the tool failed.

import anthropic

MODEL = "claude-sonnet-5-5"
client = anthropic.Anthropic()  # reads ANTHROPIC_API_KEY
SYSTEM = (
    f"You manage the user's to-do list. Today is {date.today().isoformat()}. "
    "Use the tools for every fact about tasks; never invent task data."
)

def run_tool(name: str, args: dict) -> tuple[str, bool]:
    handler = HANDLERS.get(name)
    if handler is None:
        return f"Unknown tool '{name}'. Available: {', '.join(HANDLERS)}", True
    try:
        return json.dumps(handler(**args)), False
    except (TypeError, ValueError) as exc:
        return f"{type(exc).__name__}: {exc}. Fix the arguments and retry.", True

def run_agent(user_message: str, history: list[dict], max_steps: int = 8) -> str:
    history.append({"role": "user", "content": user_message})
    for _ in range(max_steps):
        response = client.messages.create(
            model=MODEL, max_tokens=2048, system=SYSTEM,
            tools=TOOLS, messages=history,
        )
        history.append({"role": "assistant", "content": response.content})
        if response.stop_reason != "tool_use":
            return "".join(b.text for b in response.content if b.type == "text")
        results = []
        for block in response.content:
            if block.type == "tool_use":
                output, is_error = run_tool(block.name, block.input)
                results.append({"type": "tool_result", "tool_use_id": block.id,
                                "content": output, "is_error": is_error})
        history.append({"role": "user", "content": results})
    return f"Stopped after {max_steps} steps without a final answer."

Try it:

history: list[dict] = []
print(run_agent("Add a task to send the March invoice, due 2026-10-09.", history))
print(run_agent("What is still open?", history))

A few details that matter more than they look:

  • Append the full response.content, not just the text. The tool-use blocks (and any thinking blocks) have to be in the history so the next request can match each tool_result to its tool_use.
  • Errors go back to the model, not up the stack. A bad date becomes a tool_result with is_error: true and an instructive message. The model usually corrects the arguments and tries again. The docs recommend error text that says what went wrong and what to try next, instead of a bare "failed".
  • Every stop reason is a decision. Besides tool_use and end_turn, the API can stop with max_tokens, refusal or model_context_window_exceeded. My loop treats anything that is not tool_use as "done", which is fine for a demo. In production, log the stop reason and handle max_tokens (raise the limit or ask for a shorter answer) separately from a refusal.

Step 4: add a custom tool

With the registry pattern above, a new capability is three small changes: a function, a schema, and one entry in HANDLERS. Here is complete_task, which also shows a habit I keep for every tool that changes data: validate against real state, not just the schema.

def complete_task(task_id: int) -> dict:
    tasks = _load()
    for task in tasks:
        if task["id"] == task_id:
            if task["status"] == "done":
                raise ValueError(f"task {task_id} is already done")
            task["status"] = "done"
            _save(tasks)
            return task
    raise ValueError(f"no task with id {task_id}; call list_tasks to find the id")

TOOLS.append({
    "name": "complete_task",
    "description": (
        "Mark one task as done by its numeric id. Call list_tasks first if you do "
        "not know the id. Fails if the task does not exist or is already done."
    ),
    "input_schema": {
        "type": "object",
        "properties": {"task_id": {"type": "integer"}},
        "required": ["task_id"],
    },
})
HANDLERS["complete_task"] = complete_task

Notice the error message points the model at the tool that fixes the problem ("call list_tasks to find the id"). That one sentence saves a lot of confused retries. The same idea applies when you wrap a real API: return only the fields the model needs to decide the next step, and use stable identifiers. If you want the LangChain flavour of this, I covered it in Extending LangChain with custom tools.

Step 5: memory and state

There are two kinds of memory in this agent, and it helps to keep them separate in your head.

Short-term memory is the history list. It is what lets "What is still open?" follow on from the previous request. It grows with every tool call, and every request re-sends it, so long sessions get slower and more expensive. The simplest fix is to trim old turns, but you must never cut between a tool_use and its tool_result, or the API rejects the request. Trimming at plain-text user messages keeps the pairs intact:

def trim_history(history: list[dict], keep_turns: int = 6) -> list[dict]:
    turn_starts = [
        i for i, m in enumerate(history)
        if m["role"] == "user" and isinstance(m["content"], str)
    ]
    if len(turn_starts) <= keep_turns:
        return history
    return history[turn_starts[-keep_turns]:]

Long-term state is tasks.json. The agent does not "remember" tasks; it looks them up through a tool every time. I prefer this design for anything that matters: the source of truth lives in a store you control, and the conversation is disposable. For summaries, user preferences and retrieval over documents, I wrote a separate walkthrough on LLM agent memory basics.

Step 6: guardrails that live in code

Prompts are suggestions. Code is a guarantee. These are the guardrails I put into every agent, roughly in order of importance:

  • A step limit. max_steps stops a confused model from looping forever. Pick a number that fits the task and log every time it is hit, because a hit usually means a bad tool description.
  • Input validation in the tool. The JSON Schema guides the model, but the function still checks ranges, formats and existence. Anthropic also documents strict: true on tool definitions to guarantee inputs match your schema; it does not replace business-rule checks like "is this task already done".
  • An allowlist of tools per task. Only pass the tools the job needs. A read-only question gets list_tasks and nothing that writes.
  • Confirmation for destructive actions. For deletes, payments or emails, have the tool return a preview and require a second, explicit user confirmation before it executes.
  • Treat tool output as untrusted. If a tool returns web pages, emails or user uploads, that text can contain instructions aimed at your agent. The docs recommend keeping such content inside tool_result blocks and never pasting it into the system prompt.
  • Budget and logging. Each response carries token usage. Log it per run along with tool names, arguments and stop reasons. When something goes wrong you will want the trace; I go through how I read those traces in Debugging AI agents.

Do not skip the step limit

An agent without a step limit and with a write tool is the classic way to produce fifty duplicate records overnight. The limit costs one line.

Swapping the model (OpenAI, DeepSeek and others)

The loop does not care much which model sits inside it, which is the main reason I like writing it myself. The original version of this post mentioned OpenAI and DeepSeek. Both work with the same shape of loop: OpenAI has its own tool-calling format in the Responses API, and DeepSeek documents an OpenAI-format base URL (https://api.deepseek.com) and an Anthropic-format one (https://api.deepseek.com/anthropic). Message formats differ slightly between providers, so put the provider call behind one small function (call_model(history, tools)) and keep the tool registry, validation and guardrails provider-neutral. Then run your own test prompts through each model before switching; compatibility layers rarely support every parameter.

When to use LangChain or LangGraph instead

I write my own loop when the agent has a handful of tools, a single model, and runs inside an app I already own. The code above is easy to read, easy to test and has no dependency churn.

I reach for a framework when the problem stops being "a loop" and becomes "a long-running process":

  • LangChain gives you prebuilt agent loops and a large catalogue of integrations (models, vector stores, document loaders). The LangGraph docs themselves recommend starting with LangChain's prebuilt agents if you are new to this.
  • LangGraph describes itself as a low-level orchestration framework and runtime for long-running, stateful agents. It earns its place when you need persistence (resume after a crash), human-in-the-loop approval steps, a mix of fixed and agentic steps, or several agents handing work to each other.

My rule: if you are about to build your own checkpointing, resumable state machine and approval queue, stop and use LangGraph. If you are about to install a framework to call one model with three tools, write the loop. I compare the options in more detail in LangChain vs LlamaIndex vs a vanilla stack.

Where n8n fits: no-code orchestration

Not every agent needs to be code. n8n has an AI Agent node: you connect a chat model and one or more tool sub-nodes, and the agent decides which tool to call. n8n's docs note that you must connect at least one tool sub-node, and you can add memory as well. The Google Sheets node, for example, supports operations such as Append Row, Get Row(s) and Update Row, so "a task list in a spreadsheet" is a few nodes instead of a JSON file.

I use n8n when the agent mostly glues SaaS tools together (Sheets, Gmail, Slack, a CRM), when someone non-technical needs to see and adjust the flow, or when triggers and schedules matter more than custom logic. I keep custom Python when the tools need real validation, tests, or code that does not fit in a node. The two also combine well: a Python service with its own loop can sit behind an n8n HTTP Request node. For a hands-on start, see Create your first AI planning agent in n8n, and for a larger real build, my n8n Content Machine.

Checklist before you ship

  • Every tool has a description that says what it does, when to use it, and what it will not do.
  • Tools validate their inputs and return instructive errors with is_error: true.
  • The loop has a step limit and logs every stop reason.
  • Write tools are allowlisted per task, and destructive ones need confirmation.
  • History is trimmed at turn boundaries, never between a tool call and its result.
  • Tool output from outside sources is treated as untrusted.
  • Token usage, tool calls and arguments are logged per run.

Frequently asked questions

Do I need a framework to build an AI agent?

No. An agent is a model, tools, a loop and some state. The loop in this post is about thirty lines. Use a framework when you need what it adds, such as persistence, human approval steps or a large set of ready integrations.

Which model should I use for an agent?

Use one that supports tool calling and test it with your own tools and prompts. I used Claude Sonnet 5.5 here because it balances speed and capability. Check the provider's models page for current IDs, because they change.

How do I stop an agent from looping forever?

Cap the number of model calls per task in code (max_steps), and improve tool descriptions and error messages when you see the cap being hit.

How is this different from the Node.js agent post?

The Node.js guide gets a first agent running quickly in JavaScript. This one is Python and focuses on the loop internals, custom tools, memory trimming and guardrails.

Can the agent use Google Sheets instead of a JSON file?

Yes. Either write a tool that calls the Google Sheets API, or build the agent in n8n and use the Google Sheets node as the tool. The loop does not change, only the functions behind the tools.

Sources

Verified against the sources below on October 3, 2026. Products and docs change often: check the linked sources if something looks different.

  1. Claude models overview (model IDs)
  2. Claude API: Define tools
  3. Claude API: Handle tool calls
  4. Claude API: Handling stop reasons
  5. LangGraph overview
  6. n8n AI Agent node
  7. n8n Google Sheets node
  8. DeepSeek API docs (base URLs and models)
Share
Discussion

Comments

Loading comments...

Add a comment

Comments are reviewed before they appear.