Day 14/15 — Production Architecture: Governance, Monitoring & Error Handling
Governance for environments, tools and approvals; monitoring that joins Salesforce's API logs to your own tool-call log; an error taxonomy with a retry policy; limits and cost; and a typed error-handling module plus a runbook for Acme.

The pilot works. Neha asks about Acme Global Tech, Claude calls getAccountHealth, and a follow-up task appears after she clicks Approve. Now Arjun has to answer the questions a demo never asks. Who decides when a new tool goes on the server? How does a change reach production? When Neha says "it gave me a wrong answer on Tuesday", where does he look? What happens at nine on the last morning of the quarter, when her token has expired, the model service is busy and the org is close to its daily API limit? And what will all of this cost?
By the end of today you will have a governance matrix for Acme, a monitoring plan that joins Salesforce's logs to your app's, a typed error-handling module with a retry policy, and a runbook that an on-call engineer can follow.
Day 13 ended with an architecture decision record for Neha's assistant. Today we make that decision operable. This is the last architecture day of the series on Salesforce Headless 360 (now called AIforce), and it assumes the pieces from earlier days: hosted Model Context Protocol (MCP) servers, an External Client App, the custom tools from Day 12 and the approval step from Day 9.
Governance: environments, tools and approvals
Governance answers three questions: where changes are made, which tools exist, and who may use and approve them.
Environments: sandbox first, always
Custom servers can "Be deployed via Metadata API between sandboxes and production", so treat the McpServerDefinition like any other metadata: source control, review, deploy. Three details catch teams out when they move from a Developer Edition org to a sandbox pipeline:
- URLs differ per environment. Sandbox and scratch orgs use
/sandbox/in the hosted server URL. Keep URLs in configuration, as the capstone does withSF_MCP_SOBJECT_URLandSF_MCP_CUSTOM_URL. - External Client Apps don't always survive a refresh. "Local external client apps aren't copied to a new sandbox when you clone or refresh a sandbox. Only packaged external client apps are copied to the sandbox." Package the app, or plan to recreate it, and allow up to 30 minutes before a new one works.
- Beta features stay out of production paths. For the Headless 360 MCP Server, Salesforce's advice is: "make configuration changes in a sandbox or Developer org before you apply them in production."
Which tools, and for whom
Standard servers are "disabled by default and require explicit administrative action to enable", which gives you an inventory to govern. Three controls do most of the work.
- The External Client App decides who connects. One app per client ("one for Claude, one for ChatGPT, one for Cursor, etc."), Admin approved users are pre-authorized, a permission set, and a refresh-token lifetime of "30 days or less" with rotation.
- Permissions decide what each tool can do. Salesforce is explicit that "You cannot restrict access to a specific MCP server through the ECA configuration, but you control access to the tools that compose those servers." Class access on your Apex tools and object permissions on the standard ones are the real gates.
- Your client decides which tools the model may call right now. Anthropic's MCP connector docs recommend denylisting write tools "when you want a human confirmation step before state changes", which is exactly what the capstone does. They also warn that "An MCP server can change its tools at any time". The newer
mcp-client-2026-09-15beta header "records the tool list each server returns and lets you pin it".
If you run many MCP servers, not only Salesforce's, a gateway helps: "Products like MuleSoft AI Gateway provide centralized governance across all your MCP servers, not just Salesforce."
Who approves what
This is Acme's matrix. The roles are the point, not the names.
| Change | Proposed by | Approved by | Proven in | Evidence |
|---|---|---|---|---|
| Activate a standard server | Arjun (admin) | Security architect | Sandbox | Server inventory updated |
| New custom tool, or a changed schema | Developer | Tool owner and security | Sandbox, runAs tests |
Tests, description review |
| New MCP client or External Client App | Arjun | Security architect | Sandbox | Day 11 app checklist |
| Enable a write tool in the app | Product owner | Security and the business owner | Sandbox | Approval flow and write audit |
| Any beta server or feature | Arjun | Architecture board | Sandbox or Developer org only | An ADR with a review trigger |
| Model or system prompt change | Developer | Product owner | Staging, fixed test prompts | Before-and-after answers |
What changes in November (announced, not shipped)
Today every hosted MCP call runs as a person: Salesforce's docs say "There are no service accounts, no machine-to-machine flows, and no autonomous operation outside of user context." A Salesforce knowledge article from September 17 announces a change: "Targeting November, Salesforce will be releasing a new set of security controls, agentic registration, and a new billing model as part of the Headless Toolkit." Registration "carves out a discrete identity for each agent instead of letting it operate under the identity of the person it assists", and admins "can grant a narrower, reduced permission set to the agentic identity". Customers will need to "reconfigure your MCP clients (and agents that use APIs) to present the new OAuth credentials associated with the registered agent". The article adds: "Use of any Salesforce MCP server requires agent registration within three months' notice from Salesforce. This is for all customers."
No Setup path or release date has been published, so treat all of this as announced and subject to change. The practical step today is an inventory of every MCP client and External Client App you run, so that registration becomes a checklist rather than an investigation.
Monitoring and audit

What Salesforce records
Salesforce records the Salesforce side for you. Its security docs say "agent actions appear in standard Salesforce API logs with full user attribution. You can identify MCP traffic by filtering for records where API_CLIENT_CATEGORY = SALESFORCE_HOSTED_MCP." The security blog gives the path: "navigate to Setup and search for Event Log File Browser. Filter Event Type on "API Total Usage"", then check STATUS_CODE, USER_NAME and CLIENT_IP. Logs "can be downloaded as CSV files on a daily basis or retrieved programmatically." Salesforce doesn't say which editions or licenses generate this event type, and my Developer Edition org generated none: after dozens of hosted MCP calls, API Total Usage for 3 to 7 October gave "No event log files found" (2026-10-08). Confirm it in your production org; until then, API Requests and your own turn log are what you have. In my org the attribution held on the record itself too: a Task created through the capstone app showed the signed-in user as Created By (2026-10-04).

Three more places complete the picture:
- API usage: Setup → Company Information → API Requests, Last 24 Hours. "MCP tool calls consume API calls against your org's daily API quota." In my Developer Edition org the counter only went from 1 to 15 (of 15,000) on 2026-10-04, after dozens of hosted MCP tool calls, deploys and Apex test runs. Treat that as one observation, not a rule, and measure your own org.
- Tokens: "go to Setup, search for OAuth Usage, select your ECA, and revoke individual tokens or run a bulk revoke operation." Quick Find shows two pages with that name: use External Client Apps → OAuth Usage (click the user count to revoke). The older Connected Apps OAuth Usage page doesn't list External Client Apps.
- Records: "If the agent updates a record, the authenticated user's name appears in the audit trail as the editor."

What your app must log
Salesforce's governance guide names the gap: "Salesforce monitoring stops at the Salesforce boundary, but headless requests often cross multiple systems ... Because no single log captures the entire path, tracking requests across these systems is essential." Your app owns the rest of the path. For every turn, log a correlation ID, the Salesforce user ID, the model, Anthropic's request ID (every Claude API response carries a request-id header), the stop reason, token usage, and each tool call's server, name and error flag. Log approvals and the write audit too.
Don't log tokens, prompts that contain record data, or tool output. That matters twice here, because of Anthropic's retention terms for the connector: "The MCP connector is not covered by ZDR arrangements. Data exchanged with MCP servers, including tool definitions and execution results, is retained according to Anthropic's standard data retention policy." Your own logs shouldn't add a second copy.
This helper builds one structured line per turn from the Messages API response. It records names, IDs and counts only:
// lib/turn-log.ts
import type { ClassifiedError } from "./errors";
import { classifyStopReason, classifyToolResult, type McpToolResultLike } from "./errors-tools";
type Block =
| { type: "mcp_tool_use"; id: string; name: string; server_name: string }
| ({ type: "mcp_tool_result" } & McpToolResultLike)
| { type: "text" | "thinking" | "tool_use" };
export interface TurnLog {
correlationId: string; // also sent to the browser and written by every hop
userId: string; // the Salesforce user Id, not a name or email
model: string;
stopReason: string | null;
usage: { inputTokens: number; outputTokens: number; cacheReadInputTokens: number };
toolCalls: Array<{ id: string; server: string; tool: string; isError: boolean }>;
problems: Array<Pick<ClassifiedError, "kind" | "source" | "detail">>;
}
/** One structured line per turn: names, IDs and counts. Never tool inputs or record data. */
export function buildTurnLog(
meta: { correlationId: string; userId: string; model: string },
final: { stop_reason: string | null; content: Block[] },
usage: TurnLog["usage"],
): TurnLog {
const calls = new Map<string, TurnLog["toolCalls"][number]>();
const problems: TurnLog["problems"] = [];
for (const block of final.content) {
if (block.type === "mcp_tool_use" && "server_name" in block) {
calls.set(block.id, { id: block.id, server: block.server_name, tool: block.name, isError: false });
} else if (block.type === "mcp_tool_result" && "tool_use_id" in block) {
const call = calls.get(block.tool_use_id);
if (call) call.isError = block.is_error;
const problem = classifyToolResult(call?.server ?? "unknown", call?.tool ?? "unknown", block);
if (problem) problems.push({ kind: problem.kind, source: problem.source, detail: problem.detail });
}
}
const stop = classifyStopReason(final.stop_reason);
if (stop && stop.retry.action !== "continue") problems.push({ kind: stop.kind, source: stop.source, detail: stop.detail });
return { ...meta, stopReason: final.stop_reason, usage, toolCalls: [...calls.values()], problems };
}
Salesforce doesn't document a way to carry your correlation ID into its API logs, so join the two sides on user and time window. That is enough to answer "what happened to Neha's request at 10:42".
Alert on a few signals rather than dozens: a jump in auth problems (an app policy changed, or tokens were revoked), API requests approaching the org's daily limit, any executed write that doesn't match its approval, and spikes in refusal or max_tokens stops.
Error handling: a taxonomy and a retry policy
Every failure in this system falls into one of six kinds. Naming them is what makes the retry policy simple.
| Kind | Typical signals | Retry? | Neha sees |
|---|---|---|---|
| Auth | Hosted server answers JWT Token is required or Invalid token; a refresh fails with invalid_grant; Claude API 401 |
Refresh once, then sign in again; never loop | "Sign in again" |
| Permission | A tool reports that Salesforce denied access; Claude API 403 | Never | "Your user can't do that" |
| Validation | Validation rules and triggers fire, as on any entry point; a tool returns validation errors; Claude API 400 | Not automatically; a write needs a new approval | What to change |
| Limits | Org's daily API quota; query and search result caps; Claude API 429 or 413 | Only when retry-after is present |
"Busy" or "limit reached" |
| Model | Stop reasons max_tokens, refusal, model_context_window_exceeded |
No; pause_turn is continued instead |
A short explanation |
| Transient | Claude API 500, 504, 529; connection errors | Back off with jitter, idempotent calls only | "Try again in a moment" |
A few rules turn the table into a policy:
- One retry layer per call. Anthropic's SDK "automatically retries transient failures (such as connection errors, rate limits, and 5xx server errors) with exponential backoff, twice by default, honoring the
retry-afterheader when present." Don't wrap SDK calls in a second loop; tunemaxRetriesinstead, and keep your own retry helper for other calls such as a token refresh. - Retry only what is safe to repeat. Reads, yes. Writes only if the tool is idempotent by design, like
createFollowUpTask, and the approval covers exactly the same arguments. - A 429 without
retry-aftermeans stop. Anthropic's docs: "A tier spend-cap 429 has noretry-afterheader and keeps failing until access resumes." - A 200 isn't proof of success. With streaming, "an error can occur after the API returns a 200 response." If a stream fails after a write tool ran, read the record back before doing anything else, as on Day 9.
pause_turnis not an error. For that stop reason the docs say: "Send the assistant content back to continue." The capstone does this up to three times per turn.- Quota errors are a monitoring failure. Salesforce doesn't document the exact error a hosted server returns when the daily quota runs out, and "Each tool invocation counts as one or more API calls depending on the tool." Watch API requests and alert well before the limit.
Limits and cost
Salesforce API usage. Every tool call is at least one API call against the org's daily quota. Result sizes have limits too: soqlQuery allows a "Maximum 50,000 total records per transaction across all queries" and asks you to "Always include a WHERE clause to filter results and a LIMIT clause to control result size"; find "Returns a maximum of 2,000 records total." The beta Archive Connect tools add their own, such as a "Maximum 50 requests per hour per org" for unarchive.
Salesforce billing. Use the official wording, because it is still moving. The hosted MCP docs say the servers "are intended only for customers with Flex Credits and you may be billed for server usage." The September knowledge article adds that every successful call a registered agent makes, "whether it arrives over Model Context Protocol (MCP) or direct API—is metered as a Headless Platform Interaction (HPI)", that "The multipliers for this usage type haven't been announced yet, so it is currently listed as "TBA" and agentic usage isn't metered", and that "Sandboxes, scratch orgs, and Developer Edition orgs don't incur HPI metering through Digital Wallet". Calls to the Data 360 MCP Server "count as calls to the underlying Connect API endpoints for limits and Flex Credit usage", and Prompt Builder templates used as tools consume credits "at invocation time". Salesforce's own caveat applies to all of it: "Pricing and packaging are subject to change."
Model tokens. Each response reports input_tokens, output_tokens, cache_creation_input_tokens and cache_read_input_tokens; total input is the sum of the three input fields. Prompt caching helps an assistant like Neha's, because tool definitions and the system prompt repeat on every turn. The cache prefix is built from tools, then system, then messages, and an MCP toolset accepts cache_control. The default cache lifetime is five minutes, and cache reads are billed at a fraction of the base input price (0.1 times by default; check the pricing table for your model). Keep tool definitions and the system prompt stable, and ask for filtered, small tool results.
Where the connector runs. The Claude API MCP connector is beta on the Claude API, Claude Platform on AWS and Microsoft Foundry, and not available on Amazon Bedrock or Google Cloud. Check that before procurement picks a platform.
Hands-on: an error-handling module and a runbook
The module has four files, about 230 lines in all. I typechecked it with tsc --noEmit under strict mode, with the capstone's compiler options, and ran 16 small test cases against it, 15 for classification and retries and one for the turn log, using real error objects from Anthropic's TypeScript SDK. It has no dependencies, so it drops into the capstone's web/lib/ folder or any other TypeScript backend.
Step 1: one vocabulary
Every failure becomes a ClassifiedError with a kind, a retry decision, a message that is safe to show, and a detail line that is safe to log. ApiErrorLike lists the fields the SDK's APIError exposes, so the module needs no import from the SDK.
// lib/errors.ts: one vocabulary for everything that can fail in an AI + Salesforce request.
export type ErrorKind = "auth" | "permission" | "validation" | "limits" | "model" | "transient" | "unknown";
export type Source = "anthropic" | "salesforce-oauth" | "mcp-tool" | "app";
export type Retry =
| { action: "none" } // a person or a code change must fix it
| { action: "backoff"; retryAfterMs?: number } // safe to try again later
| { action: "sign-in" } // send the user through OAuth again
| { action: "continue" }; // pause_turn: send the assistant content back
export interface ClassifiedError {
kind: ErrorKind;
source: Source;
retry: Retry;
userMessage: string; // safe to show Neha
detail: string; // for your logs: never tokens, never record data
status?: number;
requestId?: string; // Anthropic's request-id, for support tickets
}
/** The fields the Anthropic TypeScript SDK's APIError exposes. */
export interface ApiErrorLike {
status?: number;
headers?: Headers;
type?: string | null;
requestID?: string | null;
message?: string;
}
const NONE: Retry = { action: "none" };
const BACKOFF: Retry = { action: "backoff" };
function retryAfterMs(headers?: Headers): number | undefined {
const raw = headers?.get("retry-after");
if (raw === null || raw === undefined || raw.trim() === "") return undefined;
const seconds = Number(raw);
return Number.isFinite(seconds) && seconds >= 0 ? Math.round(seconds * 1000) : undefined;
}
Step 2: classify Claude API errors
The status codes and their meanings come from Anthropic's errors page. The one judgment call is the 429 without retry-after, which is treated as a limit to wait out, not to retry.
// lib/errors.ts (continued)
/** Errors thrown by the Messages API client (HTTP status codes from Anthropic's errors page). */
export function classifyAnthropicError(error: unknown): ClassifiedError {
const e = (error ?? {}) as ApiErrorLike;
const base = {
source: "anthropic" as const,
status: e.status,
requestId: e.requestID ?? undefined,
detail: `${e.type ?? "no_type"}: ${e.message ?? String(error)}`,
};
switch (e.status) {
case 400:
return { ...base, kind: "validation", retry: NONE, userMessage: "The assistant sent a request the model service rejected. It has been logged." };
case 401:
return { ...base, kind: "auth", retry: NONE, userMessage: "The assistant can't reach the model service right now. An admin has been alerted." };
case 403:
return { ...base, kind: "permission", retry: NONE, userMessage: "The assistant isn't allowed to use that model feature. An admin has been alerted." };
case 413:
return { ...base, kind: "limits", retry: NONE, userMessage: "This conversation is too large. Start a new chat." };
case 429: {
const after = retryAfterMs(e.headers);
// A 429 without retry-after can be a spend cap that keeps failing: don't loop on it.
return after === undefined
? { ...base, kind: "limits", retry: NONE, userMessage: "The assistant has reached a usage limit. Try again later." }
: { ...base, kind: "limits", retry: { action: "backoff", retryAfterMs: after }, userMessage: "The assistant is busy. Retrying shortly." };
}
case 500:
case 504:
case 529:
return { ...base, kind: "transient", retry: BACKOFF, userMessage: "The model service is busy. Try again in a moment." };
case undefined:
// No HTTP status: a connection error or timeout before any response.
return { ...base, kind: "transient", retry: BACKOFF, userMessage: "The model service could not be reached. Try again in a moment." };
default:
return { ...base, kind: "unknown", retry: NONE, userMessage: "Something went wrong. It has been logged." };
}
}
Token refresh failures follow OAuth 2.0: invalid_grant means the refresh token "is invalid, expired, revoked, does not match the redirection URI used in the authorization request, or was issued to another client". Only a new sign-in fixes that.
// lib/errors.ts (continued)
/** Errors from Salesforce's OAuth token endpoint, for example during a refresh (RFC 6749 section 5.2). */
export function classifyTokenError(status: number, body: { error?: string; error_description?: string }): ClassifiedError {
const base = { source: "salesforce-oauth" as const, status, detail: `${body.error ?? "no_error_code"}: ${body.error_description ?? ""}`.trim() };
if (status >= 500) {
return { ...base, kind: "transient", retry: BACKOFF, userMessage: "Salesforce sign-in is unavailable. Try again in a moment." };
}
if (body.error === "invalid_grant") {
// Refresh token expired, revoked, or issued to another client: only a new sign-in helps.
return { ...base, kind: "auth", retry: { action: "sign-in" }, userMessage: "Your Salesforce session ended. Sign in again." };
}
return { ...base, kind: "auth", retry: NONE, userMessage: "Salesforce sign-in is misconfigured. An admin has been alerted." };
}
Step 3: classify tool results and stop reasons
Tool failures arrive as data inside mcp_tool_result blocks, not as exceptions. The auth strings are the bodies a hosted server returns without a valid token, checked against the live endpoint on September 30; the others are the messages the Day 12 Apex tools write. Add your own tools' messages as you build them, and replace the text matching with structured fields wherever your tools return them.
// lib/errors-tools.ts
import type { ClassifiedError } from "./errors";
/** The mcp_tool_result block shape from the Claude API's MCP connector. */
export interface McpToolResultLike {
tool_use_id: string;
is_error: boolean;
content: Array<{ type: string; text?: string }>;
}
// Bodies a Salesforce hosted MCP server returns without a valid token (observed 2026-09-30).
const AUTH_TEXT = ["JWT Token is required", "Invalid token"];
// Messages written by the capstone's own Apex tools (Day 12).
const PERMISSION_TEXT = ["Salesforce denied access", "Salesforce denied the request", "do not have access"];
const VALIDATION_TEXT = ["Task not created."];
/** Classify one tool result. Returns null when nothing went wrong. */
export function classifyToolResult(server: string, tool: string, result: McpToolResultLike): ClassifiedError | null {
const text = result.content.map((part) => part.text ?? "").join("\n");
const has = (needles: string[]) => needles.some((needle) => text.includes(needle));
const base = { source: "mcp-tool" as const, detail: `${server}/${tool} (${result.tool_use_id})` };
if (has(AUTH_TEXT)) {
return { ...base, kind: "auth", retry: { action: "sign-in" }, userMessage: "Your Salesforce session ended. Sign in again." };
}
if (has(PERMISSION_TEXT)) {
return { ...base, kind: "permission", retry: { action: "none" }, userMessage: "Salesforce didn't allow that for your user. Ask your admin if you need access." };
}
if (has(VALIDATION_TEXT)) {
return { ...base, kind: "validation", retry: { action: "none" }, userMessage: "Salesforce rejected the values. The assistant will explain what to change." };
}
if (result.is_error) {
return { ...base, kind: "unknown", retry: { action: "none" }, userMessage: "A Salesforce tool failed. The assistant can see the error and will explain it." };
}
return null;
}
// lib/errors-tools.ts (continued)
/** Stop reasons that need handling (values from Anthropic's stop-reason docs). */
export function classifyStopReason(stopReason: string | null): ClassifiedError | null {
const base = { source: "anthropic" as const, detail: `stop_reason: ${stopReason}` };
switch (stopReason) {
case "pause_turn":
return { ...base, kind: "model", retry: { action: "continue" }, userMessage: "" };
case "max_tokens":
return { ...base, kind: "model", retry: { action: "none" }, userMessage: "The answer was cut off. Ask for a shorter answer." };
case "model_context_window_exceeded":
return { ...base, kind: "limits", retry: { action: "none" }, userMessage: "This conversation is too long. Start a new chat." };
case "refusal":
return { ...base, kind: "model", retry: { action: "none" }, userMessage: "The assistant can't help with that request." };
default:
return null; // end_turn, tool_use, stop_sequence
}
}
Step 4: retry only what is safe
withRetry is for calls the SDK doesn't already retry, such as a token refresh or your own REST calls. It refuses to retry anything that isn't marked idempotent, honors retry-after, and otherwise uses exponential backoff with full jitter.
// lib/retry.ts
import type { ClassifiedError } from "./errors";
export class ClassifiedFailure extends Error {
constructor(public readonly classified: ClassifiedError, options?: { cause?: unknown }) {
super(classified.detail, options);
this.name = "ClassifiedFailure";
}
}
export interface RetryOptions {
/** Only operations that are safe to repeat are ever retried. */
idempotent: boolean;
classify: (error: unknown) => ClassifiedError;
maxAttempts?: number; // including the first try
baseDelayMs?: number;
maxDelayMs?: number;
onRetry?: (classified: ClassifiedError, attempt: number, delayMs: number) => void;
sleep?: (ms: number) => Promise<void>;
random?: () => number;
}
export async function withRetry<T>(operation: (attempt: number) => Promise<T>, options: RetryOptions): Promise<T> {
const {
maxAttempts = 3,
baseDelayMs = 500,
maxDelayMs = 8_000,
sleep = (ms: number) => new Promise<void>((resolve) => setTimeout(resolve, ms)),
random = Math.random,
} = options;
for (let attempt = 1; ; attempt += 1) {
try {
return await operation(attempt);
} catch (error) {
const classified = options.classify(error);
const retry = classified.retry;
if (retry.action !== "backoff" || !options.idempotent || attempt >= maxAttempts) {
throw new ClassifiedFailure(classified, { cause: error });
}
// Honor retry-after when the server sent one; otherwise exponential backoff with full jitter.
const ceiling = Math.min(maxDelayMs, baseDelayMs * 2 ** (attempt - 1));
const delayMs = retry.retryAfterMs ?? Math.round(random() * ceiling);
options.onRetry?.(classified, attempt, delayMs);
await sleep(delayMs);
}
}
}
Step 5: check it
The checks build errors with the SDK's own APIError.generate, so they break if the SDK's shape drifts. Two of them:
// tests/errors.test.ts (excerpt)
// Real SDK error objects, built the way the SDK builds them from an HTTP response.
function sdkError(status: number, type: string, headers: Record<string, string> = {}) {
return Anthropic.APIError.generate(status, { type: "error", error: { type, message: `${type} message` } }, undefined, new Headers(headers));
}
await check("429 without retry-after is not retried (spend cap)", () => {
const c = classifyAnthropicError(sdkError(429, "rate_limit_error"));
assert.deepEqual(c.retry, { action: "none" });
});
await check("withRetry never retries a non-idempotent operation", async () => {
let calls = 0;
await assert.rejects(
withRetry(async () => { calls += 1; throw sdkError(529, "overloaded_error"); }, { idempotent: false, classify: classifyAnthropicError, sleep: async () => {} }),
(error: unknown) => error instanceof ClassifiedFailure && error.classified.kind === "transient",
);
assert.equal(calls, 1);
});
npx tsc --noEmit
npx tsc --outDir dist && node dist/tests/errors.test.js # "15 checks passed"
In the capstone, the natural home is the /api/chat route: pass thrown errors through classifyAnthropicError, send userMessage to the browser, and write buildTurnLog for every finished turn.
Step 6: write the runbook
A runbook is what the on-call person opens at 2 a.m. Keep it short and specific:
# Runbook: <service name>
Owner: <team> · On call: <rota link> · Last reviewed: <YYYY-MM-DD>
## Dependencies
| Dependency | Where to check | Owner |
| --- | --- | --- |
| Salesforce org <My Domain> | Setup → Company Information | <admin> |
| Hosted MCP servers <names, URLs> | Setup → MCP Servers | <admin> |
| External Client App <name> | Setup → External Client App Manager | <admin> |
| Claude API, model <id> | Claude Console | <developer> |
## Signals
- App turn log: <link>, filter by correlationId or Salesforce user ID
- Event Log File Browser → API Total Usage, API_CLIENT_CATEGORY = SALESFORCE_HOSTED_MCP
- Company Information → API Requests, Last 24 Hours
## Playbooks
### <symptom>
- Detect: <alert or report>
- Check: <the first three things to look at>
- Fix: <steps>
- Escalate: <who, and when>
## Kill switches
- <how to stop writes, remove a tool, deactivate a server, revoke tokens>
Acme's playbooks, in short:
| Symptom | Check first | Action |
|---|---|---|
| Many users asked to sign in again | auth problems in the turn log; recent app policy changes |
Confirm the policy; users sign in again; if tokens may be exposed, bulk revoke in OAuth Usage |
| "The assistant is busy" for minutes | 429 and 529 rates, with request IDs | Let backoff work; lower concurrency; a 429 without retry-after means checking the Claude Console's spend limits |
| The org nears its daily API limit | API Requests; API Total Usage by USER_NAME |
Find the heaviest user or loop; narrow queries; talk to Salesforce about limits |
| A tool disappeared or changed | The server's tool list in Setup; recent deploys | Roll back the tool configuration; pin tool lists in the app |
| A write nobody approved | The write audit; the record's audit trail | Remove the tool from the server, revoke that user's tokens, look for injected instructions in recently read records |
Acme's kill switches: remove Headless_Assistant_User from users to cut access to the Apex tools; deactivate a server in Setup (Salesforce Help says to select "an active server that isn't in use", so test how that behaves with live sessions); bulk revoke the app's tokens in OAuth Usage; and an app-level switch that rejects every approval.
What can go wrong
| Symptom | Cause | Fix |
|---|---|---|
| A short outage turns into a flood of 429s | Retries stacked in the SDK, your code and the browser | One retry layer per call |
| Two tasks after a timeout | A write replayed after an ambiguous failure | Idempotent write tools; read back before any retry |
| Customer data in your log platform | Logging prompts or tool output | Log names, IDs and counts only |
| The assistant breaks after a sandbox refresh | A local External Client App wasn't copied | Package the app or recreate it; update the sandbox URLs |
| API calls stop mid-afternoon | The daily quota ran out unnoticed | Alert on API requests, not on errors |
| An unexpected bill | Assuming hosted MCP is free in production | Budget with Salesforce's official billing wording; revisit when HPI multipliers are published |
Production security note
Tokens are the production risk that grows quietly. Keep refresh tokens server-side, shorten their lifetime, revoke them on sign-out, and review OAuth Usage as part of your access reviews. Resist the urge to "simplify" with a shared integration user: Salesforce calls a service account used across all sessions "an anti-pattern that customers should avoid."
Today's checklist
- Every MCP change moves sandbox to production, with URLs in configuration.
- Acme's approval matrix names an approver for servers, tools, clients, write tools, beta features and prompts.
- I have an inventory of MCP clients and External Client Apps, ready for agent registration.
- I know where to find MCP traffic: API Total Usage,
API_CLIENT_CATEGORY, API requests and OAuth Usage. - My app logs one structured line per turn, with no tokens or record data.
- Errors are classified into six kinds, with one retry layer per call.
- The runbook has playbooks and working kill switches.
Frequently asked questions
Does every MCP tool call count against the org's API limit?
Salesforce's wiki says they do: each invocation "counts as one or more API calls depending on the tool". In my org the counter barely moved (1 to 15 after dozens of calls), so measure API Requests, Last 24 Hours yourself before you plan around either answer.
Will Salesforce charge for hosted MCP usage?
Salesforce says hosted MCP servers are intended for customers with Flex Credits and usage may be billed. It is targeting November for a new model that meters registered agents' calls as Headless Platform Interactions, with multipliers still "TBA". Sandboxes, scratch orgs and Developer Edition orgs aren't metered that way.
Should my app retry a failed tool call?
Only if the call is safe to repeat. Reads can be retried with backoff; writes only when the tool is idempotent and the user approved exactly those arguments. Let the SDK handle its own retries rather than adding another layer.
What is agent registration?
A change Salesforce announced for November 2026: each registered agent gets its own identity and, optionally, a narrower permission set, instead of acting only as the user it assists. No Setup path or date has been published, so treat it as future and subject to change.
What's next
Everything is in place: hosted servers, an External Client App, custom tools, an architecture, and the operations around it. On Day 15, the final day, we build the complete headless Salesforce AI assistant in Next.js: sign-in with PKCE, Claude calling sobject-all and Acme's custom tools as the signed-in user, a live tool trace, and the approve-before-write step, deployed with the Salesforce CLI and hardened with today's checklist.
Sources
Verified against the sources below on September 30, 2026. Salesforce ships Headless 360 changes often: check the linked docs if a screen looks different.
- Security Best Practices (Salesforce Hosted MCP Servers)
- How to Secure Salesforce Hosted MCP Servers — Event Log File Browser, OAuth Usage
- Security and Governance in Headless Experiences
- Understand AIforce Impact — Salesforce knowledge article, September 17, 2026: agent registration and billing
- Salesforce Hosted MCP Servers: Get Started (billing considerations)
- Known Limitations (mcp-hosted wiki)
- Claude API errors
- MCP connector (Claude API)
Everything from today on one page. Tap to zoom, or download it for later.
All 15 days in this series
- Day 11Security Architecture: OAuth, Permissions & Least Privilege
- Day 12Custom MCP Tools + Apex/Flow Business Logic
- Day 13Headless 360 + Agentic AI Architecture
- Day 14Production Architecture: Governance, Monitoring & Error Handling




Comments
Loading comments...