Best Practices for Deploying LangChain Apps to Production
A practical checklist for running LangChain and LangGraph agents in production in 2026: where to host them now that LangServe is deprecated, how to keep conversation state per user with checkpointers, when an LLM cache helps, LangSmith tracing and evaluation, secrets, and the timeouts that break long agent runs.

An agent that works in a notebook has only proved that the prompt and tools can work once, for one person, on a fast connection. Production adds everything a notebook hides: many users at the same time, several server processes, slow or failing model providers, costs that grow with traffic, and bugs you have to diagnose from logs after the fact.
This guide is the checklist I would work through before putting a LangChain or LangGraph agent in front of real users. It is written against LangChain 1.x and the LangChain, LangSmith and AWS documentation as of October 3, 2026. There are no benchmarks here and no numbers I have not taken from the docs; where something depends on your workload, the advice is to measure it yourself.
1. Pick how you will serve the agent
The old answer was LangServe: wrap a chain with add_routes on a FastAPI app. The LangServe README now says it has been deprecated since November 18, 2024 and recommends LangGraph Platform for new projects. In the current docs that hosted offering is LangSmith Deployment, which runs your graph on LangSmith's Agent Server.
That leaves three realistic options:
- LangSmith Deployment (cloud). Managed infrastructure for stateful, long-running agents, with persistence and background runs. The docs also list hybrid, standalone server and self-hosted options if your data has to stay in your own cloud.
- Your own API around the agent. A FastAPI (or similar) service that calls
agent.invokeoragent.astream. More work, full control, fits into an existing backend with its own auth. - A queue and workers. For runs that take minutes, accept the request, put a job on a queue, and let workers run the agent and store the result. The client polls or receives a webhook.
For LangSmith Deployment, the project is described by a langgraph.json file that names your dependencies, the graphs to expose and the environment file. The application structure page shows this shape:
{
"dependencies": ["langchain_openai", "./your_package"],
"graphs": {
"my_agent": "./your_package/your_file.py:agent"
},
"env": "./.env"
}
The agent it points to can be the object returned by create_agent. Do not commit the .env file itself; see the secrets section below.
My rule of thumb: if the agent is the product and you do not already run a Python backend, start with LangSmith Deployment. If the agent is one feature inside an existing application, put it behind your own API so it shares that application's auth, logging and deployment pipeline.
2. Stream the response
Agent runs are slow compared with normal API calls: a model call, a tool call, another model call. Users tolerate that far better when they see progress. LangChain agents support stream and astream; with stream_mode="messages" you receive tokens as the model produces them, along with metadata about which step produced them, so the client can show text as it arrives and a status line while a tool runs.
On your own API, send these chunks as Server-Sent Events or over a WebSocket. Two practical points: make sure every proxy in the path (load balancer, CDN, API gateway) passes streaming responses through instead of buffering them, and decide what the client should see when a run fails halfway through a stream.
3. Keep conversation state outside the process
The most dangerous notebook habit is keeping memory in a module-level variable. In a web server that variable can be shared by every request in the process, so one user's conversation can bleed into another's, and it disappears on every restart or scale-down.
In LangChain 1.x, short-term memory is LangGraph state saved by a checkpointer and keyed by a thread_id you pass in the config. InMemorySaver is fine for development. For production, the short-term memory guide uses Postgres:
pip install -U langgraph-checkpoint-postgres "psycopg[binary]"
import os
from langchain.agents import create_agent
from langgraph.checkpoint.postgres import PostgresSaver
DB_URI = os.environ["CHECKPOINT_DB_URI"]
with PostgresSaver.from_conn_string(DB_URI) as checkpointer:
checkpointer.setup() # creates the checkpoint tables on first run
agent = create_agent(
"openai:gpt-5.5",
tools=[],
checkpointer=checkpointer,
)
config = {"configurable": {"thread_id": "user-123:conversation-7"}}
agent.invoke({"messages": [{"role": "user", "content": "Hello"}]}, config)
For an async server, the persistence docs also list AsyncPostgresSaver. A few rules I would follow:
- Build the
thread_idon the server from the authenticated user and conversation, never trust one sent by the client, and keep it under 255 characters (the persistence docs call out that limit for Postgres). - Long threads grow. Add
SummarizationMiddlewareor trimming so old turns are condensed before they fill the context window. - If you deploy on the Agent Server, the persistence docs say you do not need to configure checkpointers yourself; the server handles it.
Implementing memory in LangChain conversational AI goes deeper on what to keep, what to summarise and what to forget.
4. Add caching where requests really repeat
An LLM cache stores the response for a given prompt and model settings and returns it next time instead of calling the provider. The langchain-redis package provides RedisCache (exact match) and RedisSemanticCache (similar prompts, using embeddings):
pip install -U langchain-core langchain-redis redis
import os
from langchain_core.globals import set_llm_cache
from langchain_redis import RedisCache
set_llm_cache(RedisCache(redis_url=os.environ["REDIS_URL"], ttl=3600))
Be realistic about what this catches. An exact-match cache only hits when the whole prompt is identical, and in a multi-turn agent the prompt includes the conversation so far, so hits are rare. Caching pays off on stateless steps: classifying a ticket, summarising the same document, answering a fixed FAQ. A semantic cache can return an answer to a question that is similar but not the same, so keep it away from anything personal or account-specific, and set a TTL so stale answers expire. Measure your hit rate before deciding the cache is worth the extra moving part.
5. Set timeouts, retries and rate limits deliberately
Every layer has its own timeout, and the shortest one wins.
- The model call.
init_chat_modelacceptstimeout(seconds) andmax_retries(default 6). Retries on a long agent run multiply the worst-case time, so set both on purpose. LangChain also has an optional in-memory rate limiter you can pass to the model to stay under a provider's request limits. - The agent loop. Cap the number of model and tool steps so a confused agent cannot loop indefinitely, and put a timeout on every tool that calls an external system.
- The platform. On AWS, the API Gateway REST API integration timeout can be set anywhere from 50 milliseconds to 29 seconds. The quota table says it can be raised above 29 seconds for Regional and private APIs (not edge-optimized ones), and that doing so may require a reduction in your account's Region-level throttle quota. Lambda can stream responses through function URLs, the
InvokeWithResponseStreamAPI or an API Gateway proxy integration; the Lambda docs say native streaming is supported on Node.js managed runtimes, while Python needs a custom runtime or the Lambda Web Adapter.
My preference for agents that can run longer than a typical HTTP request is a container service with streaming, or the queue-and-worker pattern, rather than fighting gateway limits.
6. Trace everything from the first deploy
Normal application logs tell you that a request failed. They rarely tell you what the rendered prompt was, which tool the model chose, what arguments it passed or what came back. Tracing does. For LangSmith, the setup is environment variables:
export LANGSMITH_TRACING=true
export LANGSMITH_API_KEY="<your-langsmith-api-key>"
export LANGSMITH_PROJECT="support-agent-prod"
Two settings from the tracing docs are easy to miss. If your LangSmith account is in a region other than the US, also set LANGSMITH_ENDPOINT to your region's API URL, or the key will not authenticate. In serverless environments, set LANGCHAIN_CALLBACKS_BACKGROUND=false so traces finish sending before the function ends.
If you cannot send prompts to a third-party service, keep the same habit with your own tooling: log the inputs and outputs of every model and tool call with a run ID you can search on. Debugging LangChain with callbacks shows how to hook into each step yourself.
7. Evaluate before every change
Changing a prompt, a model or a tool description can quietly break behaviour that used to work. LangSmith splits evaluation into offline evaluation (run a curated dataset before you ship) and online evaluation (run evaluators on real production traffic). Evaluators can be code rules, an LLM acting as a judge, pairwise comparisons or human review.
A workable starting point:
- Collect real inputs that cover your main use cases and the failures you have already fixed, each with a reference answer or a rule for what "correct" means.
- Write checks in code wherever you can (the right tool was called, the JSON parses, a required field is present), and use an LLM judge only for things code cannot check.
- Run the set on every prompt, model or tool change, and compare the results with the previous version before deploying.
- Add production traces that went wrong to the dataset, so the same mistake cannot ship twice.
8. Handle secrets and data properly
- Keep provider keys, database URLs and the LangSmith key in your platform's secret store or a secrets manager, injected as environment variables at runtime. Never commit them, and never put them in prompts or tool outputs.
- Give tools the narrowest credentials that work. A tool that reads orders should not hold a key that can also issue refunds.
- Decide what may leave your system. Traces contain full prompts and outputs, so check what personal data they include and what your tracing provider's retention settings are.
- Validate tool arguments in code. The model choosing a tool is not authorisation to run it with any input.
A minimal production checklist
- Agent served through LangSmith Deployment or your own API, not LangServe.
- Streaming works end to end, through every proxy.
- Checkpointer backed by a real database;
thread_idbuilt on the server. - Timeouts and retry limits on the model, the tools and the agent loop.
- Cache only on stateless, repeatable steps, with a TTL.
- Tracing on from day one, with the right region and serverless settings.
- An evaluation set that runs before every change.
- Secrets in a secret store; least-privilege credentials for tools.
If you are still deciding whether you need LangChain at all, start with What is LangChain? When to use it and when to write your own code.
Frequently asked questions
Is LangServe still a good way to deploy LangChain apps?
Not for new projects. Its README says it has been deprecated since November 18, 2024 and recommends LangGraph Platform, now documented as LangSmith Deployment. For an existing LangServe app, plan a move to LangSmith Deployment or to your own FastAPI service.
Can I deploy a LangChain agent on AWS Lambda?
Yes, with care. API Gateway REST APIs default to a 29-second integration timeout, which can be raised for Regional and private APIs. Lambda response streaming is native on Node.js runtimes; Python needs a custom runtime or the Lambda Web Adapter. For long agent runs, a container service or a queue with workers is usually simpler.
How do I store conversation history in production?
Use a LangGraph checkpointer backed by a database, such as PostgresSaver from langgraph-checkpoint-postgres, and pass a thread_id per conversation in the config. On LangSmith's Agent Server, persistence is handled for you.
Does an LLM cache work for chat agents?
Mostly not for the conversational turns, because the prompt includes the conversation history and rarely repeats exactly. It works well for stateless steps such as classification or summarising the same document. langchain-redis provides both exact-match and semantic caches.
Which environment variables turn on LangSmith tracing?
LANGSMITH_TRACING=true and LANGSMITH_API_KEY, optionally LANGSMITH_PROJECT. Add LANGSMITH_ENDPOINT for non-US regions and LANGCHAIN_CALLBACKS_BACKGROUND=false in serverless functions.
Do I need LangSmith to run LangChain in production?
No. LangChain runs without it. You do need some form of tracing and evaluation, and LangSmith is the most direct option for LangChain apps. If data rules prevent it, log model and tool calls yourself and run your evaluation set in CI.
Sources
Verified against the sources below on October 3, 2026. Products and docs change often: check the linked sources if something looks different.
- LangServe README (deprecation notice)
- LangGraph: Deployment
- LangSmith: Application structure (langgraph.json)
- LangGraph: Persistence
- LangChain: Short-term memory
- LangChain: Models (timeouts, retries, rate limiting)
- LangChain: Streaming
- Redis cache integration for LangChain
- langchain-redis on PyPI
- LangSmith: Trace with LangChain
- LangSmith: Evaluation
- Amazon API Gateway: REST API quotas
- AWS Lambda: Response streaming



Comments
Loading comments...