# LangSmith Tracing for Production AI Agents

# LangSmith Tracing for Production AI Agents: 15 Things Every AI Engineer Should Know

Building an AI agent is only half the job.

In development, we usually see:

```text
User → Agent → LLM → Response
```

But in production, an agent can execute multiple LLM calls, tools, retrieval steps, retries, and validations.

When something goes wrong, simply seeing the final answer is not enough.

This is where **LangSmith tracing** becomes useful.

* * *

# 1\. What Is LangSmith Tracing?

LangSmith tracing records the execution of your AI application so you can inspect what happened during a request.

Instead of:

```text
Request → Response
```

you can see:

```text
Request
   ↓
Agent
   ├── Tool
   ├── LLM
   ├── Retrieval
   └── Final Response
```

This makes an AI agent much easier to debug.

* * *

# 2\. Why Do Production Agents Need Tracing?

Imagine a user reports:

> "The agent gave me the wrong answer."

Without tracing, you may not know whether the problem came from:

*   the prompt
    
*   the LLM
    
*   a tool
    
*   retrieval
    
*   incorrect tool arguments
    
*   an agent decision
    
*   a retry
    
*   application logic
    

Tracing lets you inspect the execution path.

* * *

# 3\. Trace vs Run

A simple mental model:

```text
Trace
│
└── Agent Run
     ├── Tool Run
     ├── LLM Run
     └── Validation Run
```

**Trace** = the complete execution of a request.

**Run** = an individual operation inside that execution.

This parent-child structure is especially useful for multi-step agents.

* * *

# 4\. Install LangSmith

Create a Python project:

```bash
mkdir langsmith-tracing-demo
cd langsmith-tracing-demo

python -m venv .venv
source .venv/bin/activate

pip install langsmith python-dotenv
```

* * *

# 5\. Configure LangSmith

Create a `.env` file:

```env
LANGSMITH_TRACING=true
LANGSMITH_API_KEY=your_api_key
LANGSMITH_PROJECT=production-agent-demo
```

Never commit your real API key.

Add this to `.gitignore`:

```gitignore
.env
.venv/
__pycache__/
```

* * *

# 6\. The Simplest Trace

LangSmith provides the `@traceable` decorator for instrumenting functions.

```python
from langsmith import traceable


@traceable
def hello_agent(user_input: str):
    return f"Agent received: {user_input}"


print(hello_agent("Explain LangGraph"))
```

The important part is:

```python
@traceable
```

Now the function execution can appear as a trace in LangSmith.

* * *

# 7\. Trace a Tool

Agents frequently use tools.

Let's create a simple knowledge-search tool:

```python
from langsmith import traceable


@traceable(
    name="search_knowledge",
    run_type="tool"
)
def search_knowledge(query: str) -> str:

    knowledge = {
        "langgraph":
            "LangGraph is a framework for building stateful agent workflows.",

        "langsmith":
            "LangSmith provides observability for AI applications.",

        "rag":
            "RAG combines retrieval with generation."
    }

    query = query.lower()

    for key, value in knowledge.items():
        if key in query:
            return value

    return "No relevant information found."
```

Now the tool becomes visible inside the trace.

* * *

# 8\. Trace the Model Call

We can keep the application model-agnostic.

```python
@traceable(
    name="model_call",
    run_type="llm"
)
def model_call(prompt: str) -> str:

    # Connect your preferred model here.
    # OpenAI, Gemini, Ollama, Hugging Face, etc.

    return f"Model response for: {prompt}"
```

The tracing architecture doesn't have to depend on a particular model provider.

* * *

# 9\. Trace the Complete Agent

Now combine the tool and model:

```python
from langsmith import traceable


@traceable(
    name="agent",
    run_type="chain"
)
def agent(user_input: str):

    context = search_knowledge(user_input)

    prompt = f"""
User question:
{user_input}

Retrieved context:
{context}

Answer the user using the context.
"""

    response = model_call(prompt)

    return response


if __name__ == "__main__":
    result = agent("What is LangGraph?")
    print(result)
```

LangSmith can represent this execution as:

```text
agent
│
├── search_knowledge
│
└── model_call
```

This is the foundation of agent tracing.

* * *

# 10\. Add Metadata

Metadata gives additional context about a run.

```python
@traceable(
    name="production_agent",
    run_type="chain",
    metadata={
        "environment": "production",
        "agent_version": "1.0.0"
    }
)
def agent(user_input: str):
    ...
```

This becomes useful when debugging different deployments.

For example:

```text
environment = production
agent_version = 1.0.0
```

versus:

```text
environment = production
agent_version = 2.0.0
```

* * *

# 11\. Add Tags

Tags help categorize traces.

```python
@traceable(
    name="production_agent",
    run_type="chain",
    tags=[
        "production",
        "customer-support",
        "agent-v2"
    ]
)
def agent(user_input: str):
    ...
```

Think of it like:

```text
Metadata → information about the run

Tags → labels used to categorize runs
```

* * *

# 12\. Debug Production Errors

Suppose the production API returns:

```text
500 Internal Server Error
```

Your normal application log might only show:

```text
Request failed
```

A trace can expose:

```text
Agent
│
├── Search Tool
│      └── ERROR
│
└── Model
```

Now you know where to start investigating.

Instead of guessing, you can inspect the failing operation.

* * *

# 13\. Find Performance Bottlenecks

Imagine an agent takes 8 seconds.

The trace might show:

```text
Agent              8.0s
│
├── Retrieval       0.2s
├── Tool            0.3s
├── LLM #1          2.8s
├── LLM #2          4.5s
└── Finalization    0.2s
```

Now you have evidence that most of the latency comes from the model calls.

You can then investigate:

*   unnecessary LLM calls
    
*   prompt size
    
*   model selection
    
*   sequential execution
    
*   agent loops
    

The principle is:

> **Measure first. Optimize second.**

* * *

# 14\. Trace Complex Agent Workflows

A production agent might look like:

```text
                 User
                  │
                  ▼
               Agent
                  │
        ┌─────────┼─────────┐
        ▼         ▼         ▼
     Planner     Tool    Retriever
        │         │         │
        └─────────┼─────────┘
                  ▼
                 LLM
                  │
                  ▼
              Validator
                  │
                  ▼
             Final Answer
```

Without tracing, this execution can become difficult to understand.

With tracing, we can inspect the individual operations and their relationships.

This becomes especially useful with:

*   LangGraph
    
*   RAG
    
*   tool-calling agents
    
*   multi-agent systems
    
*   retries
    
*   conditional workflows
    

* * *

# 15\. The Production Mindset

The biggest lesson is simple.

A production AI engineer shouldn't only ask:

> **"Does my agent work?"**

They should also ask:

> **"Can I understand what my agent did?"**

A useful production loop is:

```text
User Request
     ↓
Agent Execution
     ↓
LangSmith Trace
     ↓
Inspect
     ↓
Find Root Cause
     ↓
Fix
     ↓
Run Again
     ↓
Compare
```

That is the real value of tracing.

* * *

# Complete Example

Here is the complete minimal example in one file:

```python
from langsmith import traceable


@traceable(
    name="search_knowledge",
    run_type="tool"
)
def search_knowledge(query: str) -> str:

    knowledge = {
        "langgraph":
            "LangGraph is a framework for building stateful agent workflows.",

        "langsmith":
            "LangSmith provides observability for AI applications.",

        "rag":
            "RAG combines retrieval with generation."
    }

    query = query.lower()

    for key, value in knowledge.items():
        if key in query:
            return value

    return "No relevant information found."


@traceable(
    name="model_call",
    run_type="llm"
)
def model_call(prompt: str) -> str:

    # Replace this with your preferred LLM.
    return f"Model response for: {prompt}"


@traceable(
    name="production_agent",
    run_type="chain",
    tags=["agent", "demo"],
    metadata={
        "environment": "development",
        "agent_version": "1.0.0"
    }
)
def agent(user_input: str) -> str:

    context = search_knowledge(user_input)

    prompt = f"""
User question:
{user_input}

Retrieved context:
{context}

Answer the user using the available context.
"""

    return model_call(prompt)


if __name__ == "__main__":

    result = agent(
        "What is LangGraph?"
    )

    print(result)
```

The resulting execution can conceptually be viewed as:

```text
Trace
│
└── production_agent
      │
      ├── search_knowledge
      │
      └── model_call
```

* * *

# Final Takeaway

An AI agent without observability can become a black box.

LangSmith tracing helps turn:

```text
"Something went wrong."
```

into:

```text
"Retrieval returned poor context,
which caused the downstream model
to generate an incorrect response."
```

That difference matters in production.

**Build the agent.**

**Trace the agent.**

**Understand the agent.**

**Then improve the agent.**

* * *

## What I Learned

For me, the most important shift is:

```text
Development mindset:
"Does it work?"

Production mindset:
"Can I observe, debug and understand it?"
```

That is where tracing becomes an important part of building reliable AI agents.
