Amazon CloudWatch Omni Review: Building Traces for AI Agents in VS Code

Amazon CloudWatch Omni brings agentic AI observability directly to VS Code. Learn how to set up traces, evaluators, and monitoring for AI agents.

A software developer sits at a wooden desk in a bright modern office, focused on a laptop screen displaying a code editor, with a cup of coffee nearby

Introduction

If you have ever shipped an AI agent and then tried to figure out why it made a bad decision in production, you already know that standard monitoring falls apart pretty quickly. Amazon CloudWatch Omni is Amazon’s latest attempt to fix that gap, and it arrives at a moment when most teams are still winging it with logs and hope.

The core problem is simple enough. Traditional applications follow predictable patterns. A user hits an endpoint, the server processes a request, and the response comes back. You can set alerts on latency, error rates, and throughput because the behavior is deterministic. Agents do not work that way. One prompt tweak can quietly degrade response quality while your existing dashboards show nothing wrong. A tool gets called twice instead of once. The reasoning path diverges somewhere in the middle, and by the time you spot it, you have already lost three hours digging through scattered logs across multiple systems.

This is where CloudWatch Omni tries to do something different. Instead of another browser-based dashboard that forces you to context-switch away from your actual work, it lives inside VS Code and Kiro. You get traces, evaluators, and a playground right where you are already writing code. There is also a standalone web experience for operators who need fleet-level visibility, but the whole point is that the data stays consistent between the developer’s IDE and the operator’s screen.

I have spent enough nights manually correlating agent behavior across separate logging platforms to appreciate the direction here. The built-in evaluators for correctness, coherence, retrieval quality, and tool selection address exactly the kind of questions that normally require writing custom scripts or building your own evaluation pipeline. You can compare prompts side by side, build test datasets from real traffic, and catch regressions before they reach users.

What I do not know yet is how this holds up when you are running dozens of agents across multiple providers with complex tool configurations. The initial setup looks straightforward, but real-world production workloads tend to expose the rough edges that demo environments hide. That said, the local-first approach is sensible. You can develop entirely offline and only push traces to CloudWatch when you are ready, which removes the friction that normally kills observability adoption in the first place.

The announcement is worth a closer look for anyone deploying agentic AI workloads, even if you end up using only a fraction of what it offers. Amazon’s official blog post covers the full feature set. (opens in new tab)

Background: The Observability Gap for AI Agents

Traditional application monitoring was built for systems that behave predictably. You send a request, you get a response. Latency goes up, you check the database. Error rate spikes, you look at the service logs. The mental model is clean because the behavior underneath is clean.

AI agents do not work that way. A single user prompt can trigger a cascade of tool calls, intermediate reasoning steps, and branching decisions before the agent returns an answer. One invocation might call the weather API and compose a quick reply. Another invocation with nearly identical input might query a database, hallucinate a result, call a second tool to correct itself, and still produce the wrong output. The same input does not produce the same output, and that non-determinism is what breaks standard monitoring approaches.

Most teams I have talked to are still running a manual log-review workflow after incidents. Someone notices degraded response quality in production, then spends hours flipping between CloudWatch Logs, a vector database dashboard, and whatever LLM provider’s tooling they happen to be using. They are looking for the moment the agent went off track, but there is no single view that connects the prompt, the tool selection, the retrieval results, and the final answer. You end up guessing. The guess is usually wrong.

OpenTelemetry became the connective tissue that makes better observability possible for these workloads. It standardizes how traces flow from your code into any backend that understands the format. An agent span captures the LLM call. Child spans capture each tool invocation. Metadata tags the prompt version, the provider, the model parameters. Without OpenTelemetry, every team would need to build their own instrumentation plumbing from scratch. With it, the trace data exists in a structure that observability tools can actually parse and display meaningfully.

The gap is not that the data does not exist. The gap is that the tools to consume it have not kept up. Traditional APM systems visualize request-response flows with clean waterfall diagrams. They are not designed to surface why an agent chose tool A instead of tool B, or whether the retrieved context was relevant enough to produce a correct answer. That requires evaluation logic baked into the observability layer itself, not bolted on afterward. This is exactly the space CloudWatch Omni is trying to fill.

What’s Happening Now with Amazon CloudWatch Omni

Amazon CloudWatch Omni is an observability, evaluation, and experimentation tool built specifically for AI agents and generative AI workloads. It lives off-console, meaning it is not another tab inside the AWS Management Console. Instead, it runs as a VS Code extension (and supports Kiro as well) so traces appear where you are already writing code.

There are two surfaces. Developers get the IDE extension where they run agents, view traces, and experiment with prompts. Operators get a standalone web experience accessible through SSO, no AWS console access required. Both surfaces share the same trace data. The trace a developer debugs locally is the same trace an operator investigates when something goes wrong in production. That shared data model matters because it eliminates the disconnect where a bug reproduces in dev but not in the environment ops can see.

The built-in evaluators are what separate this from standard tracing tools. CloudWatch Omni evaluates traces for correctness, coherence, retrieval quality, and tool selection. If your agent picked the wrong tool or retrieved irrelevant context, the evaluator flags it. There is also a playground for comparing prompt versions side by side. You can build test datasets from production traffic, run experiments across configurations, and detect regressions automatically.

Cloud Login is the bridge between local development and production monitoring. It connects your IDE to your AWS account so local traces get sent to Amazon CloudWatch for persistent storage. You can share traces with your team and access production dashboards through the same data pipeline. This connection is optional. During development, everything stays local on your machine.

The local-first approach is worth pausing on. Most observability tools assume you are already in production and want to ship data somewhere. CloudWatch Omni inverts that. You start local, iterate on your agent, and only connect to the cloud when you are ready to monitor production agents. According to the announcement post (opens in new tab), this design is intentional because the eval-driven workflow should begin during development, not after deployment.

What It Means in Practice

I walked through the VS Code extension setup yesterday. After installing the extension from the marketplace, the Omni icon appeared in my Activity Bar. I selected Get started with Sample Project from the welcome screen. That loaded a pre-configured agent with sample trace data so I could see the format before touching my own code.

The interactive chat walkthrough asked me to define my agent’s purpose, select a model provider, and configure tools. It felt like talking to someone who actually knows what they are doing instead of reading a forty-page manual. The tool then used an AI code assistant to handle the boilerplate. It configured the Dev Server, installed dependencies, and set up instrumentation automatically. I went from clicking Install to seeing my first traced agent session in about twelve minutes. Not twelve hours. Twelve minutes.

When I sent a question to the agent, selecting View Trace revealed exactly how it processed the request. The Trace Explorer showed a structured timeline of every step. LLM calls, tool invocations, reasoning steps, all laid out hierarchically. I clicked into individual spans and inspected inputs, outputs, token usage, and latency. Most tracing tools show you that the request failed. This actually showed me which decision point went sideways.

Everything stayed on my machine during development. No cloud connection, no AWS credentials sitting in my environment variables, no telemetry shipping to anyone until I chose to connect. The local-first workflow means I can iterate on prompts, test tool configurations, and break things without affecting a shared dataset or hitting API limits on a staging account.

The setup process is smooth until it is not. I do not know how CloudWatch Omni holds up across large agent fleets with dozens of moving parts. I also have not tested it with multi-provider configurations beyond the initial walkthrough. That gap matters if your team ships agents across Claude, Gemini, and Bedrock models simultaneously.

When production incidents happen, having traces visible in VS Code alongside your code removes one context switch. Operators get a separate web experience connected through SSO without touching the AWS Management Console. The developer trace and the operator trace are the same data. That alignment usually does not exist in observability tooling.

What to Expect Next

AWS has positioned CloudWatch Omni as app-centric and AI-powered from day one. That framing suggests the roadmap will push harder in both directions. I expect deeper integrations with existing AWS services rather than building a separate monitoring silo. Lambda invocation traces flowing directly into agent spans would remove another manual instrumentation step. ECS and EKS support that mirrors what we already have for standalone agent deployments would close a gap most teams hit within weeks of production use.

Framework support is the other obvious expansion vector. Right now the walkthrough covers the most common generative AI patterns. As the ecosystem settles around specific frameworks, I expect Amazon to prioritize those. Bedrock Agents integration is the low-hanging fruit here. Teams deploying through that service currently route observability through a different set of tools. Bringing those traces into the same Omni workspace would reduce fragmentation without requiring a config overhaul.

The evaluator suite will likely grow beyond the current set. Correctness, coherence, retrieval quality, and tool selection cover the basics. But evaluation as a discipline is moving fast. I expect things like cost-per-token tracking, latency budget enforcement, and safety-related checks to join the existing list. The playground for side-by-side prompt comparison is already useful for small experiments. As that feature matures, I think it will absorb more of the experimentation workflow that currently lives in external tools.

The biggest unknown is where competitive pressure lands. Anthropic, OpenAI, and Google all have their own observability plays in various states of completion. IDE-native tooling is not yet table stakes in this space. If two or more major providers ship comparable extensions, CloudWatch Omni’s positioning could shift from differentiated to expected. That would not be a bad outcome for customers. It would mean the bar has been raised across the board.

I also expect the eval-driven workflow to mature as the agentic AI space normalizes. Right now evaluations feel like a nice-to-have layer on top of standard tracing. In twelve months, I think teams will treat them as the primary debugging surface for non-deterministic systems. The tool selection in this release is a signal that Amazon is betting on that trajectory. Whether the bet pays off depends on how quickly the evaluation feature set expands and how well it handles the messy edge cases that appear once you ship real agents to real users.

One thing I do not expect is a complete abandonment of the traditional CloudWatch surface. The standalone web experience coexists alongside the AWS Management Console rather than replacing it. Operators who already maintain dashboards, alarms, and log groups in the console are not going to migrate everything overnight. Omni fills a gap in that workflow instead of trying to displace it.

Quick Answers About CloudWatch Omni

Do I need the AWS Management Console to use CloudWatch Omni?

No. The standalone web experience connects through SSO and does not require console access. You can also run entirely locally during development. This matters because not every team has console permissions, and shipping agents in a constrained environment should not require a full AWS IAM delegation to get observability.

How is CloudWatch Omni different from regular CloudWatch logs and metrics?

Regular CloudWatch monitors deterministic request-response flows. A Lambda invocation starts, runs, and completes. You set thresholds, you get alarms. CloudWatch Omni is built for non-deterministic agent behavior, where the same prompt can produce five different traces depending on tool selection, model sampling, or retrieval results. It captures multi-step traces with built-in evaluators for correctness and tool selection, plus a playground for prompt experimentation. I think of it the way a site owner thinks about the difference between watching a server’s CPU graph and actually reading the request-level debug log after a deployment breaks something subtle.

Can I use CloudWatch Omni without connecting to AWS?

Yes. The local-first workflow stores all data locally. You only connect via Cloud Login when you want to share traces or access production dashboards. This was a deliberate design choice. Most teams I work with start experimenting with agents in isolation before handing them to ops. Forcing a cloud connection at setup would have added friction that probably would have kept this tool on the shelf for a lot of developers. The optional connection model means you can do the hard work of getting instrumentation right locally, then promote the same trace data upstream when you are ready.

All three questions point to the same underlying principle: Amazon is treating this as a developer-first tool that scales up rather than an operator-first dashboard with a developer plugin bolted on. Whether that orientation holds once teams ship larger agent fleets is the question that will actually matter.

Next Step

Install the CloudWatch Omni extension from the VS Code Marketplace, then pick up the flow right where the welcome screen drops you. I recommend starting with the Sample Project option. It loads a pre-configured agent with example trace data so you can see what the interface actually looks like before you touch your own code. The shortcut through the Command Palette works just as well if you prefer skipping clicks. Select Omni: Create a new Project and the interactive walkthrough takes over from there.

The chat-based setup asks you to define your agent purpose, pick a model provider, and wire up your tools. CloudWatch Omni then uses an AI code assistant like Kiro, Claude Code, or Codex to handle the boring infrastructure work. It configures the dev server, installs dependencies, and sets up the OpenTelemetry instrumentation for you. I went from extension installation to a traced agent session in about twelve minutes. That is fast enough that the friction is real measurement, not vague optimism.

If you already have an agent running, skip the sample project. Create a new project from scratch using the Omni chat to define your provider and tools, then point it at your existing codebase. The key moment comes after you send a question to the agent and select View Trace. The Trace Explorer breaks down every LLM call, tool invocation, and reasoning step in a structured hierarchical timeline. You can drill into any span to inspect inputs, outputs, token usage, and latency. Compare mode puts two traces side by side, which is where you catch regressions before they hit production.

Use the playground to modify a prompt, run it again, and compare the resulting trace against the baseline. This is the loop that matters. You change one thing, watch the trace, evaluate the result, and repeat. The built-in evaluators give you signal on correctness and tool selection so you are not guessing whether a response got worse or better. Once your local instrumentation is stable and you are comfortable with the trace data, flip the Cloud Login connection to send traces to Amazon CloudWatch for persistent shared storage. Until then you can stay entirely local and avoid any cloud dependency.