Intercept and inspect Coding Agent API traffic from Claude Code, Codex CLI, Gemini CLI, Cursor CLI, OpenCode, Kimi/Kimi Code, Pi, and Hermes in a local trace viewer.
-
Updated
Sep 22, 2026 - Python
Intercept and inspect Coding Agent API traffic from Claude Code, Codex CLI, Gemini CLI, Cursor CLI, OpenCode, Kimi/Kimi Code, Pi, and Hermes in a local trace viewer.
OpenTelemetry-native SDK for AI agent observability, tracing, evaluation, debugging, governance, and runtime policy enforcement. Framework-agnostic and built for OpenAI Agents, LangGraph, CrewAI, and LLM applications in production.
Audit trail database for AI agents. Tamper-evident, checksum-backed decision logs with session replay and time-travel debugging to support EU AI Act Article 12 record-keeping. Drift detection, MCP, Python & TypeScript SDKs. Self-host or cloud.
【EMNLP 2026 Demo】A debugging framework for agentic AI systems: diagnose failures, attribute root causes, recover with evidence, and validate fixes through reruns.
Local-first runtime for observable, debuggable AI-agent workflows with slidable autonomy and multi-agent/tool orchestration.
Local open-source dev tool to debug, secure, and evaluate LLM agents. Provides static analysis, dynamic security checks, and runtime monitoring - integrates with Cursor and Claude Code.
Root cause analysis for AI agents. Detects agent loops, retry storms, and optimization opportunities in LangSmith, Langfuse, Arize Phoenix, and OpenTelemetry traces.
Local replay debugger for Browser Use failures with screenshots, model I/O, failed-step timelines, and public-safe HTML exports.
A git-diff-style debugger for agent memory.
Local-first agent debugging for the "it worked yesterday" moments. Track AI costs, replay runs, and catch non-deterministic failures.
Explain why your agent failed — root-cause debugging, memory attribution, and run divergence for LLM agents.
A real-time observability and debugging layer for AI agents.
AI agents fail like junior teammates, looping on bad ideas, ignoring feedback, and escalating commitment. vstack ports 34 of the most-cited organizational-behavior frameworks so you can diagnose your agents the same way you'd diagnose your team.
ChainWatch is a flight data recorder for multi-step AI systems. It's a CLI-based tool that records every step in an AI decision chain, links them together in order, prevents tampering, and allows you to verify the chain's integrity and replay the full decision flow.
A truth-first visual observatory for understanding, replaying, comparing, and monitoring AI agent runs.
Freeze, rewind, and attribute AI agent failures with replayable public traces.
Failure attribution for agent pipelines — find which span caused the failure and what kind of fix it needs.
Time-Travel Debugger for AI agents — record agent executions, inspect every step, rewind state, and fork/replay execution paths with cached tool outputs through an interactive React Flow debugger.
Developer platform for tracing, visualizing, and debugging AI agent executions with evidence-linked investigations across model, tool, retrieval, and workflow spans.
To associate your repository with the agent-debugging topic, visit your repo's landing page and select "manage topics."