Freestructured workflowmedium risk

Capability · rank 15

agent-observability-reviewer

Find missing traces, weak evaluations, and operational blind spots in agent systems.

Reviews agent observability: traces, tool calls, eval hooks, cost/latency signals, and incident readiness.

Listed owner

skillforge/ai-engineering

Package source

Kairaxis manifest

Category

AI Engineering

Package validatedCreator reviewedAI Engineering
$npx skills add skillforge/ai-engineering --skill agent-observability-reviewer

Overview

Reviews agent observability: traces, tool calls, eval hooks, cost/latency signals, and incident readiness. Review the sandbox comparison, limitations, permissions, and pinned version before install.

Outcome

Find missing traces, weak evaluations, and operational blind spots in agent systems.

Who it is for

Platform teams operating multi-step agents in production.

What it is not for

Generic APM setup with no agent-specific workflows.

Runtime export

codex_skill · markdown_export

Sandbox comparison
Generic baseline vs agent-observability-reviewer on the same sample input · pinned v0.9.1

Sample input

Agent has request logs but no tool-call spans or cost attribution.

Model: gpt-4.1-mini · temp 0.2

Generic baseline

42/100
Add more logging and a dashboard.

Skill-enhanced

89/100
Gaps: no tool spans, no token/cost labels, no eval sampling, no failure taxonomy.
Plan: structured traces per step, redaction policy, cost dimensions, weekly eval slice, on-call runbook links.

Rubric scores

+47 lift
Coverage of failure modes25
Actionable recommendations35
Measurable next steps24
Risk / limitation awareness24

Assumptions

  • Sample input is representative of a real production task
  • Same model and temperature for baseline and skill-enhanced runs

Comparison runs through the shared sandbox engine on registry sample cases (deterministic v0 path).

Usage

  1. 1. Inspect trust badges, limitations, and permissions on this page.
  2. 2. Run the sandbox comparison to see baseline vs skill-enhanced output.
  3. 3. Install the pinned version with the command below for your agent runtime.
  4. 4. Agents can also fetch /api/v1/skills/agent-observability-reviewer/manifest.
$npx skills add skillforge/ai-engineering --skill agent-observability-reviewer

Example prompts & use cases

1 / Prompt

Review our agent tracing plan before launch

2 / Prompt

Identify observability gaps in this tool-calling workflow

Limitations

  • Cannot access private telemetry unless provided in the sandbox input.

Audience

Platform teams operating multi-step agents in production.

Not for

Generic APM setup with no agent-specific workflows.

Version history

v0.9.1

Jun 25, 2026current

Beta reviewer pack