Skill Doctor: Autonomous diagnostics for model tooling and observability
A diagnostic design for checking agent tool contracts, permissions, dependencies, and live availability before failures spread through long workflows.
Research status: System design and benchmark proposal. Any performance or reliability results will be published separately with task definitions, fixtures, and reproducible traces.
An agent can form a valid plan and still fail because the tool boundary changed. An SDK renames a field, an API removes an endpoint, a credential loses scope, or the local runtime no longer matches the documentation in context. The model often sees only the final error and may spend several turns changing a correct plan.
Skill Doctor is a diagnostic layer for that boundary. It checks what a tool claims to support, what the active environment can reach, and what happened during execution. The output is structured evidence that a planner can act on without granting the diagnostic system broad repair permissions.
Tool health is a versioned contract
Each tool registration should include more than a name and prose description. We define a contract with:
- input and output schemas;
- authentication and permission requirements;
- side-effect classification;
- timeout, retry, and rate-limit behavior;
- runtime and dependency versions;
- safe probe operations;
- known error categories; and
- the time and environment of the last successful check.
interface ToolHealthRecord {
toolId: string;
contractVersion: string;
environmentDigest: string;
checkedAt: string;
checks: Array<{
id: string;
status: 'pass' | 'fail' | 'unknown';
evidence: string[];
safeToRetry: boolean;
}>;
}
The record is evidence for one version in one environment. It should expire when the tool, credential scope, or runtime changes.
Diagnostic pipeline
The proposed architecture has four stages.
[Agent Planner] -> [Contract Gate] -> [Scoped Probe Runner] -> [Tool]
| |
v v
[Contract Diff] [Trace Store]
\ /
-> [Health Record] ->
1. Static contract audit
Before execution, the auditor compares the registered schema with the active client and available documentation. It checks required parameters, enum values, response shapes, dependency versions, credentials declared by name, and missing configuration.
Static analysis can identify a stale function signature without making a network request. It cannot confirm that the service or credential is active.
2. Scoped dynamic probes
The probe runner executes only operations marked non-destructive. Examples include reading service metadata, listing a bounded resource, validating a token without mutation, or running a fixture through a local adapter.
The runner uses dedicated low-privilege credentials, strict timeouts, and a separate rate budget. A diagnostic system should never convert “check whether deletion works” into a real deletion.
3. Runtime trace analysis
During a workflow, the diagnostic proxy records tool name, contract version, sanitized inputs, latency, retry count, response shape, error category, and correlation identifiers. Sensitive values are removed before storage.
The trace lets the system separate several failures that may produce similar model-visible text:
- invalid arguments generated by the planner;
- a stale adapter that serialized valid arguments incorrectly;
- authentication or permission failure;
- rate limiting or temporary unavailability;
- a response that violates the registered schema; and
- a downstream scientific or business rule rejecting the request.
4. Recovery recommendation
Skill Doctor returns a bounded recommendation, not an unrestricted code change. The planner may correct an argument, select a compatible tool version, wait under the retry policy, use an approved fallback, or ask for human intervention.
Changes to code, schemas, credentials, or external state remain subject to the normal review boundary. Automatic repair is appropriate only when the allowed transformation is narrow, reversible, and tested against fixtures.
Observability that agents can consume
Traditional dashboards are designed for operators. Agents need a compact health view that can enter planning context without replacing the underlying evidence.
A useful view answers:
- Is the tool available in this environment?
- Does the active interface match the registered contract?
- Which operations are permitted with the current credential?
- What failed most recently, and is a retry safe?
- Which fallback is approved for this operation?
- Which evidence supports the health status?
The status should include unknown. Absence of a recent failure is not proof that a tool works.
Benchmark design
We propose a fixture-based benchmark with controlled faults. Each task package starts from a working tool integration, then introduces one change at a time:
- renamed or removed parameter;
- incompatible response field;
- expired or under-scoped credential;
- missing runtime dependency;
- deterministic rate-limit response;
- temporary network failure; or
- unsafe operation that must not be probed.
The benchmark measures fault classification accuracy, unnecessary tool calls, unsafe actions attempted, recovery success, time and tokens spent, and false positives against healthy integrations. Results should be reported per fault class; a single aggregate score would hide the failures that matter most.
Security boundary
Diagnostics increase access to schemas, logs, and service state. The system must redact secrets, restrict probe credentials, cap response storage, and prevent untrusted tool output from rewriting its own contract. Health records should be signed or stored in an append-only audit stream when they influence high-impact automation.
Relationship to the research stack
Skill Doctor covers operational evidence at the tool boundary. Deterministic evaluations define how to compare agent versions under controlled tools. Recursive context defines how verified failures and repairs can inform later runs without becoming permanent folklore.
The next publishable result is not a claim that the system self-heals. It is a reproducible benchmark showing which tool failures it detects, which recoveries it performs, and where it stops for review.
Cite this research note
BibTeX entry for academic citations, literature trackers, and generative research synthesizers:
@article{kalaris2026_skill_doctor,
title = {Skill Doctor: Autonomous diagnostics for model tooling and observability},
author = {Chowdhury, Sayan},
journal = {Kalaris Labs Research Notes},
year = {2026},
url = {https://kalarislabs.com/research/skill-doctor}
}
