anthropics / anthropics/claude-code

Supported isolated auth and observable skill-to-agent dispatch for native skill evaluations

Ouverte
#92,726 0 commentaires 0 réactions 0 personnes assignées Voir sur GitHub
area:permissions area:skills question
Langage dominant
Python
Étoiles
145k
Forks
23.1k
Métriques de merge des PR
Métriques de PR en attente

Description

We are preparing ten synthetic native Claude Code skill cases, each repeated
three times. They cover natural activation, near misses, a competing procedure,
multi-step action and authority boundaries. Candidate skill metadata permits
natural invocation; explicit invocation would be a different test.

We need a supported method to run these cases while preserving normal skill
discovery and the real generated Claude-to-Codex integration, with independently
verifiable filesystem confinement and authentication isolation. CLI tool
allowlists, strict MCP configuration and disabling session persistence are
useful settings, but none alone proves those other boundaries.

Please identify the supported interfaces for:

1. Authentication to a disposable evaluation session through an isolated
broker, with no copied credentials and no model/worker credential access.
2. Reading the actual installed skill identity and invocation event without
injecting the skill text into the test prompt or forcing explicit invocation.
3. Observing real generated Codex dispatch, its selected route and effective
permissions, and correlating it with the frozen work packet.
4. Retaining only invocation/tool-name/outcome metadata and independently
checked artifact hashes, without keeping raw transcripts or auth material.

We will first run one synthetic canary after these prerequisites are verified.
Thirty case descriptions, static metadata tests or disabled-tool refusal cases
will not be represented as thirty successful native observations.

Guide de contribution

Aucun guide de contribution indexé pour ce dépôt

Piste de recherche

No files, tests, or entry points are named. Begin by locating the supported evaluation-session, skill-discovery, dispatch-observation, and metadata-retention interfaces, then verify whether they provide the requested isolation and correlation guarantees. Done means one synthetic canary produces independently checked metadata without raw transcripts or authentication material.

Rédigé par le modèle d'indexation à partir du texte de l'issue.

Évaluation

Stack technique
python
Domaine
cli, security, testing
Type d'issue
Fonctionnalité
Difficulté
5/5
Temps estimé
Plus d'une semaine
Activité
Active
Clarté
À clarifier
Accessibilité débutants
25/100

Recevez les nouvelles issues par e-mail

Un résumé court des issues GitHub adaptées aux débutants.