CommandCodeAI / CommandCodeAI/command-code
feat: Mod - Adding test harness to Mod Api to make it easier to write tests and have agent verify mod works
Dieses Issue hat noch niemand übernommen.
- Vorherrschende Sprache
- Keine Sprachdaten
- Sterne
- 4k
- Forks
- 350
- PR-Merge-Kennzahlen
- Keine gemergten PRs in 30 T.
Beschreibung
Summary
When writing and iterating on a mod, I find the the model will write a temp script which imports some of the mod functions to test the output and if it works properly, for example:
xecute Shell Command
Command Code needs to execute node --experimental-strip-types --no-warnings -e "
Promise.all([
import('./.commandcode/mods/erd.ts'),
import('./.commandcode/mods/erd/markdown.ts'),
import('./.commandcode/mods/erd/renderer.ts'),
import('./.commandcode/mods/erd/commands.ts'),
import('./.commandcode/mods/erd/tools.ts'),
]).then(() => { console.log('ALL MODULES LOAD OK'); }).catch(e => { console.error('FAIL', e); process.exit(1); });
" 2>&1 | head -20.
Press [ctrl+e] to explain this command
❯ 1. Yes
Some are more involved and check output not just if load.
Execute Shell Command
Command Code needs to execute node --experimental-strip-types --no-warnings -e "
const fs = require('fs');
Promise.all([
import('./.commandcode/mods/erd/parser.ts'),
import('./.commandcode/mods/erd/queries.ts'),
import('./.commandcode/mods/erd/markdown.ts'),
]).then(([p, q, md]) => {
const sql = fs.readFileSync('schema/schema.sql', 'utf8');
const schema = p.parseSchema(sql, 'sqlite');
console.log('listTables →', md.renderMarkdown(q.listTables(schema)).length, 'lines');
console.log('listIndexes →', md.renderMarkdown(q.listIndexes(schema)).length, 'lines');
console.log('listForeignKeys →', md.renderMarkdown(q.listForeignKeys(schema)).length, 'lines');
console.log('findColumn →', md.renderMarkdown(q.findColumn(schema, 'id')).length, 'lines');
console.log('schemaSummary →', md.renderMarkdown(q.schemaSummary(schema, 'sqlite')).length, 'lines');
console.log('describeTable →', md.renderMarkdown(q.describeTable(schema, 'users', 'sqlite')).length, 'lines');
console.log('ALL QUERY FUNCTIONS RENDER CLEANLY');
}).catch(e => { console.error('FAIL', e); process.exit(1); });
" 2>&1 | head -20.
Press [ctrl+e] to explain this command
❯ 1. Yes
2. Yes, don't ask again for this exact command in this project
3. No, tell Command Code what to do differently
Expected Behavior
Would be useful as mods get more complex an refactor to have a proper test harness or suite like vitest that can be run easily manually or within agent as part of th mod-builder skill.
Actual Behavior
Mode has to cobble together one off scripts and runs that have to be permitted each Tim since the command is unique often per run so hard to write permissions to allow.
Can see from example above the shell command on-line a script and not sure I'd want to permit blanket node --experimental-strip-types --no-warnings -e allow permissions here.
Steps to reproduce the issue
- build mod
- try to test it
Command Code Version
1.7.0
Operating System
macOS
Terminal/IDE
ghostty
Shell
zsh
Additional context
No response
Beitragsleitfaden
Für dieses Repository ist kein Beitragsleitfaden indexiert
Erste Schritte
- Lies das ganze Issue und danach den Beitragsleitfaden des Projekts.
- Schreib ins Issue, dass du es übernimmst — das erspart doppelte Arbeit.
- Forke das Repository und arbeite in einem Branch.
- Öffne einen Pull Request, der die Issue-Nummer nennt.
Rechercherichtung
Beginne mit der Überprüfung der Mod API und des th-mod-builder-Skills und vergleiche dann die Beispiel-Node-Befehle im Issue mit dem aktuellen Mod-Workflow. Definiere, was der Harness abdecken sollte, wie er manuell oder über den Agenten ausgeführt werden sollte und wie erfolgreiche Modul-Ladevorgänge und Ausgabenprüfungen aussehen; es sind keine spezifischen Implementierungsdateien oder Tests angegeben.
Vom Indexierungsmodell aus dem Issue-Text verfasst.
Bewertung
- Tech-Stack
- node.js, typescript
- Bereich
- cli, developer-experience, testing
- Issue-Typ
- Feature
- Schwierigkeit
- 5/5
- Geschätzter Aufwand
- Über eine Woche
- Aktivitätsstatus
- Ruhig
- Klarheit
- Muss geklärt werden
- Anfängerfreundlichkeit
- 35/100