Repeated cost-control instruction failures after corrective plan; 10,750.48-credit account decrease
Nobody has claimed this yet.
- Dominant language
- Rust
- Stars
- 125k
- Forks
- 19.4k
- PR merge metrics
- PR metrics pending
Description
Feedback reference and environment
- App feedback submitted by the user; supplied session reference:
01a08389-23ef-7880-97af-8f64e25eef20. - Codex desktop on Windows with PowerShell, operating a delegated MakeCode Arcade tutorial-development workflow.
- Subscription: Pro 20× usage, followed by credits after weekly allowance exhaustion.
- Recorded models included
gpt-6-astrafor coordination/directors andgpt-5.6-solfor the continuing learner and runtime support. Exact desktop build is not recorded in this report.
This is an assistant-prepared incident summary based on the user's conversation and retained accounting/diagnostic records. It requests investigation of instruction-following and cost control, not a finding of intentional misconduct or a predetermined financial remedy.
Problem
The user repeatedly required economical execution, stage-level cost awareness without per-prompt reporting, preservation of learner progress, and prevention of known operator mistakes. After an earlier forensic review identified avoidable coordination and retry overhead, an owner-approved learner rework explicitly required stopping equivalent retries that produced neither coverage nor repair.
Despite that correction, expensive operational loops continued and a previously diagnosed input-timing failure recurred. The tutorial remained unfinished. The recorded account credit balance fell from 24,473.333035 to 13,722.851162: a decrease of 10,750.481873 credits in approximately ten hours on September 16, 2026. Of that decrease, 8,093.523480 credits occurred after acceptance of the revised shared method.
These are observed account balance changes, not an estimate of exact wasted credits. Token counters identify the dominant workload but do not establish per-task billed charges. The user's authorization to work explicitly included economic and known-error constraints; it was not authorization to consume the remaining balance without enforcing them.
Concrete recurrence after correction
- On September 15, a retained diagnosis established that the game's alternating petting gesture resets after 1,800 ms between strokes. Separate tool/model operations actually delivered strokes 3,250–4,966 ms apart. Requested short waits did not establish actual input timing.
- The later operator guide expressly required keeping a timed input sequence inside one bounded operation.
- On September 16, the continuing learner again delivered separate commands with actual gaps of 3,085 / 3,078 / 3,124 ms, despite requesting 500 ms waits. It also stood beside the wrong creature.
- Correct positioning and one composite gesture delivered approximately 413 / 416 / 423 ms gaps and succeeded on the same project. Before/after Blocks and TypeScript were byte-identical. No game-code repair was justified.
This is a specific previously understood operator failure followed by additional work to establish that unchanged code worked, not an inference that all testing cost was waste.
Earlier measured context
Retained structured tracking begins September 13. Recorded account weekly usage increased through 39%, 42%, 46%, 50%, 52%, 54%, 56%, 60%, 65%, 69%, 71%, 74%, and 77% while contemporaneous stage reports repeatedly identified avoidable waiting analysis, substantial coordination overhead, and unresolved efficiency problems. The user stated this was the only consuming project for the relevant interval. The ensuing forensic review explicitly concluded: “The system measured overruns but did not change its behavior.”
Further completion-stage counters were retained through September 15, followed by the approved rework and the September 16 credit decrease above. These selected records are not a gap-free billing ledger; allowance percentages, token estimates and observed credits have not been added into a synthetic total.
Expected behavior and requested investigation
Explicit corrective instructions should govern subsequent execution: reuse established working methods, retain useful progress, and change method or stop when known costly failures recur. Activity alone should not establish that autonomous work remains economical.
Please investigate why the acknowledged constraints failed to govern execution and whether model behavior, delegation, context/compaction, tool routing or workflow controls contributed. Those causes are not isolated by the local evidence.
Production is paused and automatic continuation has been deleted. Original evidence remains preserved locally. A redacted copy of the detailed report and all 17 selected evidence records is now attached below. Inclusion of the files in the user's app feedback remains unconfirmed. The feedback reference above is provided for correlation; it is not a claim that an investigation has already been assigned.
Supporting evidence
Download the redacted evidence ZIP
The archive includes the full incident report, historical stage reviews and accounting snapshots, approved corrective plan, petting/water diagnoses, redaction notes, and a SHA-256 manifest.
Personal account names and local user-directory prefixes have been removed. Internal task/session and browser-context identifiers use consistent aliases; use this issue's feedback reference for internal correlation. The private alias mapping is not included. Timestamps, credit balances, token counts and diagnostic measurements are preserved; numeric values in all nine JSON evidence files were checked against the originals.
Archive SHA-256: 25c002eaddd0cefd45205e00332f8c8743dcafe90ef37304630c85df9de4f156.
Historical references to files not yet being uploaded describe the earlier state when those records were written. This is a redacted evidence copy, not an anonymous report; the GitHub issue itself identifies its author.
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start with the attached redacted evidence ZIP and the feedback reference 01a08389-23ef-7880-97af-8f64e25eef20. Review the incident report, historical stage reviews, accounting snapshots, and approved corrective plan first. Done means identifying whether model behavior, delegation, context or compaction, tool routing, or workflow controls caused the acknowledged constraints to be ignored, without inferring causes the evidence cannot isolate.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- powershell, rust
- Domain
- ai, devtools, tooling
- Issue type
- Bug
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Active
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100