Describe the feature or problem you'd like to solve
When a compact happens matters a lot for cost. Done while the prompt cache is still warm, it mostly reads the context from cache, which is cheap. Done after the 5-minute cache window, the whole context has to be written to cache again, which costs several times more. The agent usually knows the best moment: a phase of work just ended, the task is wrapping up, or we're about to wait on CI, reviewers, or a manual test. But only the user can type /compact , so the moment is often missed. The agent can remind me in its reply, but if I've stepped away or don't notice, the cache expires and the compact gets expensive.
Proposed solution
Let the agent propose a compact as an approval prompt inside the CLI, like the tool-permission prompts. The prompt shows suggested focus text and offers Yes, Edit, and No. Yes runs /compact with that focus text right away, while the cache is still warm. The user stays in control, and the agent never types into its own terminal.
Optional extras:
• A setting that allows the agent to compact automatically when a phase ends.
• A notification (sound or system toast) when a compact is suggested, in case the user is in another window.
• The estimated credit cost of compacting now versus after the cache expires.
Example prompts or workflows
- The root cause of a bug has been found, and work moves on to the fix. The CLI shows: "Good time to compact. Focus: keep the root cause, changed files, and next steps. [Yes / Edit / No]". The user presses Y.
- The user says "I'm heading to lunch." Before the reply ends, the agent offers a compact so the session is small and cheap to resume.
- A PR has been opened and is waiting on CI and reviewers. The agent offers a compact focused on the PR status and pending checks before the session sits idle.
Additional context
The agent can't drive its own terminal; Computer Use is blocked from controlling Windows Terminal, and that's the right call. My current workaround is a personal script that plays a sound and shows a pop-up with the /compact command, copying it to the clipboard so I can paste it. It works, but it relies on me noticing in time, and an approval prompt inside the CLI would be cleaner and safer. In my usage, a compact on a session that had sat idle for about an hour cost about 35 AI credits, and most of that was writing the context back to cache.
Describe the feature or problem you'd like to solve
When a compact happens matters a lot for cost. Done while the prompt cache is still warm, it mostly reads the context from cache, which is cheap. Done after the 5-minute cache window, the whole context has to be written to cache again, which costs several times more. The agent usually knows the best moment: a phase of work just ended, the task is wrapping up, or we're about to wait on CI, reviewers, or a manual test. But only the user can type /compact , so the moment is often missed. The agent can remind me in its reply, but if I've stepped away or don't notice, the cache expires and the compact gets expensive.
Proposed solution
Let the agent propose a compact as an approval prompt inside the CLI, like the tool-permission prompts. The prompt shows suggested focus text and offers Yes, Edit, and No. Yes runs /compact with that focus text right away, while the cache is still warm. The user stays in control, and the agent never types into its own terminal.
Optional extras:
• A setting that allows the agent to compact automatically when a phase ends.
• A notification (sound or system toast) when a compact is suggested, in case the user is in another window.
• The estimated credit cost of compacting now versus after the cache expires.
Example prompts or workflows
Additional context
The agent can't drive its own terminal; Computer Use is blocked from controlling Windows Terminal, and that's the right call. My current workaround is a personal script that plays a sound and shows a pop-up with the /compact command, copying it to the clipboard so I can paste it. It works, but it relies on me noticing in time, and an approval prompt inside the CLI would be cleaner and safer. In my usage, a compact on a session that had sat idle for about an hour cost about 35 AI credits, and most of that was writing the context back to cache.