Copilot Context Window Showing ~40% Reserved Output Even With Minimal Prompt #188691
Replies: 31 comments 18 replies
|
What you're seeing isn't a bug or a configuration error; it is actually a standard safety feature of how GitHub Copilot manages its 192k token capacity. Think of the "Reserved Output" as a guaranteed parking space for the AI's response. Even if you only say "hi," the system immediately sets aside about 30% of the total window (roughly 60k tokens) to ensure that if you later ask for a massive code refactor or a long explanation, the AI has enough "room" to finish writing without being cut off mid-sentence. To answer your specific points:
In short, you still have plenty of room (over 100k tokens) for your code and prompts. The "Reserved Output" is just the system making sure it can always talk back to you effectively. |
|
Thank you for sharing that @zckLab, it makes more sense having that context. I was confused when I saw the UI element as well. I think it might be worth considering rendering the UI as a meter of available context and context actually used --subtracting out the reserved context from the total. This would allow the user to see roughly how much strain is on the model that they can do something about (compression, being more selective with tools, beginning to wrap up the session, etc.). eg I'm wondering now if I've aborted sessions prematurely because of the way the UI element represented the situation. |
|
@TheNotary I agree with you; it confuses me too. I will prefer to see the context window UI percentage only based on the context used in my chat window. "Reserved Output" Should be only informative and shouldn't affect the calculations of the current context window. I usually watch the context window size when I'm working on some feature, and when it exceeds 50%, I always switch to a new chat window. Right now I am very confused because I don't know when my actual conversation exceeds 50% and when it is because there is some reserved output. @zckLab could you help here? |
|
The breakdown you're seeing (system instructions + tool definitions + reserved output eating 40%) is a good argument for treating the instruction layer as a first-class budget concern, not an afterthought. Prose system prompts expand silently — there's no natural stopping point, so they tend to grow until someone notices the context is gone. Typed instruction blocks (role, constraints, examples, output format as separate, bounded fields) let you reason about what's consuming context before you send anything. You can prune a specific block without rewriting the whole prompt. I've been building flompt for exactly this, a visual prompt builder that decomposes prompts into 12 semantic blocks and compiles to Claude-optimized XML. Open-source: github.com/Nyrok/flompt The structured format also tends to be more token-efficient than equivalent prose because there's less ambiguity for the model to resolve. |
|
I don't see the purpose of reserved context windows. Let's say there's 30% of this. Without the reserved window, when I reached 70%, there's 30% left for me, so the output won't be cut off. With the reserved window, when I reached 70%, there's 30% reserved, so the output won't be cut off. There's little difference, and reserving only adds to confusion. |
|
This is expected behavior. The context window in Copilot includes more than just your visible prompt. Even in an empty chat, tokens are already used by:
Because of this, you may see around 30–40% usage even with a simple message like So the usage you’re seeing is normal and not a configuration issue. |
|
There's clearly an issue. I am getting the same results as others. Yesterday I had no issues with long conversations and coding tasks without ever getting close to the context window needing compaction. Today, the context window compacts nearly every other message. In one 15 minute session it's compacted 6x already. |
|
second this, when I use claude opus 4.6 to analyze a run log of 500 row and the execution plan, it reaches 200k upper limit immediately, and then it start to took long time to compacting conversation, which make the model barely unusable, as compacting conversation is taking very long time even tho the conversation just started. |
|
Agree with these. It's useless right now |
|
I was quite pleased with Claude Opus 4.6, but it became literally unusable, even GPT 5.3 with context window twice as large is pretty much unusable after a short while. |
This comment was marked as spam.
This comment was marked as spam.
|
The problem that I observed is that it was working just fine until like 2 days ago, that's when I started to hit Compacting conversation every few minutes... |
|
Same thing for me. It was working great a week ago, and after the new update, it keeps compacting and big prompts take much, much, much longer. |
|
Quick Resolution: This is expected behavior, not an issue to fix. The ~30-35% reserved output space is intentionally allocated to ensure Copilot can complete long responses without truncation. What You Can Do: Accept it's normal – The reservation is by design and cannot be reduced Optimize your input space instead: Start fresh chats for new topics Use /clear to reset conversation history Reference files with #file rather than pasting content Keep prompts concise Bottom line: No action needed – your available input space is still ~120k tokens, which is plenty for complex tasks. |
|
So far I am seeing a much reduced reserve on Opus 4.6 today (15-30%). Looks like the 60% on opus was perhaps unintended yesterday. |
|
Oh! That's why it was auto compacting after only 80-100k tokens in other tools like opencode for models like Codex 5.3 via Github Copilot, because the backend is nerfed. I am pretty sure it was not like this just a few weeks ago. |
|
Copilot started to forget instructions and context extremely fast. Disappointing degradation |
This comment was marked as low quality.
This comment was marked as low quality.
|
I'm not sure what the best approach is here. Perhaps totally leaving the 'reserved output' out of the UI? VSCode users are generally only concerned with input (at least I am) not output. So maybe the percent should just be for input? It's interesting to know how much is reserved for output, but the user who started this thread found the high percent before even starting a chat confusing, and myself and perhaps nop1984 find the low percent before context compaction confusing. So perhaps the best fix is to leave 'reserved output' out of the percent calculation. If the 'Reserved Output' is a static, non-configurable budget that we cannot use for input, then perhaps it shouldn't be part of the percentage calculation at all. Just a thought. |
|
So im not new here, just want to know how to get out of this situation.. |
|
I've stopped using copilot. That solved the problem for me. Thank you everyone! |
|
Hello! Just wanted to follow up and close this discussion as there is outdated information. Reserved output is capacity set aside for the model’s response and can’t currently be reduced or configured. It does not mean those tokens have already been consumed. Copilot has since moved from the request-based model introduced in 2025 to usage-based billing for most plans in June 2026. Billing is based on input, output, and cached tokens actually consumed, not the maximum reserved capacity. Individual users can review their consumption in the usage dashboard. See Models and pricing and Usage-based billing for individuals. If you have additional questions on usage-based billing, please open a new discussion. |













Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
Select Topic Area
Question
Copilot Feature Area
VS Code
Body
Title: Copilot Context Window Showing ~40% Reserved Output Even With Minimal Prompt
Hello everyone,
I recently noticed the update where GitHub Copilot's context window was increased from 128k tokens to 192k tokens, which is great.
However, I am experiencing an issue related to the context window usage.
Even when I open a completely empty chat and send a very simple message like:
the Context Window indicator already shows around 40% usage, and most of that appears to be labeled as Reserved Output.
Example from the Copilot UI:
This happens even when there is no meaningful conversation history, which makes it feel like a large portion of the context window is already consumed before doing any real work.
My questions:
If anyone from the team or community has insight into how the reserved output allocation works or how to optimize the available context window, I would really appreciate the clarification.
Thanks.
Screenshot:

All reactions