What Claude can see
The context window is everything Claude can see: the whole current conversation, re-read from the top every time you send a message, and nothing outside it. Lesson 1 said the model reads the whole conversation every time. This lesson is what follows from that, and it is the single most useful mechanical fact about using Claude.
There is a limit on how much conversation fits. It is large, and you will not hit it on a normal afternoon. But long before you hit it, something else happens, and that is the part worth understanding.
A conversation is a room that fills up
Picture everything the model can see as a room with a fixed amount of floor space. Into it goes your first message, its reply, your second message, its reply, every document you attached, and every long thing it produced.
Two consequences, and the second one surprises people.
The room costs money to carry. Everything in it is re-sent on every message. Message forty is paying for messages one through thirty-nine. This is why the same plan gets you much further if you start a new chat per task.
A crowded room answers worse. By message forty, the thing you care about is one paragraph among fifty, competing with a document you attached for a different reason, an answer you rejected, and a tangent you abandoned. There is nothing marking which of those still matters.
That second one is the real cost, and it is invisible. Nothing goes wrong. The answers just get a bit more generic, a bit more likely to reference something you moved on from, a bit less sharp. If you have ever thought “it seemed better earlier”, this is why.
What that looks like
The tells, roughly in the order you will meet them:
- It brings back a constraint you dropped ten messages ago.
- It gives you a version of something you already rejected.
- It starts hedging and summarising instead of answering.
- It answers a slightly different question from the one you asked.
None of these are the model going wrong. All of them are the room being crowded.
The fix is a new chat
That is it. That is the technique.
Start a new conversation when you change task. Not when things break, when you change task, which is much more often than most people do it.
The objection is always the same: but it will not know what we already established. Correct, and that is the point. Everything you established that still matters, you can say in three lines, and those three lines are worth more than forty messages of history because they contain only the parts that survived.
The move that works:
Summarise everything from this conversation that I would need to carry into a new chat
to continue: what we decided, what we ruled out, and what is still open. Be brief.
Copy that. New chat. Paste it at the top. You have kept the conclusions and left the wreckage.
What to attach and when
Attaching a document puts the whole thing in the room.
For one document you are working on, that is exactly right and you should stop reading this section.
For six documents when the answer is in one, it is expensive and it makes the answers worse in the way described above. Attach the one. If you do not know which one, attach them in a project instead, which is lesson 5, and where a project earns its keep.
The same logic applies to pasting. A whole page pasted so it can answer a question about one paragraph is four pages of room for one paragraph of relevance.
What survives between conversations?
Worth knowing, because it explains a specific confusion:
- The current conversation: everything, until it is too long.
- Your standing instructions from lesson 2: applied to every conversation, always.
- Memory, if you have it on: facts it saved, carried across.
- The previous conversation: nothing. It is a different room.
So a thing you want in every conversation goes in your instructions, and a thing you say inside one conversation dies with it. That is the whole of why lesson 2 spent time on the instructions box.
Your turn
You are going to make it degrade on purpose, because reading about it is not the same as seeing it.
Open a new chat and give it a constraint at the top that is easy to check:
For this whole conversation, answer everything in exactly two sentences. Never more.
Then have a long conversation. Fifteen or twenty messages, on anything: a plan for something, a topic you know about, questions about a document. Wander. Change the subject twice.
Every five messages or so, ask something ordinary and count the sentences.
Check: somewhere in the conversation you get three sentences, or five, and the two-sentence rule quietly stops holding. Note roughly which message it happened at. That is the point where the room got crowded enough for your instruction to stop being the loudest thing in it.
Now do the recovery. Ask for the summary:
Summarise everything from this conversation I would need to carry into a new chat.
Be brief.
Paste it into a new chat with your two-sentence rule at the top, and carry on.
Check: you are back to two sentences, and you did not lose anything you needed. Both halves of this lesson, in ten minutes, on your own screen.
Recap
Everything the model can see is one room, and it is re-read from the start on every message. A crowded room costs more and answers worse, and the answering-worse part is silent.
The tells are old constraints coming back, rejected versions returning, and hedging. The fix is a new chat, at every change of task rather than at every disaster.
Carry the conclusions across with a summary, not the history. Attach the document you need rather than all of them.
A thing you want in every conversation belongs in your standing instructions. A thing you say in one conversation dies with it.
Next: the place to put the context you want on every conversation about a particular piece of work.