A model has no state between calls. Every turn, it receives the system prompt, the full prior conversation, retrieved documents, tool definitions and your new message, reads all of it, answers, and retains nothing.
Roughly a million tokens is now standard at the frontier. But advertised size and reliably usable size differ: material in the middle of a very long input is used less reliably than material at the edges.
