The system prompt is the cheapest and most reversible way to change model behaviour, and it resolves a surprising share of problems teams try to solve with fine-tuning.
It consumes context on every call, which is why prompt caching and keeping it disciplined both matter once an application is running at volume.
