All guides

How to reduce your ChatGPT and Claude token usage

Seven practical habits that cut the tokens you send to ChatGPT, Claude, and Gemini without making the answers worse, plus how to automate them.

October 5, 2026 · 6 min read

Every message you send to ChatGPT, Claude, or Gemini is split into tokens, small chunks of text. In English, a token is roughly four characters, so a 100-word prompt is somewhere around 130 tokens. Tokens are what you pay for on an API bill, and they're what quietly eats into the usage limits on subscription plans.

The good news: a lot of what we type is filler the model doesn't need. These habits cut it without making the answers any worse.

1. Drop the politeness and hedging

“Could you please help me,” “if it's not too much trouble,” “thank you so much in advance.” Models don't need any of it to do a good job. Lead with the verb.

Beforecan you please help me understand what an API is and how it works
AfterExplain what an API is and how it works.

2. Say each requirement once

Long prompts often repeat themselves: the same constraint stated at the start, in the middle, and again at the end “just to be sure.” Write requirements once, as a short list, and the model follows them just as well.

Beforeexplain machine learning to me like im 5 years old please and thank you
AfterExplain machine learning like I'm 5.

3. Paste only what's relevant

Pasting a whole file when the bug is in one function, or a full log when three lines matter, is the biggest token sink in coding workflows. Trim pasted code and logs to the part the question is actually about.

4. Don't re-paste context every turn

Within a conversation, the model already has what you sent earlier. Refer back to it (“in the function above”) instead of pasting it again.

5. Start a new chat for a new topic

Each new message is processed along with the conversation so far, so a long thread makes every reply heavier. When you switch to an unrelated task, open a fresh chat.

6. Ask for the length you need

Replies cost tokens too. If you want three bullet points, say so. “In two sentences” or “code only, no explanation” keeps the answer to what you'll actually use.

7. Automate it

These habits are easy to forget mid-task. TerraWatchdoes the trimming for you: it sits next to the message box in ChatGPT, Claude, Gemini, and Perplexity, rewrites your prompt to use fewer tokens with one click, keeps every requirement, and tracks what you've saved. For API and coding workflows, the twc CLI does the same in your terminal.

Curious what it's worth for you? Try the savings calculator.