Token costs in production can surprise you fast. That’s especially true with agentic workloads, where a single user action can fan out into five, ten, or more LLM calls under the hood. Without visibility, you won’t know there’s a problem until the bill arrives. By then, the damage is done.

image of a woman holding a large amount of dollars and handing one over to the screen

The good news: there are concrete, low-effort steps you can take right now. Here are four of them:

1. Set TPM limits on your deployments

Tokens per minute (TPM) limits cap runaway usage at the source. Set them per deployment so a rogue agent run can’t burn through your budget before anyone notices.

2. Alert on TokensConsumed in Azure Monitor

The TokensConsumed metric is available per model and per deployment. Set a threshold alert so you’re notified before you hit your budget ceiling, not after.

3. Log token counts against your run IDs

The SDK response already gives you prompt and completion token counts. Log them against your agent run IDs and you can trace any expensive call back to exactly where it happened.

4. Trim your system prompt

Your system prompt runs on every single call. Cutting 200 tokens from a prompt called 10,000 times a month saves 2 million tokens. Audit it ruthlessly and remove anything that isn’t load-bearing.

The PTU exception

caution yellow tape over the top of a greenfield backdrop

One thing that catches teams out: if you’re on a PTU (provisioned throughput unit) deployment, token counts are the wrong thing to monitor. With PTU, you’re paying for reserved capacity regardless of how much you use it. Watching token consumption doesn’t tell you whether you’re getting value from what you’ve bought.

On PTU, the metric that matters is utilisation. You want to know whether you’re running hot (saturating your provisioned capacity and potentially queuing requests) or cold (underutilising and paying for headroom you’re not using).

Save PTU for production. Use pay-as-you-go for dev and test. Provisioned throughput is cost-effective when you have a predictable, high-volume workload. It’s wasteful for development and testing, where usage is bursty and unpredictable. Run your dev and test environments on standard pay-as-you-go deployments and you’ll avoid paying for capacity that sits idle between runs.

Bottom line: cost visibility isn’t a nice-to-have once you’re in production, it’s a requirement. TPM limits, Azure Monitor alerts, run-level token logging, and prompt hygiene are four things you can act on today. None of them take long, and together they’ll give you genuine control over what you’re spending.

Tags: , , , ,