Using a hosted LLM, you pay per token, usually with output tokens costing more than input. Costs scale with prompt size (context), output length, and model tier — so a chatbot with a huge system prompt on every call gets expensive fast. Levers: smaller models where adequate, prompt caching, trimming context, capping output, and batching. Track cost per request like any other unit economic.