Training frontier models and serving billions of queries use significant electricity and water (for cooling), concentrated in data centers. Inference at scale can outweigh one-time training over a model's life. The field is responding with efficiency (better hardware, quantization, MoE, smaller task-fit models) and cleaner energy sourcing. It's a real cost and externality to weigh, not a reason for paralysis — efficiency is both green and cheaper.