Open-weight models (Llama, Mistral, Qwen) can be downloaded, run privately, fine-tuned, and audited — control, privacy, no per-token fee, but you handle infra and they often trail the absolute frontier. Closed models (Claude, GPT, Gemini) offer top capability via API with no ops burden, but you send data out, pay per token, and can't inspect or self-host. Choose by privacy, cost, control, and capability needs.