AI systems are growing more complex, more embedded, and more consequential all while the practices for evaluating these systems remain strikingly thin. Models now shape high-stakes decisions across private and public sectors, yet are often assessed through static benchmarks, internal demos, or short-term testing that reveal little about real-world performance, distributional effects, or downstream impact. This Article argues that effective AI governance must be built around legally grounded, ongoing evaluation rather than episodic or purely technical assessment. Drawing on public law frameworks governing performance measurement, evidence-based policymaking, and procurement, we show that existing law already supplies the mandate, authority, and framework for more rigorous, system-level evaluation of AI deployments. For evaluation to produce meaningful oversight, central institutional challenges of independence and resourcing must be overcome, coupled with necessary technical advances. Risk-tailored evaluation that integrates pilot testing, randomized trials, quality assurance, and performance measurement will also be central. We conclude by outlining institutional models through which AI researchers, funders, governments, and private actors can embed evaluation as a core feature of AI governance.