AI Lawmaking empirical benchmark
AI Lawmaking: Empirical and GPT Baseline Comparison
This report evaluates twenty final bills produced by the AI Lawmaking pipeline against real legislative regimes with known empirical results. The cases include historically effective practices and historically weak or adverse practices, from the Acid Rain Program and EITC to Prohibition, Three Strikes, demonetization, and No Child Left Behind.
A separate comparison uses GPT-5.5 Extra high outputs for the same normative requests. The result highlights a substantive difference in regulatory priority: GPT often selected familiar legislative templates, while AI Lawmaking more often designed around consequences, failure modes, side effects, actors, constraints, and causal mechanisms.