The problem of computable normative choice
Regulation changes people's opportunities, relationships, and conditions for action. Given a normative commitment, the quality of regulation depends on how fully and accurately the reality being regulated enters into the decision. JudgeAI aims to build a normative apparatus in which empirical data directly constitute the content of the normative criterion. Better understanding of reality would thereby become a mechanism for improving normative decision-making.
Advances in analysis and prediction alone leave open how new knowledge changes decisions. Recent research in decision-focused learning demonstrates this issue in optimisation: better predictive metrics may not produce better decisions. This motivates training and evaluating empirical models in relation to the consequences of the decisions they support. Zhang et al., NeurIPS 2025; Ren et al., UAI 2024.
Normative decision-making requires specifying how empirical characteristics constitute normative relationships. When data serve only as context for a subsequent judgment, their influence on the result remains dependent on that interpretation. The research problem is to build an apparatus in which data directly constitute the content of the normative criterion, with a computable and testable influence on the outcome. Representing the situation, identifying relevant dependencies, constructing possible states, and comparing alternatives form a single process. Normative choice begins within the analysis itself.
JudgeAI makes its foundational normative commitment explicit: preserving participants' agency—their capacity to act. Empirical data determine what this commitment means in the situation being regulated and how alternative regulations change the conditions for fulfilling it. Regulation also changes participants' behaviour and the environment for subsequent decisions. Research on performative prediction and interacting agents examines the importance of this feedback. Góis et al., AISTATS 2025.
Why this could improve on human normative decisions
Given the same normative commitment, a better decision is one that fulfils it more accurately under actual conditions. Human capacity to jointly account for relationships between possible actions and their consequences is limited by the computational complexity of the task. In experimental tasks involving the choice of means to achieve goals, more complex dependencies reduced the quality of human decisions even when the number of elements remained unchanged. Reichman et al., 2023.
JudgeAI investigates how thousands of consequences and many alternative states can enter directly into normative computation. When empirical content constitutes the criterion itself, computational capacity to analyse reality becomes capacity to account for its normative significance in determining the decision. The potential advantage is the ability to identify preferable regulations whose determination requires jointly processing more dependencies than unaided human reasoning can handle.
This potential rests on a more complete and consistent computation of how alternatives fulfil the adopted normative commitment. With a sufficiently reliable model of reality, it creates the possibility of choosing regulation that fulfils that commitment better. The research question is how far this potential translates into better decisions in specific tasks, compared with human decisions under comparable initial conditions.
What a solution must achieve
Traceability must make it possible to establish which empirical grounds shaped the normative relationships and the result, what needs further investigation, and how refining the model changes the decision. Refining data, testing dependencies, and expanding alternatives must form a single cycle of improvement of the normative apparatus. With the complete grounds and the version of the computational procedure held constant, the result must be reproducible. New evidence must allow normative relationships to be recomputed, whether this preserves or changes the preferred alternative.
The claimed connection between grounds and results must itself be tested. OpenAI studies the reliability of reasoning analysis as a way to evaluate model behaviour. Anthropic tests explanations by making counterfactual changes to inputs and measuring the results. These studies offer methodological guidance for testing explanations and dependencies within a normative system. OpenAI, 2025; Anthropic, 2026.
A new regulatory situation usually has no ready-made, generally accepted benchmark for the correct normative decision. Evaluation must therefore cover empirical grounding, reproducibility, sensitivity to relevant changes, and the actual consequences of regulation as they become observable. Human decisions provide a basis for comparison. Agreement with them does not by itself establish normative correctness.
JudgeAI develops this approach as a long-term research programme in autonomous normative choice. The goal is decision-making without human involvement in the normative computation, with progressively more accurate and complete empirical content within the normative apparatus, and tests of how these improvements affect the quality of regulation.