SYNDICAL QUALITY EVIDENCE — PUBLICATION POLICY Version 1 · 3 October 2026 Purpose Our ambition is to develop systems that outperform human-developed software in correctness, security and maintainability. We start with our own repositories. That ambition is not a claim that a comparative result has been achieved. What we publish Publish the scope, date, source revision, command, outcome and limitations of each measurement. Preserve failures, blocked checks and unmeasured areas. Label a working-tree measurement explicitly and retain its source digest. Keep historical snapshots and identify corrections instead of silently rewriting earlier results. Never combine missing checks into a passing score. Evidence boundaries A successful build is build evidence. A passing test suite covers the tests that ran. Neither establishes full correctness, security or maintainability. Distinguish source review, local execution, hosted CI and production results. Do not claim deployment, certification or independent audit from local checks. Publish sanitized summaries and hashes; keep credentials, private logs, account identifiers, customer data and local paths out of the public evidence. Changes and reproduction Keep the published JSON under version control with the page. Each measurement must name its tool commands and revision. When a check, dependency or source changes, record a new measurement; do not relabel an older pass as current. Preserve the policy version used for each comparison. Commands and policy are provided publicly. Rerunning source checks requires access to the measured repositories; these downloads do not publish their source. Controlled comparison — proposed method 1. Specify representative tasks, acceptance criteria, primary outcomes and sample size before examining comparative results. Record selection rules. 2. Use the same initial code and requirements. Record human experience, tools, AI models, time budgets, compute costs and human interventions. Randomize task assignment or use paired tasks where practical. 3. Use held-out correctness and security checks. Assess maintainability with independent reviewers blinded to authorship where practical. Record any overlap between development and evaluation data. 4. Report all attempts, failures and exclusions, sample sizes, uncertainty and variation. Measure defects after release over a declared observation window. Limit conclusions to the evaluated tasks, conditions and quality dimensions. Current status No controlled comparison against human development has been completed or published in this baseline. This page presents our own engineering evidence and does not claim independent certification. Published measurements are self-assessed unless explicitly stated.