# Syndical quality evidence programme

Public edition, 3 October 2026. Based on methodology at revision `e1c835c958f8e7e62ac61d720fd25666efd6a04f`. Repository-relative references are adapted for this download. Commands and policy are provided publicly; rerunning source checks requires access to the measured repositories.

Syndical's objective is to produce software that performs better than expert human development on specified, independently measured outcomes. This is a research and engineering objective. It becomes a defensible public claim only after the controlled evaluation described below. Repository checks, passing tests and lower duplication alone cannot establish it.

## First proving ground: Syndical

Syndical measures its own source with the same checked-in command in local development, GitHub Actions and Azure Pipelines:

```sh
npm ci
npm run quality:test
npm run quality:check -- --output .syndical-local/quality/first-run
```

The first command uses the lockfile. The second exercises the reporting safeguards, including a real scanner fixture. The third runs every configured measurement sequentially and writes `public/report.json` and `public/report.md`, even when a completed check fails. Existing output directories are refused to preserve earlier evidence. Commands and tool-version probes have no application deadline by default. A caller can explicitly request a per-command limit with `--timeout-ms`; the report records that request. Explicit cancellation still terminates the active process tree and retains completed results.

The policy is [downloadable measurement policy](./policy.json). The runner is `tools/quality/run.mjs` in the measured repository (source access required). They are hashed together in the report. The scanner dependency is pinned to jscpd 5.4.0 in the lockfile. Tool versions, operating system, architecture, OS-vault integration setting, start/end times, source commit, dirty state and a digest of the actual source contents are recorded. The source digest includes tracked and non-ignored untracked files, including local changes; a commit hash alone is insufficient for a dirty checkout. Published report exports under `docs/quality/reports/` are excluded to avoid a report hashing itself.

The source digest is checked again after the run. If source changed, the report retains the measurements but says they are not current proof, and the command fails. These files do not replace Syndical's application evidence store or establish a host execution receipt. A product gate should use the existing gate/assessment protocol and hash-pinned adapter boundary when consuming them.

| Measurement | Scope and interpretation | Enforcement in this first pilot |
| --- | --- | --- |
| .NET build | Entire solution in Debug; repository warnings-as-errors policy applies | Required |
| .NET tests | Projects named `*.Tests.csproj`, declaring `IsTestProject`, or referencing the .NET Test SDK; each runs sequentially after the build. TRX counts must exist and contain tests | Required; missing receipts remain errors and later projects still run |
| TypeScript typecheck | Existing workspace typecheck commands | Required |
| TypeScript tests | Both existing workspaces, including the IPC sidecar round trip; independent JSON test receipts | Required; requires successful .NET build |
| Desktop build | Existing desktop production build command | Required |
| Duplication | Git-known production C#, TypeScript and JavaScript under `apps/` and `libs/`; test/fixture, generated declaration and packaged resource paths excluded. jscpd token matching, minimum 5 lines and 50 tokens, mild mode, one worker; the scanner size setting is derived from the largest actual snapshot file so no selected file is excluded by size | Measurement required; the finding level is advisory pending a reviewed baseline policy |
| Estimated complexity | Same selected production source; jscpd's token-based estimate, total and file-level median/p95/max | Measurement required; advisory ranking signal, not verified function complexity or a quality grade |
| npm dependency advisories | Installed project and development dependencies using the current npm advisory service | Advisory; positive findings are recorded as failed, service/parse faults as measurement errors |

The scanner reads an exact private copy of the selected production bytes and an explicit empty configuration plus policy arguments. This prevents generated or ignored files, local jscpd configuration and a moving working tree from silently changing its scan. The report distinguishes selected files, files actually measured by the scanner and exclusions; very small files can fall below its token threshold. The scanner size setting is the largest immutable snapshot byte length plus one, recorded in each scan, so its built-in default cannot silently exclude larger source. Every measured filename is validated against the selected inventory and retained privately; a digest of that actual inventory is published. Raw reports retain detailed findings for remediation. Lines and tokens describe measurement scope and denominators, not quality.

The [jscpd documentation](https://github.com/kucherenko/jscpd/blob/master/docs/rust.md#summary) describes its complexity estimate as token-based, without a parser. It is useful for selecting code to inspect; it does not establish architectural quality, function-level complexity, correctness or a safe universal limit. [npm audit](https://docs.npmjs.com/cli/v11/commands/npm-audit) covers known dependency advisories and depends on the service's current database.

The report explicitly leaves coverage, mutation testing, independent maintainability review, SAST/secrets/.NET dependency scanning, production outcomes and human comparison **unmeasured**. There is no aggregate health score. A required advisory measurement means the instrument ran successfully, not that the source met an invented quality threshold. An npm advisory failure does not currently block the build; this initial policy must not be described as full security enforcement.

## Baselines and comparisons

Keep the baseline checkout immutable while measuring. Install its own lockfile dependencies, and invoke the current runner against it so both revisions use the same measurement policy and scanner:

```sh
# In the baseline checkout first: npm ci
node tools/quality/run.mjs \
  --repository /path/to/baseline-checkout \
  --output .syndical-local/quality/baseline-run

# After implementation, with the working tree held stable:
npm run quality:check -- \
  --output .syndical-local/quality/candidate-run \
  --baseline .syndical-local/quality/baseline-run/public/report.json
```

An output path is resolved against the caller's working directory; it must be outside the measured checkout or ignored by its Git rules. The measurement runner and jscpd come from the runner checkout. Build, tests, typecheck and installed workspace dependencies come from `--repository`. This separation must be preserved in the experiment record.

For a quick advisory scan, use `--checks duplication,complexity,npm-audit`. Every omitted check remains visible as unmeasured. For executable verification, select `dotnet-build,dotnet-tests,typescript,typescript-tests,desktop-build`, or omit `--checks` for the full policy. Test checks require `dotnet-build` to pass in that same run. They do not silently trust an older sidecar or test binary.

The comparison produces signed numeric deltas only when report schema, policy/runner digest, tool/environment versions, unchanged-during-run source state and the check's selected file inventory match. Scanner comparisons additionally require matching actual measured-file inventory digests; a file falling below the tokenizer threshold cannot disappear from the comparison. Older reports lacking that identity are not comparable. A changed inventory or toolchain is reported as **not comparable**, requiring a new baseline or an explicitly designed broader study. This prevents scope removal from being presented as improvement. It does not prohibit new files; it prevents claiming an equivalent-scope metric delta from them. npm audit results never get causal improvement deltas because the advisory database is not pinned. Test-count increases demonstrate more executed tests, not improved test effectiveness.

Once comparable baselines exist, introduce reviewed change policies: reject newly introduced confirmed critical findings, protect changed-code test effectiveness, and inspect any increase in duplication or complexity. Preserve accepted historical debt and its ownership; never reset a baseline automatically to make a failing change appear acceptable. A policy or exclusion change requires a new policy digest, rationale and comparable remeasurement.

## Publication and custody

Raw logs, source snapshots, TRX/JSON receipts and detailed scanner output live in the ignored `.syndical-local/quality/` directory. They can contain private paths, source snippets and runtime diagnostics. The public reports contain a fixed set of aggregate metrics, check outcomes, source/policy/tool identity and a digest of the retained raw evidence. The CI jobs upload only `public/`, including on check failure. They do not upload raw logs or automatically publish a successful marketing claim.

Public results for the product belong in the `syndical.site` repository and must be sourced from the generated sanitized aggregates. Publish failed and unmeasured outcomes alongside passing ones, the selected scope and policy, the exact measured revision/content digest, limitations and a reproducible command. Keep each dated result rather than silently replacing an earlier result. The report is evidence for the measured snapshot; rerun after source or policy changes before describing it as current. Each other target repository owns its own adapters, baselines, diagnostic evidence and publication content.

## Controlled evaluation before a human-superiority claim

1. **Register the claim and experiment before implementation.** Choose a defined population of tasks and repositories, a representative held-out task set and primary outcome. Specify acceptable functional, security, performance and maintainability requirements, sample size, stopping rules and the statistical analysis before results are available. Start with Syndical task pilots, then replicate on repositories with different languages and architectures.
2. **Use a matched control.** Give experienced human developers and Syndical the same starting source, written task, documentation, tools, environment and acceptance constraints. Record participant experience, model/tool versions, elapsed time, active human effort, model/compute cost, review effort and number of attempts. Randomise assignment or counterbalance matched tasks. Describe existing authorship as mixed or unknown unless it is actually established; an older commit is not automatically a human control.
3. **Keep evaluation independent.** Freeze hidden acceptance, adversarial and security tests before solution authors see them. Separate builders from evaluators. Blind reviewers to author and condition; use a rubric for architecture, readability, maintainability and test relevance, retain individual judgments and disagreements, and report agreement. Use executable outcomes wherever possible.
4. **Measure whether tests detect defects.** Compare mutation detection and seeded-regression detection, not just test count or line coverage. Include deployment/integration acceptance, meaningful performance workloads and sustained escaped-defect and reliability follow-up. Record false positives and defects independently confirmed by reviewers.
5. **Count every assigned attempt.** Publish failures, timeouts, abandoned tasks, exclusions with reasons and remediation effort as well as successful deliveries. Do not select only tasks where the system wins. Report uncertainty and paired effect sizes for the registered primary outcome; do not combine unrelated metrics into an unvalidated score.
6. **Constrain the conclusion to the evidence.** A claim should name the evaluated tasks, control population, budget, outcome and uncertainty. A duplication decrease is a duplication result. Superiority in accepted-task success at a stated budget does not prove universal software superiority. Replicate an initial result on the held-out set and across relevant stacks before broadening the claim.

The delivery programme should progress through truthful measurement and enforced current evidence, independent review and defect-detection checks, guarded repair loops with retained findings, then the controlled evaluation. Every stage should produce reproducible results on Syndical and publish its limitations. The mission remains better software; the acceptance criterion is independently demonstrated outcomes.

## Product enforcement in this release

Newly compiled squads with a code reviewer downstream of an implementer pin the `code-quality-v1` review contract in their accepted snapshot. The reviewer must give separate, evidence-backed spec and standards outcomes. Missing or unverified outcomes block completion. A failed outcome needs a concrete finding on that axis and routes through the existing bounded repair path. The Feature template includes this code reviewer alongside its separate test reviewer, after measured gates and before acceptance. Existing accepted snapshots retain their original contract; diagnostic reports and planning documents use their own review requirements.

These outcomes are model judgments about the supplied change and evidence. They do not establish independent human evaluation, mutation effectiveness or execution of a command. A reviewer must reuse retained measurements and must not claim execution merely from reading source. The existing finding ledger and repair budgets remain authoritative.

Coverage gates now reject absent, malformed and unsupported measurement evidence even when a command exits successfully. A proven threshold violation is a failure; missing instrumentation or an unverifiable denominator is blocked, requiring setup or reviewed adapter evidence. Coverage contract versioning invalidates older coverage receipts. This validates a configured coverage gate; it does not claim Syndical's own coverage has been measured in the pilot.

Repository Quality and Metrics pages label recorded results as historical. Missing commands, missing readings, changed definitions and legacy readings without definition identity remain visible. A partial set of readings cannot produce a complete pass-rate percentage. Current-source delivery verification continues to use the existing immutable execution receipts and source identity; the repository overview is not a substitute for that gate.
