Skip to content
YEGFacts.ca
On this page

Methodology

Methodology changes

This is the public record of changes to how findings are produced. Each version has a quick summary, the main changes, links for context, and the complete change note. Every claim records the version that produced it.

  1. v1.9

    Framing check checkability and first use

    First use of delegated briefs: both panels halted on framing concerns the check had passed. The check now sees the output schema, must name a published source for every threshold, and must verify that every instrument a brief relies on exists.

    What changed

    • Both first briefs passed the framing check, then the research panel halted synthesis on a MATERIAL FRAMING CONCERN. Infill: the brief relied on a mature-neighbourhood instrument in force on the freeze date, and none exists since the overlay was retired on 2024-01-01. Low-density history: a density rule the check itself had demanded could not be computed from the 1926 record, and the brief demanded variant verdicts the schema cannot carry.
    • The framing checker now receives prompts/review-schema.json and may not ask for outputs it cannot carry; must name the published source a reviewer would compute any threshold from, or drop it; and must verify that every instrument, boundary, dataset or definition exists as described on the as-of date.
    • The reviewer prompt says to write MATERIAL FRAMING CONCERN only to raise one; the synthesis script ignores a negated mention, since a seat wrote 'no MATERIAL FRAMING CONCERN' and a plain text match would have halted on it.
    • The panel quality ledger gains a framing stage and a framing-miss event; six events from the two first runs are recorded, four valid catches by seats and two misses by the check.
    • Both stories are being rerun under revised briefs: infill with the dwelling unit as the verdict unit and a frozen historical overlay boundary, low-density history with classification by each era's own terminology and documented housing form, counts from published neighbourhood records. Evidence from the halted runs stays in the registry.
    • Date erratum: the first run directories and the v1.7 and v1.8 entries were dated 2026-09-02; the actual date was 2026-09-01. Entry dates are corrected; committed directory names are kept and each run's errata records it. The runner's pinned Gemini command could not run as written and is corrected.
    Full change note

    The framing check was designed to keep a brief fair; on first use it kept the brief fair and made it unresearchable, because the checker optimised definitions without knowing what the output schema carries or what the historical record can supply. The panel caught both cases in round one, which is the method working as designed, and the cost was two halted runs. Three changes follow: the checker sees the schema, the checker must name the published source behind every threshold it asks for, and the checker must confirm that the instruments a brief relies on exist. The halted runs, their concerns, the revised briefs and every check report are committed beside each other so a reader can see the loop.

  2. v1.8

    Framing check escalation and bound

    Framing-check disputes now resolve with Stew as editor, on the record; the check no longer asks for alternatives to alternatives.

    What changed

    • A finding still open after two revisions is decided in writing by Stew, the editor responsible for content, with the checker's standing objection kept in the committed record. The founder is not asked to arbitrate; he stays accountable and can revert.
    • A cutoff justified from an identified standard, or stated with one reasonable alternative and results under both, is sufficient. The checker may not ask for an alternative to the alternative.
    • First use: both infill briefs reached check 3 still at REVISE; the open findings and the editor's resolution of each are committed beside the briefs.
    Full change note

    v1.7 sent unresolved framing disputes to the founder. On first use both briefs reached that point, and the founder declined the role: editorial content is the steward's responsibility, he supplies ideas. So the escalation now ends with Stew deciding in writing, and the record shows what the checker still objected to. The founder's accountability is unchanged in kind: a standing power to revert, not a per-item approval. The second change closes a regress the first use exposed: a predeclared threshold was asked for an alternative, and the alternative would then have needed one. One identified standard, or one alternative with results under both, is now enough.

  3. v1.7

    Delegated briefs and framing check

    Stew, the AI steward, now writes review briefs; a second model from a different vendor checks the framing before a brief is frozen. Applied retroactively, the check would have sent all four published briefs back for revision.

    What changed

    • Briefs are drafted by Stew from an intake record that states where each circulating form came from, or that it was not captured. No founder approval is needed per story; the founder's accountability is unchanged.
    • Before a brief is frozen, a model from a different vendor runs prompts/framing-check.md; it looks for propositions that test a weaker or wider claim than the post made, verdict-sensitive definitions with no stated alternative, and sentences that tell reviewers where to land. The brief is frozen only on FRAME OK; after two unresolved revisions the disagreement goes to Ildar Abdulin in writing.
    • Every check report, the author's response and any resolution are committed beside the brief. A framing defect found later by the panel, the gate or a correction is logged against the check in the quality ledger.
    • Run retroactively on the four published briefs, the check returned REVISE on all four, mainly for site-written paraphrases presented as circulating forms, sentences naming which evidence 'would contradict' the claim, and folded or undefined propositions. The reports are committed beside each brief; the findings stand as published and each story will be re-framed and re-run under this version.
    • Two specialist board reviews (Evidence and Methodology on GPT, Public Trust on Gemini) reviewed the design before first use; both said revise. Their controls that were adopted are the intake record, the alternative-operationalization requirement, the mandatory re-check, the human circuit-breaker and the ledger entry. Not adopted yet: a public waiting period between freeze and panel, and a blind human-methodologist pilot.
    Full change note

    Every brief so far was drafted by an AI session and ratified by the founder at the publication gate, which D-0007 then delegated to AI audits. This entry says so plainly and closes the loop: Stew writes the brief, and the safeguard that delegation needs is a framing check by a model from a different vendor, run before the freeze, whose report is public. The check's questions are the ones a fair-minded holder of the claim and a fair-minded opponent would both ask: is this the claim I made, is it the strongest fair reading, would I have accepted these definitions before seeing the result, and does the brief tell the reviewers what to find. Run on the four published briefs, the check said all four would need revision. None of that changes a published finding, because the panel and the gate decided on evidence, but it does mean the frames the panel was given were weaker than this version requires. The reports are committed beside the briefs, and the four stories are scheduled to be re-framed and re-run under v1.7 rather than patched. The two specialist reviews are archived in the board's private record; their two unadopted asks are recorded there as open questions, with the founder's decision pending.

  4. v1.6

    Reviewer effort pinning

    Pinned the reasoning-effort setting for every panel seat and started recording it in run manifests.

    What changed

    • The Claude command now carries --effort high; the GPT and Gemini commands already pinned it.
    • Every seat runs at its vendor's "high" setting, the highest level the three CLIs share.
    • Claude and Codex offer settings above high. They are not used: the founder finds the extra cost and run time unmatched by better results, and nothing here benchmarks that.
    • The runner writes the effort into each review file and the run manifest, and the site displays it from the manifest.
    • Runs before v1.6 show "not recorded" for effort; the Claude seat's effort came from an unpublished local default and is not backfilled.
    • The run manifest now also records each seat's display name, so the site no longer reconstructs it from the provider field.
    • Tooling correction: the manifest stamper read the wrong end of the changelog and wrote methodology_version 1.0 on runs made under later versions. Fixed and tested; earlier manifests are left as written.
    Full change note

    Models are released often — Fable 5.1 shipped the day of this entry — and the reasoning-effort setting changes what a model returns as much as the model version does. Leaving it unpinned made the published command incomplete for one seat: prior Claude runs inherited an unpublished local default rather than a setting on the record. Effort is now pinned in the command, stamped onto each review file by the runner, and recorded per run in run.yaml. Every seat is pinned to "high" for two reasons: it is the highest level all three CLIs share (the Gemini CLI stops there), and while Claude and Codex offer settings above it, the founder does not use them for the panel — in his experience they cost significantly more and run longer without a matching gain. That is a judgement call, not a measurement; no benchmark of the trade-off exists here. Two smaller corrections ride along: the manifest now names each seat directly, and the script that stamps the methodology version onto manifests was reading the oldest changelog entry instead of the newest, so runs made under v1.1 through v1.5 are labelled "1.0" in their manifests. The stamper is fixed; the earlier labels stay as written, because a manifest is a record of what ran, including its mistakes.

  5. v1.5

    Reviewer model refresh

    Updated the Claude panel seat from Fable 5 to Fable 5.1 for new review runs.

    What changed

    • New review runs use Claude Fable 5.1 with the pinned model ID claude-fable-5-1.
    • Historical review artifacts retain the model identity recorded when they ran.
    Full change note

    The Claude panel seat moved from Fable 5 to Fable 5.1. The pinned reviewer command now uses claude-fable-5-1. Historical review artifacts remain unchanged.

  6. v1.4

    Verdict semantics and freshness

    Tightened the rule for "Partially supported" findings and added a freshness check before publication.

    What changed

    • A Partially supported finding must name the part of the normalized proposition that fails. Evidence-quality doubts affect confidence, not the finding.
    • The e-bus procurement claim was re-run against archived evidence under the tightened rule.
    • Freshness and completeness are now a third publication-gate audit, including a retroactive pass on the four published stories.
    Full change note

    Two changes from the external review's P0 findings. (1) "Partially supported" now structurally requires naming which part of the normalized proposition fails; evidence-quality doubts (untested allegations, litigation figures) affect confidence, never the finding. The one published claim decided under the looser reading (e-bus procurement failure) is re-run against its archived evidence under the tightened rule; any change ships as a logged verdict-change. (2) The publication gate gains a third audit: freshness & completeness (prompts/freshness-audit.md) — an independent re-search for newer or stronger sources and material events between evidence gathering and publication, run before every publication and retroactively on the four published stories.

  7. v1.3

    Synthesis basis and vocabulary

    Moved synthesis to independent round-one verdicts, replaced canonical confidence with panel agreement, and opened a quality record.

    What changed

    • The synthesis matrix now uses locked round-one verdicts; round two remains an error-documentation channel.
    • Panel agreement (Unanimous, Adjacent, Split) replaces canonical confidence; per-reviewer confidence remains in AI review.
    • The cautious synthesis rule and its reason are now published, and the panel quality record is open.
    Full change note

    Four decisions from the 2026-09-01 methodology panel, resolved with written second and third opinions (dispositions in methodology/reviews/2026-09-01-panel/DISPOSITIONS.md). (1) Synthesis now consumes the locked ROUND-1 verdicts. Round 1 is the only round in which the three reviewers are independent; synthesising from round 2, where each has read the other two, treated a deliberative round as the vote. Round 2 keeps its job as an error-documentation channel: every final position is carried into synthesis.json as round2_positions and rendered, and a material catch there (fabricated citation, wrong evidence) triggers a fresh blind re-run of the affected claim rather than a correction inside the run. Verified before adoption and re-verified on switching: the round-1 and round-2 multisets give identical canonical findings on all six published claims, so no published verdict changed. (2) The canonical per-claim "confidence" is replaced by PANEL AGREEMENT — Unanimous, Adjacent or Split, computed from the round-1 multiset. Confidence named something the method never computed and read as a probability that the claim was true; agreement describes the panel and is published with a gloss saying so. Per-reviewer confidence stays visible in the AI review, and the matrix no longer emits a canonical confidence at all. The review.agreement value "Majority" becomes "Adjacent" so there is one vocabulary. (3) The matrix keeps its cautious lean and now publishes the reason: a qualification found by one reviewer does not disappear because two missed it, and for a fact-checking site overclaiming is the costlier error. The veto objection is answered by disclosure — the vote composition is always displayed — not by averaging. (4) A panel quality record opens at methodology/quality-ledger.yaml: adjudicated per-seat events (fabricated citations, false accusations, unsupported figures, self-corrections, catches that held up), each traceable to a committed artifact, seeded retroactively from the four published runs with the per-seat denominator stated. It is an error record, not a calibration; a summary publishes when the nine-story slate completes. Reviewer prompts, the merge, and the 20 matrix rows' findings are unchanged.

  8. v1.2

    Artifact hygiene and honesty

    Made evidence and review artifacts more explicit, and aligned the public site with the delegated publication gate.

    What changed

    • Evidence items now show fetch status and registry ID, so unverifiable citations are visible.
    • Panel identity comes from the run manifest; known artifact errors get adjacent errata files.
    • Reviewers now flag framing problems, and story pages link to the brief and gate reports.
    • Composite paraphrases are labelled as paraphrases.
    Full change note

    Changes from the first methodology/trust panel review (three models reviewing the process, not the stories; full reports in methodology/reviews/2026-09-01-panel/). (1) Evidence items in combined-evidence artifacts now carry an explicit fetch status and registry id — unverifiable citations are flagged, never silent. (2) Panel identity is rendered from the run manifest, never from a model's self-report. (3) Raw artifacts known to contain errors get an errata.md beside them; the record is annotated, never edited. (4) Reviewer prompt now instructs reviewers to flag material framing problems in the brief; a material framing flag halts synthesis until the brief is revised and round 1 rerun. (5) Site copy reconciled with the v1.1 delegated gate everywhere; gate audit reports and the frozen brief are linked from each story's AI-review section. (6) "Claims we're seeing" drops social-platform styling; composite paraphrases are visually labeled as paraphrases. Research, merge, cross-review, and synthesis rules are unchanged.

  9. v1.1

    Publication gate

    Delegated the final publication audit to a separate AI audit pair while keeping founder accountability.

    What changed

    • A source-verification audit checks every published statement against archived source bytes.
    • A privacy and copyright release check covers raw review artifacts.
    • Both reports are committed with the run; the founder remains accountable for everything published.
    Full change note

    Publication gate delegated: the founder's per-story manual review is replaced by a dedicated AI audit pair — source verification of every published statement against archived source bytes, and a privacy/copyright release check of raw review artifacts — with both reports committed alongside the run. The founder holds a standing delegation, remains accountable for everything published, and can revert any publication. Research, merge, cross-review, and synthesis stages are unchanged.

  10. v1.0

    Initial

    Established the first public methodology: three independent reviewers, deterministic synthesis, and a human publication gate.

    What changed

    • Three reviewers work independently in blind round one.
    • A deterministic merge and multiset synthesis matrix produce canonical findings.
    • A cross-review round documents errors, drafting gets a two-model faithfulness check, and a human publication gate closes the process.
    Full change note

    Initial methodology: three-reviewer panel (Claude Fable 5, GPT-5.6 Sol, Gemini 3.1 Pro via CLI), blind round 1, deterministic merge, cross-review round 2, deterministic multiset synthesis matrix, drafting with two-model faithfulness check, human publication gate. Reviewer verdicts: Supported / Partially supported / Not established / Contradicted. Canonical findings add Mixed (split panel only).