{
  "version": 1,
  "disclosure": "Fictional demonstration, not customer results. Generated skills, first drafts, and separate model review outputs are preserved unchanged. Evaluator notes flag remaining errors. Human review is pending.",
  "provenance": {
    "date": "2026-09-21",
    "model": "claude-sonnet-4-6",
    "runtime": "Node v24.17.0 native fetch; Anthropic Messages API",
    "method": "Two synthetic authors, six learning samples and two reserved samples each. One builder call generated both skills. Separate baseline and kit calls generated the same six briefs in fresh API contexts, three per author. The baseline received the same learning writing with an ordinary voice request; the kit received generated skills and their learning evidence. Held-out writing was absent from all generation requests. A separate review call per author received the actual voice-check instructions, that author's generated skill and evidence, the same briefs, and both original branches. The final review response schema puts factual assessment before revision fields. Model settings were the same throughout. All output strings are preserved exactly. Agent evaluator notes are separate.",
    "humanReview": "pending",
    "settings": {
      "temperature": 0.2,
      "maxOutputTokens": 8192,
      "responseFormat": "JSON schema structured output"
    },
    "rawOutputHashes": {
      "builder": "937330d15a1ef75d484ad77f7e87603037877c1f6e7508084dcfdf30bae5e77c",
      "baseline": "695cfccfe7beee43bc1170e6c7ff5979abb3dddad0506c0aaa685d09e99e869b",
      "kit": "314579aacbe071fc89a8461a4c82fe181812ae738620bbd2d43c0481c9857880",
      "review-alex": "b19f8b96ef4168e5a644f606c3782b4b84473a10b2f44de186d3d5c756a7fc94",
      "review-casey": "f7ce273e48549d8d8989add667130234e13c8e084764a54b8b9a39f573e6b4dc"
    },
    "inputHashes": {
      "builder": "9cdfb9da4f2168066483a61986d91e913bbd25e70c637caaaea702701c9100df",
      "baseline": "3dab61e19ec6b19f92170184b8a62f57d0deef31a6e764fe5ed63ea2de7c082f",
      "kit": "7e95197b0c81ae77f1dc87594f185c70472912c6083159c706f6aa9045874175",
      "review-alex": "e04e07a918f5907051d2644e0468c8aa7c44eba076173c6ed94fe598a35ae960",
      "review-casey": "eb46bcb7ce62d4ca2eb4a4cae8d119401ef0c4fa2a7e7eef5f47fa75cbfea6e3"
    },
    "requestFile": "/products/voice-mining-kit/demo/requests.json",
    "runs": {
      "builder": {
        "startedAt": "2026-09-21T08:29:53.464Z",
        "finishedAt": "2026-09-21T08:31:00.909Z",
        "reportedModel": "claude-sonnet-4-6",
        "transportNormalization": "none",
        "usage": {
          "input_tokens": 4855,
          "cache_creation_input_tokens": 0,
          "cache_read_input_tokens": 0,
          "cache_creation": {
            "ephemeral_5m_input_tokens": 0,
            "ephemeral_1h_input_tokens": 0
          },
          "output_tokens": 4002,
          "service_tier": "standard",
          "inference_geo": "global"
        }
      },
      "baseline": {
        "startedAt": "2026-09-21T08:24:06.773Z",
        "finishedAt": "2026-09-21T08:24:15.923Z",
        "reportedModel": "claude-sonnet-4-6",
        "transportNormalization": "none",
        "usage": {
          "input_tokens": 1821,
          "cache_creation_input_tokens": 0,
          "cache_read_input_tokens": 0,
          "cache_creation": {
            "ephemeral_5m_input_tokens": 0,
            "ephemeral_1h_input_tokens": 0
          },
          "output_tokens": 402,
          "service_tier": "standard",
          "inference_geo": "global"
        }
      },
      "kit": {
        "startedAt": "2026-09-21T08:31:35.426Z",
        "finishedAt": "2026-09-21T08:31:49.077Z",
        "reportedModel": "claude-sonnet-4-6",
        "transportNormalization": "none",
        "usage": {
          "input_tokens": 4983,
          "cache_creation_input_tokens": 0,
          "cache_read_input_tokens": 0,
          "cache_creation": {
            "ephemeral_5m_input_tokens": 0,
            "ephemeral_1h_input_tokens": 0
          },
          "output_tokens": 366,
          "service_tier": "standard",
          "inference_geo": "global"
        }
      },
      "review-alex": {
        "startedAt": "2026-09-21T08:41:15.823Z",
        "finishedAt": "2026-09-21T08:41:38.465Z",
        "reportedModel": "claude-sonnet-4-6",
        "transportNormalization": "none",
        "usage": {
          "input_tokens": 4527,
          "cache_creation_input_tokens": 0,
          "cache_read_input_tokens": 0,
          "cache_creation": {
            "ephemeral_5m_input_tokens": 0,
            "ephemeral_1h_input_tokens": 0
          },
          "output_tokens": 1240,
          "service_tier": "standard",
          "inference_geo": "global"
        }
      },
      "review-casey": {
        "startedAt": "2026-09-21T08:41:15.833Z",
        "finishedAt": "2026-09-21T08:41:40.756Z",
        "reportedModel": "claude-sonnet-4-6",
        "transportNormalization": "none",
        "usage": {
          "input_tokens": 4362,
          "cache_creation_input_tokens": 0,
          "cache_read_input_tokens": 0,
          "cache_creation": {
            "ephemeral_5m_input_tokens": 0,
            "ephemeral_1h_input_tokens": 0
          },
          "output_tokens": 1248,
          "service_tier": "standard",
          "inference_geo": "global"
        }
      }
    },
    "revisionHistory": [
      "The installed Codex CLI could not run its configured model. This demonstration used the Anthropic Messages API instead.",
      "Initial text responses were rejected by the harness for JSON fences before their prose was retained. Structured output was then enabled. This changed response formatting, not voice instructions.",
      "The first saved builder output invented an executable and a missing reference. An initial kit draft also added an unsupported follow-up promise. The kit instructions were corrected once and builder plus kit generation were repeated. Original saved outputs remain private development evidence.",
      "The first separate review pass per author returned the original text unchanged and claimed an edit that was absent from the returned draft.",
      "The checking instructions were clarified to preserve the original separately, restore missing required facts, remove unsupported commitments, and verify claimed edits. A second review pass on unchanged originals still left those factual corrections unapplied.",
      "The structured response schema was then reordered to assess facts before generating revised drafts. The checking skill also explicitly states that order. A third review pass used unchanged originals, the same samples, and the same model/settings. It removed the kit sharing promise and restored the kit colour status. The Casey baseline promise remains unresolved.",
      "This publication contains the final workflow outputs, not the first-ever outputs. The free baseline generation stayed unchanged. Earlier builder, draft, and review outputs remain private development evidence. All generated output strings are unedited, with evaluation notes kept separately. No further model runs were made."
    ],
    "evaluation": {
      "reviewer": "Codex agent",
      "rubric": [
        "Presence of each required brief fact and request",
        "Unsupported facts, promises, links, costs, or dates",
        "Greeting, explanation, vocabulary, paragraph structure, and closing compared with reserved examples",
        "Generated rule citations and command/reference validity"
      ],
      "limitations": [
        "Synthetic writing from only two registers; six learning samples per author.",
        "Both authors shared a builder request and all six briefs shared each draft-generation request. Review calls were separate by author. This differs from building each author in a separate conversation.",
        "Instructions and response ordering were revised after development failures and checked on the same cases. This is a workflow demonstration and regression check, not a blinded benchmark.",
        "Zero mechanical checks were enabled because the synthetic authors supplied no approved bans. A contrasting casual control also passes.",
        "The reviewed kit drafts retain the required facts on agent inspection. One reviewed baseline retains an unsupported promise, and the generated skills still contain overgeneralized observations.",
        "Agent factual review is provisional. No human preference test or customer voice-match measurement was performed.",
        "No superiority, accuracy percentage, or universal benefit is claimed."
      ]
    },
    "authorEdits": [],
    "requestFileDescription": "Public fictional input summary and hashes. Exact request bodies, including paid kit instructions, are supplied in the buyer ZIP.",
    "reviewRegression": {
      "status": "kit corrections verified; one baseline correction remains",
      "checks": [
        {
          "case": "alex-review",
          "assertion": "Reviewed kit removes the unsupplied checklist-sharing promise",
          "passed": true
        },
        {
          "case": "alex-review",
          "assertion": "Reviewed baseline removes the unsupplied promise to be in touch",
          "passed": true
        },
        {
          "case": "casey-review",
          "assertion": "Reviewed kit restores the required print-colour status",
          "passed": true
        },
        {
          "case": "casey-review",
          "assertion": "Reviewed baseline removes the unsupplied promise to be in touch",
          "passed": false
        }
      ],
      "interpretation": "These assertions inspect returned strings. Claims of an edit do not count as a correction. These targeted checks are not a general fact validator."
    }
  },
  "authors": [
    {
      "id": "alex",
      "label": "Alex / formal",
      "register": "client-email",
      "learningSamples": [
        {
          "id": "A1",
          "text": "Hello [client],\n\nThank you for sending the revised schedule. The Thursday delivery date works for our review, provided that the accessibility notes arrive with the prototype. Please let me know if that changes.\n\nKind regards,\nAlex"
        },
        {
          "id": "A2",
          "text": "Dear [client],\n\nThe attached document includes the requested cost breakdown. I have separated the initial setup from the recurring costs so that you can compare the two options accurately. Please confirm which option you would like us to develop.\n\nKind regards,\nAlex"
        },
        {
          "id": "A3",
          "text": "Hello [client],\n\nI have reviewed the test results and am comfortable proceeding with the smaller release. Two issues remain in the reporting view, although neither affects the submission process. A wider release should wait until those issues are resolved.\n\nKind regards,\nAlex"
        },
        {
          "id": "A5",
          "text": "Hello [client],\n\nThe draft is ready for your review. Please focus on the examples in the second section, where your operational knowledge will be particularly helpful. Comments received by Tuesday can be included in the version prepared for the workshop.\n\nKind regards,\nAlex"
        },
        {
          "id": "A6",
          "text": "Dear [client],\n\nI apologise for the incorrect figure in yesterday's summary. The corrected total is 42 records, and the attached report explains how the duplicate entries were removed. Your existing export remains unaffected.\n\nKind regards,\nAlex"
        },
        {
          "id": "A7",
          "text": "Hello [client],\n\nYour suggestion would make the instructions easier to follow. I propose adding a worked example immediately after the setup section, then checking it with someone who has not used the system before. Please let me know whether you would prefer a different example.\n\nKind regards,\nAlex"
        }
      ],
      "holdoutCount": 2,
      "skillExcerpt": "- **Greeting:** \"Hello [client],\" (A1, A3, A5, A7) or \"Dear [client],\" (A2, A6). Both appear; choose based on prior relationship cues in the brief.\n- **Signoff:** \"Kind regards,\\nAlex\" in every sample (A1-A3, A5-A7).\n- **Paragraph count:** Typically one body paragraph of two to four sentences (all samples).",
      "generatedSkill": "---\nname: alex-voice\ndescription: Use when drafting client emails in Alex's voice. Supports the client-email register only.\n---\n\n# alex-voice\n\nThis skill helps draft client emails that match Alex's confirmed writing habits. It covers one register: **client-email**. If the task is a different register, state that uncertainty before drafting.\n\n## 1. Read the evidence and choose the register\n\nOpen `references/evidence.md` before drafting. It lists the six learning samples (A1-A3, A5-A7), measured counts, and limitations. All samples are client-email. If a new task is materially different, ask a focused question rather than guessing.\n\n## 2. Follow explicit preferences and observed habits\n\nNo explicit preferences were recorded. Apply these observed habits from the learning samples:\n\n- **Greeting:** \"Hello [client],\" (A1, A3, A5, A7) or \"Dear [client],\" (A2, A6). Both appear; choose based on prior relationship cues in the brief.\n- **Signoff:** \"Kind regards,\\nAlex\" in every sample (A1-A3, A5-A7).\n- **Paragraph count:** Typically one body paragraph of two to four sentences (all samples).\n- **Sentence length:** Median segment is 7 words; maximum observed is 23 words. Sentences are complete and moderately formal.\n- **Structure:** Opens by acknowledging the situation or delivering the main point, adds necessary detail or a condition, closes with a clear next step or invitation to respond.\n- **Tone:** Formal but direct. Apologies are brief and factual (A6). Suggestions are framed as proposals (A7).\n- **First-person:** Uses \"I\" and \"I have\" naturally (A2, A3, A6, A7).\n- A factual brief may require longer sentences or unfamiliar words; allow that without forcing shorter phrasing.\n\n## 3. Use only facts from the brief\n\nDo not transfer names, figures, links, dates, or commitments from the learning samples into a new draft. Flag any required fact that the brief does not supply. A promise to follow up or send an update is a commitment; include it only when the brief explicitly supports it.\n\n## 4. Produce a draft without sending it\n\nReturn the draft text only. Writing style does not grant permission to send, publish, or take any external action.\n\n## 5. Check the draft\n\nAfter drafting, invoke the installed `voice-check` skill with the draft path and `ban-words.json` path from this skill's folder. The mechanical pass covers enabled rules only; see `ban-words.json` for the current config. Zero checks are enabled for Alex; the mechanical pass will not flag anything.\n\n**Manual review procedure:**\n1. Compare the draft greeting, body structure, and signoff against samples A1-A3, A5-A7 in `references/evidence.md`.\n2. Verify every fact and commitment against the brief.\n3. Ask Alex which wording sounds right before sending.\n",
      "generatedEvidence": "# Alex voice evidence\n\n## Scope and limitations\n\n- Owner: Alex\n- Register covered: client-email\n- Learning samples: 6 (IDs A1, A2, A3, A5, A7, A6)\n- Holdout samples: 2 (not used in rule derivation; reserved for calibration)\n- Independent sources: 8\n- Explicit preferences recorded: none\n- Approximate English mechanics only. Counts describe these samples, not hard writing limits. Author calibration is still required.\n\n## Measured counts (client-email, learning samples)\n\n| Metric | Value |\n|---|---|\n| Total words | 247 |\n| Approximate segments | 30 |\n| Median segment words | 7.0 |\n| Maximum segment words | 23 |\n| Complete samples | 6 of 6 |\n\n## Observed habits\n\nAll observations are tentative given the small sample size. Author confirmation is required.\n\n- **Greeting variants:** \"Hello [client],\" appears in A1, A3, A5, A7. \"Dear [client],\" appears in A2, A6. Both are established; neither is dominant enough to treat as the sole default.\n- **Signoff:** \"Kind regards,\\nAlex\" appears in all six samples (A1, A2, A3, A5, A6, A7).\n- **Body length:** One paragraph per message in all six samples. Two to four sentences per paragraph.\n- **Sentence structure:** Complete declarative sentences. Conditions stated explicitly (\"provided that\", A1). Suggestions framed as proposals (\"I propose\", A7).\n- **Apology style:** Brief, factual, followed immediately by the corrected information (A6).\n- **Closing move:** Each message ends with an invitation to respond or a clear next step (A1, A2, A3, A5, A6, A7).\n- **Capitalisation:** Standard sentence case throughout.\n- **No observed use of contractions** in any learning sample.\n\n## Redacted learning samples\n\n### A1 (fictional-thread-a1, client-email, learn)\n```\nHello [client],\n\nThank you for sending the revised schedule. The Thursday delivery date works for our review, provided that the accessibility notes arrive with the prototype. Please let me know if that changes.\n\nKind regards,\nAlex\n```\n\n### A2 (fictional-thread-a2, client-email, learn)\n```\nDear [client],\n\nThe attached document includes the requested cost breakdown. I have separated the initial setup from the recurring costs so that you can compare the two options accurately. Please confirm which option you would like us to develop.\n\nKind regards,\nAlex\n```\n\n### A3 (fictional-thread-a3, client-email, learn)\n```\nHello [client],\n\nI have reviewed the test results and am comfortable proceeding with the smaller release. Two issues remain in the reporting view, although neither affects the submission process. A wider release should wait until those issues are resolved.\n\nKind regards,\nAlex\n```\n\n### A5 (fictional-thread-a5, client-email, learn)\n```\nHello [client],\n\nThe draft is ready for your review. Please focus on the examples in the second section, where your operational knowledge will be particularly helpful. Comments received by Tuesday can be included in the version prepared for the workshop.\n\nKind regards,\nAlex\n```\n\n### A6 (fictional-thread-a6, client-email, learn)\n```\nDear [client],\n\nI apologise for the incorrect figure in yesterday's summary. The corrected total is 42 records, and the attached report explains how the duplicate entries were removed. Your existing export remains unaffected.\n\nKind regards,\nAlex\n```\n\n### A7 (fictional-thread-a7, client-email, learn)\n```\nHello [client],\n\nYour suggestion would make the instructions easier to follow. I propose adding a worked example immediately after the setup section, then checking it with someone who has not used the system before. Please let me know whether you would prefer a different example.\n\nKind regards,\nAlex\n```\n\n## Warnings\n\nNone reported by corpus_metrics.py.\n\n## Calibration status\n\nNot yet calibrated. Two holdout samples are reserved. Follow the calibration procedure in the builder skill before treating this skill as validated.\n",
      "checkerConfig": {
        "_comment": "Neutral starting profile. Enable only author-approved checks and record why. The editorial preset is optional.",
        "use_default_banned_phrases": false,
        "extra_banned_phrases": [],
        "disable": [
          "EM-DASH",
          "BANNED-PHRASE",
          "ANAPHORA",
          "STACCATO-NEGATION",
          "ANTITHESIS",
          "FLAT-RHYTHM",
          "SHORT-BEAT-RUN",
          "TRIPLET"
        ]
      },
      "ruleReview": {
        "status": "corrections needed before reuse",
        "reviewer": "Codex agent",
        "notes": [
          "The generated skill preserves the observed formal greetings, full sentences, one body paragraph, and Kind regards closing.",
          "Its evidence overstates the closing pattern: A6 ends with reassurance about an export, not an invitation or next step. The proposed greeting choice by relationship is not established by the samples.",
          "The reserved A4 and A8 also use formal greetings, one body paragraph, and the same closing. This is an agent observation on synthetic writing, not an author verdict."
        ]
      },
      "cases": [
        {
          "id": "alex-review",
          "label": "Project update",
          "brief": "Write an email to a client named Morgan. The revised handover notes are ready. Ask Morgan to confirm the contact details by Thursday. The access checklist is still being reviewed. Do not imply that a file is attached or promise when the checklist will be complete.",
          "baselineDraft": "Hello Morgan,\n\nThe revised handover notes are ready for your review. Please confirm that your contact details are correct by Thursday.\n\nThe access checklist is still being reviewed, and I will be in touch once that process is complete.\n\nKind regards,\nAlex",
          "kitDraft": "Hello Morgan,\n\nThe revised handover notes are ready for your review. Please confirm that the contact details are correct by Thursday. The access checklist is still being reviewed and I will share it once that process is complete.\n\nKind regards,\nAlex",
          "review": {
            "factualStatus": "needs revision",
            "reviewer": "Codex agent",
            "notes": [
              "Both drafts retain the ready status, Thursday request, and ongoing access review.",
              "The baseline adds a promise to be in touch. The kit draft adds a promise to share the checklist. Neither commitment was supplied in the brief."
            ]
          },
          "reviewedBaselineDraft": "Hello Morgan,\n\nThe revised handover notes are ready for your review. Please confirm that your contact details are correct by Thursday.\n\nThe access checklist is still being reviewed. Please let me know if you have any questions in the meantime.\n\nKind regards,\nAlex",
          "reviewedKitDraft": "Hello Morgan,\n\nThe revised handover notes are ready for your review. Please confirm that the contact details are correct by Thursday. The access checklist is still being reviewed.\n\nKind regards,\nAlex",
          "reviewChanges": [
            "The kit review removes the unsupplied promise to share the checklist. The required ready status, Thursday request, and ongoing review remain.",
            "The baseline review removes the promise to be in touch and adds an invitation for questions. Both reviewed drafts retain the required facts.",
            "The reviewer also claims to merge the kit body into one paragraph, although the original kit body already used one paragraph. The saved drafts show the actual edit."
          ],
          "modelReview": {
            "id": "alex-review",
            "requiredFacts": [
              {
                "fact": "Revised handover notes are ready",
                "baselinePresent": true,
                "kitPresent": true
              },
              {
                "fact": "Ask Morgan to confirm contact details by Thursday",
                "baselinePresent": true,
                "kitPresent": true
              },
              {
                "fact": "Access checklist is still being reviewed",
                "baselinePresent": true,
                "kitPresent": true
              },
              {
                "fact": "Do not imply a file is attached",
                "baselinePresent": true,
                "kitPresent": true
              },
              {
                "fact": "Do not promise when the checklist will be complete",
                "baselinePresent": false,
                "kitPresent": false
              }
            ],
            "remainingIssues": [
              "Baseline: 'I will be in touch once that process is complete' implies a promise of follow-up timing tied to completion, which the brief does not support. The brief says do not promise when the checklist will be complete; this phrasing edges toward that commitment. Revised to remove the implied timing promise.",
              "Kit: 'I will share it once that process is complete' explicitly promises to share the checklist upon completion, which the brief does not support and the brief's restriction covers. Revised to remove this unsupported commitment."
            ],
            "changes": [
              "Baseline: Removed 'I will be in touch once that process is complete' and replaced with a neutral closing invitation to respond, removing the unsupported implied timing commitment.",
              "Kit: Removed 'I will share it once that process is complete' and replaced with a neutral closing, removing the unsupported commitment to share the checklist. Merged into one paragraph to match the one-paragraph body pattern observed in all six learning samples (A1-A7). No other phrasing changes made."
            ],
            "reviewedBaselineDraft": "Hello Morgan,\n\nThe revised handover notes are ready for your review. Please confirm that your contact details are correct by Thursday.\n\nThe access checklist is still being reviewed. Please let me know if you have any questions in the meantime.\n\nKind regards,\nAlex",
            "reviewedKitDraft": "Hello Morgan,\n\nThe revised handover notes are ready for your review. Please confirm that the contact details are correct by Thursday. The access checklist is still being reviewed.\n\nKind regards,\nAlex"
          },
          "reviewedFactualStatus": "required facts present",
          "reviewedBaselineFactualStatus": "required facts present",
          "reviewedBaselineNotes": []
        },
        {
          "id": "alex-correction",
          "label": "Correction",
          "brief": "Write an email to a client named Taylor. Correct the total in your previous summary: it is 38 records, not 41. The export itself is correct and remains unchanged. Apologise for the error. Do not invent the cause or promise a new report.",
          "baselineDraft": "Dear Taylor,\n\nI apologise for the incorrect figure in my previous summary. The correct total is 38 records, not 41. Your export is accurate and remains unchanged.\n\nKind regards,\nAlex",
          "kitDraft": "Dear Taylor,\n\nI apologise for the incorrect figure in my previous summary. The corrected total is 38 records, not 41. Your existing export is correct and remains unchanged.\n\nKind regards,\nAlex",
          "review": {
            "factualStatus": "required facts present",
            "reviewer": "Codex agent",
            "notes": [
              "Both drafts preserve the correction from 41 to 38, the unchanged export, and an apology. Neither invents a cause or new report."
            ]
          },
          "reviewedBaselineDraft": "Dear Taylor,\n\nI apologise for the incorrect figure in my previous summary. The correct total is 38 records, not 41. Your export is accurate and remains unchanged.\n\nKind regards,\nAlex",
          "reviewedKitDraft": "Dear Taylor,\n\nI apologise for the incorrect figure in my previous summary. The corrected total is 38 records, not 41. Your existing export is correct and remains unchanged.\n\nKind regards,\nAlex",
          "reviewChanges": [
            "Both reviewed drafts are unchanged. Required facts remain present."
          ],
          "modelReview": {
            "id": "alex-correction",
            "requiredFacts": [
              {
                "fact": "Apologise for the error",
                "baselinePresent": true,
                "kitPresent": true
              },
              {
                "fact": "Correct total is 38 records, not 41",
                "baselinePresent": true,
                "kitPresent": true
              },
              {
                "fact": "Export is correct and remains unchanged",
                "baselinePresent": true,
                "kitPresent": true
              },
              {
                "fact": "Do not invent the cause of the error",
                "baselinePresent": true,
                "kitPresent": true
              },
              {
                "fact": "Do not promise a new report",
                "baselinePresent": true,
                "kitPresent": true
              }
            ],
            "remainingIssues": [],
            "changes": [
              "No factual or mechanical issues found in either draft. No phrasing changes made. Both drafts satisfy all brief requirements and align with the observed voice habits (Dear greeting, brief factual apology as in A6, signoff, one paragraph, no contractions, no invented cause, no promised report)."
            ],
            "reviewedBaselineDraft": "Dear Taylor,\n\nI apologise for the incorrect figure in my previous summary. The correct total is 38 records, not 41. Your export is accurate and remains unchanged.\n\nKind regards,\nAlex",
            "reviewedKitDraft": "Dear Taylor,\n\nI apologise for the incorrect figure in my previous summary. The corrected total is 38 records, not 41. Your existing export is correct and remains unchanged.\n\nKind regards,\nAlex"
          },
          "reviewedFactualStatus": "required facts present",
          "reviewedBaselineFactualStatus": "required facts present",
          "reviewedBaselineNotes": []
        },
        {
          "id": "alex-scope",
          "label": "Scope change",
          "brief": "Write an email to a client named Jordan. Their requested second integration would need a separate estimate. Ask for a representative test file before preparing the estimate. Do not quote a cost or give a delivery date.",
          "baselineDraft": "Hello Jordan,\n\nThank you for raising the second integration. That work would need a separate estimate, and I would like to prepare it as accurately as possible. Could you send a representative test file so that I can assess what is involved?\n\nKind regards,\nAlex",
          "kitDraft": "Hello Jordan,\n\nThe second integration you have requested would require a separate estimate. I propose that you send a representative test file so that I can prepare that estimate accurately. Please let me know if you have any questions in the meantime.\n\nKind regards,\nAlex",
          "review": {
            "factualStatus": "required facts present",
            "reviewer": "Codex agent",
            "notes": [
              "Both drafts ask for a representative test file before an estimate for the second integration. Neither quotes a cost or delivery date.",
              "The kit adds an invitation for questions. It does not add a delivery commitment."
            ]
          },
          "reviewedBaselineDraft": "Hello Jordan,\n\nThank you for raising the second integration. That work would need a separate estimate, and I would like to prepare it as accurately as possible. Could you send a representative test file so that I can assess what is involved?\n\nKind regards,\nAlex",
          "reviewedKitDraft": "Hello Jordan,\n\nThe second integration you have requested would require a separate estimate. I propose that you send a representative test file so that I can prepare that estimate accurately. Please let me know if you have any questions in the meantime.\n\nKind regards,\nAlex",
          "reviewChanges": [
            "Both reviewed drafts are unchanged. Required facts remain present."
          ],
          "modelReview": {
            "id": "alex-scope",
            "requiredFacts": [
              {
                "fact": "Second integration would need a separate estimate",
                "baselinePresent": true,
                "kitPresent": true
              },
              {
                "fact": "Ask for a representative test file before preparing the estimate",
                "baselinePresent": true,
                "kitPresent": true
              },
              {
                "fact": "Do not quote a cost",
                "baselinePresent": true,
                "kitPresent": true
              },
              {
                "fact": "Do not give a delivery date",
                "baselinePresent": true,
                "kitPresent": true
              }
            ],
            "remainingIssues": [
              "Kit: 'Please let me know if you have any questions in the meantime' is a closing invitation to respond, which is consistent with the observed closing-move habit (A1-A7) and is not a factual commitment. No brief restriction covers it. Retained as acceptable."
            ],
            "changes": [
              "No factual or mechanical issues found in either draft. No phrasing changes made. Both drafts satisfy all brief requirements. Baseline uses a question form for the request; kit uses a proposal form consistent with A7 ('I propose'). Both are within observed habits. No revision warranted."
            ],
            "reviewedBaselineDraft": "Hello Jordan,\n\nThank you for raising the second integration. That work would need a separate estimate, and I would like to prepare it as accurately as possible. Could you send a representative test file so that I can assess what is involved?\n\nKind regards,\nAlex",
            "reviewedKitDraft": "Hello Jordan,\n\nThe second integration you have requested would require a separate estimate. I propose that you send a representative test file so that I can prepare that estimate accurately. Please let me know if you have any questions in the meantime.\n\nKind regards,\nAlex"
          },
          "reviewedFactualStatus": "required facts present",
          "reviewedBaselineFactualStatus": "required facts present",
          "reviewedBaselineNotes": []
        }
      ]
    },
    {
      "id": "casey",
      "label": "Casey / casual",
      "register": "creator-client-email",
      "learningSamples": [
        {
          "id": "C1",
          "text": "hey [client],\n\nthe draft's ready. give the examples a look when you've got a minute.\n\nthanks,\nCasey"
        },
        {
          "id": "C2",
          "text": "hey [client],\n\ni've fixed the broken link. can you try it again?\n\nthanks,\nCasey"
        },
        {
          "id": "C3",
          "text": "hey [client],\n\nthis needs another check. i'll send an update after we've looked at the export.\n\nthanks,\nCasey"
        },
        {
          "id": "C5",
          "text": "hey [client],\n\nthat headline reads better to me. let's try it on the smaller layout before we call it done.\n\nthanks,\nCasey"
        },
        {
          "id": "C6",
          "text": "hey [client],\n\nthe new sketches are in the folder. pick the one you'd like me to take further.\n\nthanks,\nCasey"
        },
        {
          "id": "C7",
          "text": "hey [client],\n\nthis extra page changes the job a bit. can you send the final copy before i work out the extra cost?\n\nthanks,\nCasey"
        }
      ],
      "holdoutCount": 2,
      "skillExcerpt": "- **Greeting:** \"hey [client],\" in every sample (C1, C2, C3, C5, C6, C7). Lowercase throughout.\n- **Signoff:** \"thanks,\\nCasey\" in every sample. \"Casey\" is capitalised; the word \"thanks\" is lowercase.\n- **Body length:** One short paragraph per message in all six samples. Typically one to two sentences.",
      "generatedSkill": "---\nname: casey-voice\ndescription: Use when drafting creator-client emails in Casey's voice. Supports the creator-client-email register only.\n---\n\n# casey-voice\n\nThis skill helps draft client emails that match Casey's confirmed writing habits. It covers one register: **creator-client-email**. If the task is a materially different register, state that uncertainty before drafting.\n\n## 1. Read the evidence and choose the register\n\nOpen `references/evidence.md` before drafting. It lists the six learning samples (C1-C3, C5-C7), measured counts, and limitations. All samples are creator-client-email. If a new task is materially different, ask a focused question rather than guessing.\n\n## 2. Follow explicit preferences and observed habits\n\nNo explicit preferences were recorded. Apply these observed habits from the learning samples:\n\n- **Greeting:** \"hey [client],\" in every sample (C1, C2, C3, C5, C6, C7). Lowercase throughout.\n- **Signoff:** \"thanks,\\nCasey\" in every sample. \"Casey\" is capitalised; the word \"thanks\" is lowercase.\n- **Body length:** One short paragraph per message in all six samples. Typically one to two sentences.\n- **Sentence length:** Median segment is 2.5 words; maximum observed is 13 words. Sentences are brief and conversational.\n- **Tone:** Casual, direct, no formal connectors. Questions are short and action-oriented (C2, C7).\n- **Capitalisation:** Sentence-initial words and the name \"Casey\" are capitalised; most other lines begin lowercase.\n- **Contractions:** Used naturally throughout (\"the draft's\", \"i've\", \"i'll\", \"you've\", \"you'd\", \"i'd\") (C1, C2, C3, C5, C6, C7).\n- **First person:** Lowercase \"i\" throughout all samples.\n- A factual brief may require a longer sentence; allow that without forcing shorter phrasing.\n\n## 3. Use only facts from the brief\n\nDo not transfer names, figures, links, dates, or commitments from the learning samples into a new draft. Flag any required fact that the brief does not supply. A promise to send an update or complete a task is a commitment; include it only when the brief explicitly supports it.\n\n## 4. Produce a draft without sending it\n\nReturn the draft text only. Writing style does not grant permission to send, publish, or take any external action.\n\n## 5. Check the draft\n\nAfter drafting, invoke the installed `voice-check` skill with the draft path and `ban-words.json` path from this skill's folder. The mechanical pass covers enabled rules only; see `ban-words.json` for the current config. Zero checks are enabled for Casey; the mechanical pass will not flag anything.\n\n**Manual review procedure:**\n1. Compare the draft greeting, body length, and signoff against samples C1-C3, C5-C7 in `references/evidence.md`.\n2. Verify every fact and commitment against the brief.\n3. Ask Casey which wording sounds right before sending.\n",
      "generatedEvidence": "# Casey voice evidence\n\n## Scope and limitations\n\n- Owner: Casey\n- Register covered: creator-client-email\n- Learning samples: 6 (IDs C1, C2, C3, C5, C6, C7)\n- Holdout samples: 2 (not used in rule derivation; reserved for calibration)\n- Independent sources: 8\n- Explicit preferences recorded: none\n- Approximate English mechanics only. Counts describe these samples, not hard writing limits. Author calibration is still required.\n\n## Measured counts (creator-client-email, learning samples)\n\n| Metric | Value |\n|---|---|\n| Total words | 116 |\n| Approximate segments | 24 |\n| Median segment words | 2.5 |\n| Maximum segment words | 13 |\n| Complete samples | 6 of 6 |\n\n## Observed habits\n\nAll observations are tentative given the small sample size. Author confirmation is required.\n\n- **Greeting:** \"hey [client],\" in all six samples (C1, C2, C3, C5, C6, C7). Fully lowercase.\n- **Signoff:** \"thanks,\\nCasey\" in all six samples. Consistent across all messages.\n- **Body length:** One paragraph, one to two sentences, in every sample. No multi-paragraph messages observed.\n- **Sentence style:** Short, direct, often a statement followed by a single question or instruction. No formal connectors observed.\n- **Contractions:** Present in every sample. Examples: \"the draft's\" (C1), \"i've\" (C2), \"i'll\" (C3), \"you'd\" (C5), \"you'd\" (C6), \"i\" (C7).\n- **Lowercase first person:** \"i\" is not capitalised in any sample (C1, C2, C3, C5, C6, C7).\n- **Scope questions:** When a job changes, Casey asks for information before committing to cost or timeline (C7).\n- **No apology formula observed** in the six learning samples. Absence is not a ban.\n\n## Redacted learning samples\n\n### C1 (fictional-thread-c1, creator-client-email, learn)\n```\nhey [client],\n\nthe draft's ready. give the examples a look when you've got a minute.\n\nthanks,\nCasey\n```\n\n### C2 (fictional-thread-c2, creator-client-email, learn)\n```\nhey [client],\n\ni've fixed the broken link. can you try it again?\n\nthanks,\nCasey\n```\n\n### C3 (fictional-thread-c3, creator-client-email, learn)\n```\nhey [client],\n\nthis needs another check. i'll send an update after we've looked at the export.\n\nthanks,\nCasey\n```\n\n### C5 (fictional-thread-c5, creator-client-email, learn)\n```\nhey [client],\n\nthat headline reads better to me. let's try it on the smaller layout before we call it done.\n\nthanks,\nCasey\n```\n\n### C6 (fictional-thread-c6, creator-client-email, learn)\n```\nhey [client],\n\nthe new sketches are in the folder. pick the one you'd like me to take further.\n\nthanks,\nCasey\n```\n\n### C7 (fictional-thread-c7, creator-client-email, learn)\n```\nhey [client],\n\nthis extra page changes the job a bit. can you send the final copy before i work out the extra cost?\n\nthanks,\nCasey\n```\n\n## Warnings\n\nNone reported by corpus_metrics.py.\n\n## Calibration status\n\nNot yet calibrated. Two holdout samples are reserved. Follow the calibration procedure in the builder skill before treating this skill as validated.\n",
      "checkerConfig": {
        "_comment": "Neutral starting profile. Enable only author-approved checks and record why. The editorial preset is optional.",
        "use_default_banned_phrases": false,
        "extra_banned_phrases": [],
        "disable": [
          "EM-DASH",
          "BANNED-PHRASE",
          "ANAPHORA",
          "STACCATO-NEGATION",
          "ANTITHESIS",
          "FLAT-RHYTHM",
          "SHORT-BEAT-RUN",
          "TRIPLET"
        ]
      },
      "ruleReview": {
        "status": "corrections needed before reuse",
        "reviewer": "Codex agent",
        "notes": [
          "The generated skill and drafts preserve hey, a casual body, and thanks with a capitalised signature.",
          "Its capitalization rule contradicts the lowercase sample bodies. The evidence also wrongly says contractions appear in every sample and cites you'd in C5. These generated observations need correction before reuse.",
          "The reserved C4 and C8 share the lowercase body and short conversational structure. This is an agent observation on synthetic writing, not an author verdict."
        ]
      },
      "cases": [
        {
          "id": "casey-review",
          "label": "Project update",
          "brief": "Write an email to a client named Sam. The poster draft is ready. Ask Sam to check the headline by Tuesday. The print colours are still under review. Do not promise a print date or invent a link.",
          "baselineDraft": "hey Sam,\n\nthe poster draft is ready. can you check the headline by Tuesday? the print colours are still under review, so i'll be in touch on that separately.\n\nthanks,\nCasey",
          "kitDraft": "hey Sam,\n\nthe poster draft is ready. can you check the headline by Tuesday?\n\nthanks,\nCasey",
          "review": {
            "factualStatus": "needs revision",
            "reviewer": "Codex agent",
            "notes": [
              "The baseline includes the ready poster, Tuesday headline review, and ongoing colour review, but adds an unsupported promise to be in touch.",
              "The kit keeps the ready poster and Tuesday request, but omits the required fact that the print colours are still under review."
            ]
          },
          "reviewedBaselineDraft": "hey Sam,\n\nthe poster draft is ready. can you check the headline by Tuesday? the print colours are still under review, so i'll be in touch on that separately.\n\nthanks,\nCasey",
          "reviewedKitDraft": "hey Sam,\n\nthe poster draft is ready. can you check the headline by Tuesday? the print colours are still under review.\n\nthanks,\nCasey",
          "reviewChanges": [
            "The kit review restores the required fact that print colours are still under review. Its ready status and Tuesday request remain.",
            "The baseline review leaves its unsupported promise to be in touch unchanged. That comparison draft still needs correction, despite the model saying no change was warranted."
          ],
          "modelReview": {
            "id": "casey-review",
            "requiredFacts": [
              {
                "fact": "Client is Sam",
                "baselinePresent": true,
                "kitPresent": true
              },
              {
                "fact": "Poster draft is ready",
                "baselinePresent": true,
                "kitPresent": true
              },
              {
                "fact": "Ask Sam to check the headline by Tuesday",
                "baselinePresent": true,
                "kitPresent": true
              },
              {
                "fact": "Print colours are still under review",
                "baselinePresent": true,
                "kitPresent": false
              },
              {
                "fact": "Do not promise a print date",
                "baselinePresent": true,
                "kitPresent": true
              },
              {
                "fact": "Do not invent a link",
                "baselinePresent": true,
                "kitPresent": true
              }
            ],
            "remainingIssues": [
              "Kit draft omits the required fact that print colours are still under review; this is a brief requirement, not a style choice. The reviewed kit draft restores it."
            ],
            "changes": [
              "Baseline: no changes warranted; all required facts present, voice matches evidence (greeting, signoff, contractions, casual tone, one paragraph).",
              "Kit: restored the print-colours-under-review fact, which the original kit draft omitted. Phrasing kept close to the kit draft's style."
            ],
            "reviewedBaselineDraft": "hey Sam,\n\nthe poster draft is ready. can you check the headline by Tuesday? the print colours are still under review, so i'll be in touch on that separately.\n\nthanks,\nCasey",
            "reviewedKitDraft": "hey Sam,\n\nthe poster draft is ready. can you check the headline by Tuesday? the print colours are still under review.\n\nthanks,\nCasey"
          },
          "reviewedFactualStatus": "required facts present",
          "reviewedBaselineFactualStatus": "needs revision",
          "reviewedBaselineNotes": [
            "The reviewed baseline still adds 'i'll be in touch on that separately', an unsupported follow-up promise. Remove it before using this draft."
          ]
        },
        {
          "id": "casey-correction",
          "label": "Correction",
          "brief": "Write an email to a client named Riley. You sent the wrong thumbnail in your last message. The correct thumbnail is in the shared folder. Ask Riley to use that version. Apologise, but do not invent an excuse or a URL.",
          "baselineDraft": "hey Riley,\n\nsorry about that - i sent the wrong thumbnail in my last message. the correct one is in the shared folder. please use that version going forward.\n\nthanks,\nCasey",
          "kitDraft": "hey Riley,\n\nsorry about that, i sent the wrong thumbnail in my last message. the correct one is in the shared folder, so please use that version.\n\nthanks,\nCasey",
          "review": {
            "factualStatus": "required facts present",
            "reviewer": "Codex agent",
            "notes": [
              "Both drafts apologise, identify the wrong thumbnail, locate the correct one in the shared folder, and ask the recipient to use it. Neither invents an excuse or URL."
            ]
          },
          "reviewedBaselineDraft": "hey Riley,\n\nsorry about that - i sent the wrong thumbnail in my last message. the correct one is in the shared folder. please use that version going forward.\n\nthanks,\nCasey",
          "reviewedKitDraft": "hey Riley,\n\nsorry about that, i sent the wrong thumbnail in my last message. the correct one is in the shared folder, so please use that version.\n\nthanks,\nCasey",
          "reviewChanges": [
            "Both reviewed drafts are unchanged. The reviewer correctly distinguishes a spaced hyphen from an em dash, leaving punctuation preference for the author. Required facts remain present."
          ],
          "modelReview": {
            "id": "casey-correction",
            "requiredFacts": [
              {
                "fact": "Client is Riley",
                "baselinePresent": true,
                "kitPresent": true
              },
              {
                "fact": "Wrong thumbnail was sent in the last message",
                "baselinePresent": true,
                "kitPresent": true
              },
              {
                "fact": "Correct thumbnail is in the shared folder",
                "baselinePresent": true,
                "kitPresent": true
              },
              {
                "fact": "Ask Riley to use that version",
                "baselinePresent": true,
                "kitPresent": true
              },
              {
                "fact": "Apologise",
                "baselinePresent": true,
                "kitPresent": true
              },
              {
                "fact": "Do not invent an excuse",
                "baselinePresent": true,
                "kitPresent": true
              },
              {
                "fact": "Do not invent a URL",
                "baselinePresent": true,
                "kitPresent": true
              }
            ],
            "remainingIssues": [
              "Baseline uses a spaced hyphen ('sorry about that - i sent') which functions as a dash substitute. The brief bans the em dash character specifically; a spaced hyphen is not that character, so this is not a rule violation, but it is worth flagging for Casey's preference. No guidance in the evidence resolves this either way; left for author judgment.",
              "No apology formula is observed in the six learning samples, but the evidence explicitly notes 'absence is not a ban,' and the brief requires an apology. Both drafts handle this appropriately."
            ],
            "changes": [
              "Baseline: no factual or voice changes made. The spaced-hyphen flag is noted for author judgment but not altered without a preference decision.",
              "Kit: no changes warranted; all required facts present, voice matches evidence."
            ],
            "reviewedBaselineDraft": "hey Riley,\n\nsorry about that - i sent the wrong thumbnail in my last message. the correct one is in the shared folder. please use that version going forward.\n\nthanks,\nCasey",
            "reviewedKitDraft": "hey Riley,\n\nsorry about that, i sent the wrong thumbnail in my last message. the correct one is in the shared folder, so please use that version.\n\nthanks,\nCasey"
          },
          "reviewedFactualStatus": "required facts present",
          "reviewedBaselineFactualStatus": "required facts present",
          "reviewedBaselineNotes": []
        },
        {
          "id": "casey-scope",
          "label": "Scope change",
          "brief": "Write an email to a client named Drew. Adding a second poster size changes the job and needs an extra estimate. Ask Drew for the final dimensions before preparing that estimate. Do not invent a price or promise a completion date.",
          "baselineDraft": "hey Drew,\n\nadding a second poster size changes the job a bit, so i'll need to put together a separate estimate for that. can you send me the final dimensions before i get started on it?\n\nthanks,\nCasey",
          "kitDraft": "hey Drew,\n\nadding a second poster size changes the job a bit. can you send the final dimensions before i work out the extra estimate?\n\nthanks,\nCasey",
          "review": {
            "factualStatus": "required facts present",
            "reviewer": "Codex agent",
            "notes": [
              "Both drafts explain the scope change and extra estimate, then request final dimensions. Neither invents a price or completion date."
            ]
          },
          "reviewedBaselineDraft": "hey Drew,\n\nadding a second poster size changes the job a bit, so i'll need to put together a separate estimate for that. can you send me the final dimensions before i get started on it?\n\nthanks,\nCasey",
          "reviewedKitDraft": "hey Drew,\n\nadding a second poster size changes the job a bit. can you send the final dimensions before i work out the extra estimate?\n\nthanks,\nCasey",
          "reviewChanges": [
            "Both reviewed drafts are unchanged. Required facts remain present."
          ],
          "modelReview": {
            "id": "casey-scope",
            "requiredFacts": [
              {
                "fact": "Client is Drew",
                "baselinePresent": true,
                "kitPresent": true
              },
              {
                "fact": "Adding a second poster size changes the job and needs an extra estimate",
                "baselinePresent": true,
                "kitPresent": true
              },
              {
                "fact": "Ask Drew for the final dimensions before preparing the estimate",
                "baselinePresent": true,
                "kitPresent": true
              },
              {
                "fact": "Do not invent a price",
                "baselinePresent": true,
                "kitPresent": true
              },
              {
                "fact": "Do not promise a completion date",
                "baselinePresent": true,
                "kitPresent": true
              }
            ],
            "remainingIssues": [
              "Baseline includes 'so i'll need to put together a separate estimate for that' which is a forward commitment. The brief supports that an estimate is needed, so this is not unsupported; it is a reasonable restatement of the brief's intent. No issue.",
              "Both drafts align well with C7 in the evidence, which covers scope-change questions. No outstanding issues."
            ],
            "changes": [
              "Baseline: no changes warranted; all required facts present, no unsupported commitments, voice matches evidence.",
              "Kit: no changes warranted; all required facts present, voice matches evidence and closely mirrors C7 phrasing naturally."
            ],
            "reviewedBaselineDraft": "hey Drew,\n\nadding a second poster size changes the job a bit, so i'll need to put together a separate estimate for that. can you send me the final dimensions before i get started on it?\n\nthanks,\nCasey",
            "reviewedKitDraft": "hey Drew,\n\nadding a second poster size changes the job a bit. can you send the final dimensions before i work out the extra estimate?\n\nthanks,\nCasey"
          },
          "reviewedFactualStatus": "required facts present",
          "reviewedBaselineFactualStatus": "required facts present",
          "reviewedBaselineNotes": []
        }
      ]
    }
  ]
}
