diff options
| author | sillylaird <sillyfanboy@gmail.com> | 2026-09-03 00:33:59 +0000 |
|---|---|---|
| committer | sillylaird <sillyfanboy@gmail.com> | 2026-09-03 00:33:59 +0000 |
| commit | 898b52edcb47bcb3e9d6106e74ca73e74ea01e70 (patch) | |
| tree | 85c6ee5ad58b860144551184d4cf86b560c62b91 /.agents/skills/results-analysis/examples/example-stats-appendix.md | |
| download | www-main.tar.gz www-main.zip | |
Diffstat (limited to '.agents/skills/results-analysis/examples/example-stats-appendix.md')
| -rw-r--r-- | .agents/skills/results-analysis/examples/example-stats-appendix.md | 36 |
1 files changed, 36 insertions, 0 deletions
diff --git a/.agents/skills/results-analysis/examples/example-stats-appendix.md b/.agents/skills/results-analysis/examples/example-stats-appendix.md new file mode 100644 index 0000000..79ecbfc --- /dev/null +++ b/.agents/skills/results-analysis/examples/example-stats-appendix.md @@ -0,0 +1,36 @@ +# Statistical Appendix + +## Primary Metric +- Word Error Rate (WER), lower is better. + +## Sample Structure +- Unit of analysis: seed-level run, paired by shared data split and subject pool. +- Number of seeds per condition: 5. + +## Descriptive Statistics + +| Condition | Mean WER | Std | 95% CI | +|---|---:|---:|---:| +| Full fine-tuning | 31.4 | 1.9 | [29.8, 33.0] | +| Subject Adapter | 33.2 | 1.3 | [32.1, 34.3] | +| Frozen Encoder | 37.0 | 2.1 | [35.1, 38.9] | + +## Assumption Checks +- Shapiro-Wilk on paired differences: `p = 0.19` +- No strong evidence against normality for the primary contrast. +- Given small n, interpretation remains conservative. + +## Inferential Tests + +| Contrast | Test | Statistic | p | Effect size | Correction | +|---|---|---|---:|---:|---| +| Subject Adapter vs Frozen Encoder | paired t-test | `t(4) = -4.11` | 0.014 | Cohen's `d = 1.84` | Holm | +| Full fine-tuning vs Subject Adapter | paired t-test | `t(4) = -2.03` | 0.112 | Cohen's `d = 0.91` | Holm | + +## Interpretation Guardrails +- First contrast is supported after correction. +- Second contrast trends in favor of full fine-tuning but is not conclusive at this sample size. + +## Blockers / Limits +- No subject-level bootstrap yet. +- No calibration analysis available. |
