Produced alongside the img/ chart refresh via scripts/univariate_analysis.py:
per-protocol and per-gradient averages tables, summary plots, and
univariate_summary.txt (mean 5.44, median 5.48, normalized midpoint
deviation 0.11).
Full pipeline per README command flow (steps 2-6): univariate averages and
distributions, multivariate analysis (2-cluster k-means: 209 relational vs
199 institutional; PCA/t-SNE/UMAP/factor/correlation/network/importance),
LDA separation and overlaid cluster plots, and LDA classifications (trained
on the 1.2.6 run via version-matched auto-selection). Cross-version summary
is on the console report: avg Euclidean distance 9.06 vs 1.2.6 (r=0.79).
411 protocols × 23 gradients scored against bicorder.json v1.4.0 with
gpt-oss:20b-cloud (same analyst standpoint as the 1.2.6 run); the
bicorder_version column makes the run self-describing. Analysis outputs to
follow in data/synthetic_1.4.0/analysis/.
Strategy: key runs by bicorder version (not date), promote shared inputs
to analysis/data/, and stamp every output with its bicorder_version so
a re-run on edited gradients is self-describing.
Data layout:
- Promote the shared protocol inputs out of the run directory:
analysis/data/protocols_edited.csv (411 cleaned protocols)
analysis/data/protocols_raw.csv (774 un-cleaned entries)
- Rename the v1.2.6 synthetic run:
data/synthetic_20251116/ -> data/synthetic_1.2.6/
so the bicorder version it was scored against is explicit (gradient
structure changes between versions make date-based names ambiguous)
Provenance:
- bicorder_analyze.py now writes a 'bicorder_version' column into every
output readings.csv, recording which gradient structure produced it
Scripts:
- Update the real code defaults that pointed at the old run path
(bicorder_classifier.py, classify_readings.py, sync_readings.sh,
compare_analyses.py) and refresh docstring/help examples
- Remove a stray committed __pycache__/.pyc
Docs: analysis/README.md documents the new layout + how to add a run;
WORKFLOW.md, TEST_COMMANDS.md, INTEGRATION_GUIDE.md paths updated.
Remove the intermediate readings/ subdirectory level — dataset naming
(synthetic_YYYYMMDD, manual_YYYYMMDD) already encodes what the data is.
Update all path references across scripts and docs accordingly.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- Move all scripts to scripts/, web assets to web/, analysis results
into self-contained data/readings/<type>_<YYYYMMDD>/ directories
- Add data/readings/manual_20260320/ with 32 JSON readings from
git.medlab.host/ntnsndr/protocol-bicorder-data
- Add scripts/json_to_csv.py to convert bicorder JSON files to CSV
- Add scripts/sync_readings.sh for one-command sync + re-analysis of
any dataset backed by a .sync_source config file
- Add scripts/classify_readings.py to apply the LDA classifier to all
readings and save per-reading cluster assignments
- Add --min-coverage flag to multivariate_analysis.py for sparse/shortform
datasets; also applies in lda_visualization.py
- Fix lda_visualization.py NaN handling and 0-d array annotation bug
- Update README.md and WORKFLOW.md to document datasets, sync workflow,
shortform handling, and new scripts
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>