docs: command flow for analysis runs in README; keep WORKFLOW in sync

- README: new 'Command flow' section — dependency-ordered pipeline (setup →
  batch generation → univariate → multivariate → LDA/cluster visualizations →
  classify → cross-version compare → JSON/model export) with actual output
  paths per step, using data/synthetic_1.4.0/ as the example
- README: document chart production — univariate_analysis --img publishes the
  img/ summary charts; visualize_clusters added to the flow
- README: fix -b path in 'Adding a new run' (bicorder.json lives at repo root)
- WORKFLOW: script list and any-readings-CSV command block updated to match
This commit is contained in:
Nathan Schneider committed 2026-10-02 08:33:08 -06:00
1 parent 8c44064196
commit 5a84408587
2 files changed
+116 -11

No files matched your search

+19 -5
View File
@@ -21,9 +21,13 @@ The scripts automatically draw the gradients from the current state of the [bico
6. **scripts/multivariate_analysis.py** - Run clustering, PCA, correlation, and feature importance analysis on a readings CSV
7. **scripts/lda_visualization.py** - Generate LDA cluster separation plot and projection data
8. **scripts/classify_readings.py** - Apply the synthetic-trained LDA classifier to all readings; saves `analysis/classifications.csv`
9. **scripts/visualize_clusters.py** - Additional cluster visualizations
10. **scripts/export_model_for_js.py** - Export trained model to `bicorder_model.json` (read by `bicorder-app` at build time)
8. **scripts/classify_readings.py** - Apply the synthetic-trained LDA classifier to all readings; saves `analysis/classifications.csv` (training data is auto-matched to the input's recorded bicorder version)
9. **scripts/univariate_analysis.py** - Per-protocol and per-gradient averages, distributions, and summary stats (replaces the ad-hoc averages workflow)
10. **scripts/compare_analyses.py** - Compare readings CSVs to a reference (Euclidean distance, RMSE, correlation); canonicalizes renamed gradients so versions can be compared
11. **scripts/visualize_clusters.py** - Additional cluster visualizations
12. **scripts/export_model_for_js.py** - Export trained model to `bicorder_model.json` (read by `bicorder-app` at build time)
Version-agnostic helpers shared by these scripts live in **scripts/bicorder_common.py**: the historical gradient rename map, version detection (from the `bicorder_version`/`version` column, falling back to the `data/<type>_<version>/` directory convention), and training-CSV selection. When gradients are renamed in `../bicorder.json`, update `COLUMN_RENAMES` in that one module.
## Syncing a manual readings dataset
@@ -46,11 +50,21 @@ python3 scripts/multivariate_analysis.py data/manual_20260320/readings.csv \
# LDA visualization (cluster separation plot)
python3 scripts/lda_visualization.py data/manual_20260320/readings.csv
# Classify all readings (uses synthetic dataset as training data by default)
# Classify all readings (training data auto-matched to the dataset's bicorder version)
python3 scripts/classify_readings.py data/manual_20260320/readings.csv
# Univariate averages (protocol and gradient plots + summary stats)
python3 scripts/univariate_analysis.py data/manual_20260320/readings.csv
# ... add --img to also publish its three summary charts into analysis/img/ (linked from README)
# Cluster overlaid PCA/t-SNE/UMAP plots (after multivariate analysis)
python3 scripts/visualize_clusters.py data/manual_20260320/readings.csv
# Compare two runs, including across bicorder versions (renamed gradients are aligned)
python3 scripts/compare_analyses.py data/synthetic_1.4.0/readings.csv data/synthetic_1.2.6/readings.csv
```
Use `--min-coverage` (0.0–1.0) to drop dimension columns below the given coverage fraction before analysis. This is important for datasets with many shortform readings where most dimensions are sparsely filled.
Use `--min-coverage` (0.0–1.0) to drop dimension columns below the given coverage fraction before analysis. This is important for datasets with many shortform readings where most dimensions are sparsely filled. `classify_readings.py` still accepts an explicit `--training` CSV to override the automatic selection.
## Converting JSON reading files to CSV