docs: command flow for analysis runs in README; keep WORKFLOW in sync
- README: new 'Command flow' section — dependency-ordered pipeline (setup → batch generation → univariate → multivariate → LDA/cluster visualizations → classify → cross-version compare → JSON/model export) with actual output paths per step, using data/synthetic_1.4.0/ as the example - README: document chart production — univariate_analysis --img publishes the img/ summary charts; visualize_clusters added to the flow - README: fix -b path in 'Adding a new run' (bicorder.json lives at repo root) - WORKFLOW: script list and any-readings-CSV command block updated to match
This commit is contained in:
1 parent
8c44064196
commit
5a84408587
2 files changed
+116
-11
No files matched your search
+19
-5
@@ -21,9 +21,13 @@ The scripts automatically draw the gradients from the current state of the [bico
|
||||
|
||||
6. **scripts/multivariate_analysis.py** - Run clustering, PCA, correlation, and feature importance analysis on a readings CSV
|
||||
7. **scripts/lda_visualization.py** - Generate LDA cluster separation plot and projection data
|
||||
8. **scripts/classify_readings.py** - Apply the synthetic-trained LDA classifier to all readings; saves `analysis/classifications.csv`
|
||||
9. **scripts/visualize_clusters.py** - Additional cluster visualizations
|
||||
10. **scripts/export_model_for_js.py** - Export trained model to `bicorder_model.json` (read by `bicorder-app` at build time)
|
||||
8. **scripts/classify_readings.py** - Apply the synthetic-trained LDA classifier to all readings; saves `analysis/classifications.csv` (training data is auto-matched to the input's recorded bicorder version)
|
||||
9. **scripts/univariate_analysis.py** - Per-protocol and per-gradient averages, distributions, and summary stats (replaces the ad-hoc averages workflow)
|
||||
10. **scripts/compare_analyses.py** - Compare readings CSVs to a reference (Euclidean distance, RMSE, correlation); canonicalizes renamed gradients so versions can be compared
|
||||
11. **scripts/visualize_clusters.py** - Additional cluster visualizations
|
||||
12. **scripts/export_model_for_js.py** - Export trained model to `bicorder_model.json` (read by `bicorder-app` at build time)
|
||||
|
||||
Version-agnostic helpers shared by these scripts live in **scripts/bicorder_common.py**: the historical gradient rename map, version detection (from the `bicorder_version`/`version` column, falling back to the `data/<type>_<version>/` directory convention), and training-CSV selection. When gradients are renamed in `../bicorder.json`, update `COLUMN_RENAMES` in that one module.
|
||||
|
||||
## Syncing a manual readings dataset
|
||||
|
||||
@@ -46,11 +50,21 @@ python3 scripts/multivariate_analysis.py data/manual_20260320/readings.csv \
|
||||
# LDA visualization (cluster separation plot)
|
||||
python3 scripts/lda_visualization.py data/manual_20260320/readings.csv
|
||||
|
||||
# Classify all readings (uses synthetic dataset as training data by default)
|
||||
# Classify all readings (training data auto-matched to the dataset's bicorder version)
|
||||
python3 scripts/classify_readings.py data/manual_20260320/readings.csv
|
||||
|
||||
# Univariate averages (protocol and gradient plots + summary stats)
|
||||
python3 scripts/univariate_analysis.py data/manual_20260320/readings.csv
|
||||
# ... add --img to also publish its three summary charts into analysis/img/ (linked from README)
|
||||
|
||||
# Cluster overlaid PCA/t-SNE/UMAP plots (after multivariate analysis)
|
||||
python3 scripts/visualize_clusters.py data/manual_20260320/readings.csv
|
||||
|
||||
# Compare two runs, including across bicorder versions (renamed gradients are aligned)
|
||||
python3 scripts/compare_analyses.py data/synthetic_1.4.0/readings.csv data/synthetic_1.2.6/readings.csv
|
||||
```
|
||||
|
||||
Use `--min-coverage` (0.0–1.0) to drop dimension columns below the given coverage fraction before analysis. This is important for datasets with many shortform readings where most dimensions are sparsely filled.
|
||||
Use `--min-coverage` (0.0–1.0) to drop dimension columns below the given coverage fraction before analysis. This is important for datasets with many shortform readings where most dimensions are sparsely filled. `classify_readings.py` still accepts an explicit `--training` CSV to override the automatic selection.
|
||||
|
||||
## Converting JSON reading files to CSV
|
||||
|
||||
|
||||
Reference in new issue
Block a user