From 5a8440858732f42a3b5f8460c7c687ee661286ca Mon Sep 17 00:00:00 2001 From: Nathan Schneider Date: Fri, 2 Oct 2026 08:33:08 -0600 Subject: [PATCH] docs: command flow for analysis runs in README; keep WORKFLOW in sync MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit - README: new 'Command flow' section — dependency-ordered pipeline (setup → batch generation → univariate → multivariate → LDA/cluster visualizations → classify → cross-version compare → JSON/model export) with actual output paths per step, using data/synthetic_1.4.0/ as the example - README: document chart production — univariate_analysis --img publishes the img/ summary charts; visualize_clusters added to the flow - README: fix -b path in 'Adding a new run' (bicorder.json lives at repo root) - WORKFLOW: script list and any-readings-CSV command block updated to match --- analysis/README.md | 103 ++++++++++++++++++++++++++++++++++++++++--- analysis/WORKFLOW.md | 24 +++++++--- 2 files changed, 116 insertions(+), 11 deletions(-) diff --git a/analysis/README.md b/analysis/README.md index 1fa0d7a..a082839 100644 --- a/analysis/README.md +++ b/analysis/README.md @@ -14,19 +14,19 @@ Readings are organized under `data/_/`, each self-contained with Runs: - **`data/synthetic_1.2.6/`** — 411 synthetic LLM-generated readings scored against bicorder v1.2.6 (see detailed procedure below). Renamed from `synthetic_20251116` to make the bicorder version it was generated against explicit, since the gradient structure changes between versions. +- **`data/synthetic_1.4.0/`** - Same as above, but scored on v1.4.0, which has some changes to the gradients. - **`data/manual_20260320/`** — manual readings collected at [git.medlab.host/ntnsndr/protocol-bicorder-data](https://git.medlab.host/ntnsndr/protocol-bicorder-data), continuously expanding ### Adding a new run (e.g. re-running on edited gradients) 1. Make a new run directory keyed by bicorder version, e.g. `data/synthetic_1.4.0/`. -2. Point `bicorder_batch.py` at the shared input and the new output: +2. Point `bicorder_batch.py` at the shared input and the new output (`-b` defaults to `../bicorder.json` in the repo root): ```bash python3 scripts/bicorder_batch.py data/protocols_edited.csv \ - -o data/synthetic_1.4.0/readings.csv \ - -b bicorder.json \ + -o data/synthetic_1.4.0/readings.csv -b ../bicorder.json \ -m -a "" -s "" ``` -3. The output `readings.csv` now records a `bicorder_version` column, so every run is self-describing about which gradient structure produced it. Run the downstream analysis (`multivariate_analysis.py`, `classify_readings.py`, etc.) against the new `readings.csv` with the new run directory as `--output`. +3. The output `readings.csv` now records a `bicorder_version` column, so every run is self-describing about which gradient structure produced it. Run the downstream analysis (`univariate_analysis.py`, `multivariate_analysis.py`, `classify_readings.py`, etc.) against the new `readings.csv` with the new run directory as `--output`. The analysis scripts are version-agnostic: `classify_readings.py` auto-matches training data to the recorded version, and `compare_analyses.py` aligns renamed gradients across runs. ### Syncing the manual dataset @@ -54,6 +54,87 @@ Many manual readings use the shortform bicorder (9 key dimensions rather than al --- +## Command flow + +The analysis pipeline for a run, using `data/synthetic_1.4.0/` as the example. Commands run from this directory with the virtualenv activated. The flow is version-agnostic: each `readings.csv` records its `bicorder_version`, and every script reads the gradient structure from the file itself (shared helpers in `scripts/bicorder_common.py`), so the same sequence applies to any run and any future version. + +### 0. Setup (once) + +```bash +python3 -m venv .venv +source .venv/bin/activate +pip install -r requirements.txt +``` + +### 1. Generate the readings (LLM batch diagnostic) + +Queries the LLM for every gradient of every protocol, one chat each. Interrupted runs can be continued by re-running with `--resume`, which skips completed work. (For the manual dataset, `scripts/sync_readings.sh data/manual_20260320` replaces this step.) + +```bash +python3 scripts/bicorder_batch.py data/protocols_edited.csv -o data/synthetic_1.4.0/readings.csv \ + -m -a "" -s "" +``` + +→ `data/synthetic_1.4.0/readings.csv` — 411 protocols × 23 gradients, plus provenance columns (`bicorder_version`, `analyst`, `standpoint`) + +### 2. Univariate analysis (averages and distributions) + +```bash +python3 scripts/univariate_analysis.py data/synthetic_1.4.0/readings.csv + +# add --img to also publish the three summary charts into img/ (the charts linked in this README's analysis narrative) +python3 scripts/univariate_analysis.py data/synthetic_1.2.6/readings.csv --img +``` + +→ `analysis/plots/` (protocol averages, distribution histogram, gradient averages), `analysis/data/` (per-protocol and per-gradient tables), and `analysis/reports/univariate_summary.txt` (mean, median, midpoint deviation, skewness) + +### 3. Multivariate analysis (clustering, PCA, correlation, feature importance) + +```bash +python3 scripts/multivariate_analysis.py data/synthetic_1.4.0/readings.csv +``` + +→ `analysis/plots/`, `analysis/data/` (including `kmeans_clusters.csv`, **required by steps 4–5**), and `analysis/reports/analysis_summary.txt`. Use `--analyses clustering pca` to run a subset, and `--min-coverage 0.8` for sparse (shortform) datasets. + +### 4. Cluster visualizations (LDA + overlaid dimensional reductions) + +```bash +# Cluster separation histogram and projection onto the discriminant axis +python3 scripts/lda_visualization.py data/synthetic_1.4.0/readings.csv + +# Clusters overlaid on PCA / t-SNE / UMAP scatter plots +python3 scripts/visualize_clusters.py data/synthetic_1.4.0/readings.csv +``` + +→ `analysis/plots/lda_cluster_separation.png` (plus separation statistics on the console) and `pca_2d_clustered.png` / `tsne_2d_clustered.png` / `umap_2d_clustered.png` + +### 5. Classify all readings + +```bash +python3 scripts/classify_readings.py data/manual_20260320/readings.csv +``` + +Most typically applied to manual/shortform readings, with a synthetic run as training data. The training CSV is auto-selected to match the readings' recorded bicorder version (falls back to the most recent run and warns when versions differ, aligning renamed gradients); `--training` overrides. → `analysis/classifications.csv` (cluster, confidence, completeness, recommended form) + +### 6. Compare runs (including across bicorder versions) + +```bash +python3 scripts/compare_analyses.py data/synthetic_1.4.0/readings.csv data/synthetic_1.2.6/readings.csv +``` + +Canonicalizes old gradient names and compares only shared dimensions, so runs from different versions align. Console report: Euclidean distance, RMSE, correlation; with several comparison files it ranks them. With no arguments it executes the legacy model-audit comparison (see *Manual and alternate model audit* below). + +### 7. Per-protocol JSONs and model export (optional) + +```bash +python3 scripts/convert_csv_to_json.py data/synthetic_1.4.0/readings.csv +python3 scripts/export_model_for_js.py data/synthetic_1.4.0/readings.csv +``` + +→ `json/` (one bicorder.json-spec reading per protocol); `bicorder_model.json` in this directory, read by `bicorder-app` at build time (requires step 3 first; see `INTEGRATION_GUIDE.md`) + +--- + ## Purpose This analyses has several purposes: @@ -120,7 +201,7 @@ python3 scripts/bicorder_batch.py data/protocols_edited.csv -o data/synthetic_1. python3 scripts/bicorder_batch.py data/protocols_edited.csv -o data/synthetic_1.2.6/readings_gemma3-12b.csv -m gemma3:12b -a "Gemma3:12b" -s "A careful ethnographer and outsider aspiring to achieve a neutral stance and a high degree of precision" --start 1 --end 10 ``` -A Euclidean distance analysis (`python3 scripts/compare_analyses.py`) found that the `gpt-oss` model was closer to the manual example than the others. It was therefore selected to be the model used for conducting the bicorder diagnostic on the dataset. +A Euclidean distance analysis (`python3 scripts/compare_analyses.py`, with paths given as arguments: a reference CSV followed by comparison CSVs) found that the `gpt-oss` model was closer to the manual example than the others. It was therefore selected to be the model used for conducting the bicorder diagnostic on the dataset. ``` Average Euclidean Distance: @@ -137,16 +218,26 @@ python3 scripts/bicorder_batch.py data/protocols_edited.csv -o data/synthetic_1. The result was a CSV-formatted list of protocols (`data/synthetic_1.2.6/readings.csv`, n=411). +The same command, with changed version names, was used to produce the `synthetic_1.4.0` readings. + ### Further analysis #### Basic averages +For any run, `scripts/univariate_analysis.py` computes per-protocol and per-gradient averages, the summary statistics (mean, median, deviation from the midpoint, skewness), and the plots below, saving them to the run's `analysis/` directory: + +```bash +python3 scripts/univariate_analysis.py data/synthetic_1.4.0/readings.csv +``` + Per-protocol values are meaningful for the bicorder because, despite varying levels of appropriateness, all of the gradients are structured as ranging from "hardness" to "softness"---with lower values associated with greater rigidity. The average value for a given protocol, therefore, provides a rough sense of the protocol's hardness. Basic averages appear in `data/synthetic_1.2.6/readings-analysis.ods`. #### Univariate analysis +The charts linked below are the repo-level copies in `img/`, reproduced from a run with `python3 scripts/univariate_analysis.py --img` (these were produced from the 1.2.6 dataset). + First, a plot of average values for each protocol: ![Protocol averages plot](img/protocol_averages.png) @@ -178,7 +269,7 @@ Expectations: * There are some gradients whose values are highly correlated. These might point to redundancies in the bicorder design. * Some correlations might be revealing about connections in the characteristics of protocols, but these should be considered carefully as they may be the result of design or LLM interpretation. -Claude Code created a `multivariate_analysis.py` tool to conduct this analysis. Usage: +I created a `multivariate_analysis.py` tool to conduct this analysis. Usage: ```bash # Run all analyses (default) diff --git a/analysis/WORKFLOW.md b/analysis/WORKFLOW.md index a665a2e..0042d5c 100644 --- a/analysis/WORKFLOW.md +++ b/analysis/WORKFLOW.md @@ -21,9 +21,13 @@ The scripts automatically draw the gradients from the current state of the [bico 6. **scripts/multivariate_analysis.py** - Run clustering, PCA, correlation, and feature importance analysis on a readings CSV 7. **scripts/lda_visualization.py** - Generate LDA cluster separation plot and projection data -8. **scripts/classify_readings.py** - Apply the synthetic-trained LDA classifier to all readings; saves `analysis/classifications.csv` -9. **scripts/visualize_clusters.py** - Additional cluster visualizations -10. **scripts/export_model_for_js.py** - Export trained model to `bicorder_model.json` (read by `bicorder-app` at build time) +8. **scripts/classify_readings.py** - Apply the synthetic-trained LDA classifier to all readings; saves `analysis/classifications.csv` (training data is auto-matched to the input's recorded bicorder version) +9. **scripts/univariate_analysis.py** - Per-protocol and per-gradient averages, distributions, and summary stats (replaces the ad-hoc averages workflow) +10. **scripts/compare_analyses.py** - Compare readings CSVs to a reference (Euclidean distance, RMSE, correlation); canonicalizes renamed gradients so versions can be compared +11. **scripts/visualize_clusters.py** - Additional cluster visualizations +12. **scripts/export_model_for_js.py** - Export trained model to `bicorder_model.json` (read by `bicorder-app` at build time) + +Version-agnostic helpers shared by these scripts live in **scripts/bicorder_common.py**: the historical gradient rename map, version detection (from the `bicorder_version`/`version` column, falling back to the `data/_/` directory convention), and training-CSV selection. When gradients are renamed in `../bicorder.json`, update `COLUMN_RENAMES` in that one module. ## Syncing a manual readings dataset @@ -46,11 +50,21 @@ python3 scripts/multivariate_analysis.py data/manual_20260320/readings.csv \ # LDA visualization (cluster separation plot) python3 scripts/lda_visualization.py data/manual_20260320/readings.csv -# Classify all readings (uses synthetic dataset as training data by default) +# Classify all readings (training data auto-matched to the dataset's bicorder version) python3 scripts/classify_readings.py data/manual_20260320/readings.csv + +# Univariate averages (protocol and gradient plots + summary stats) +python3 scripts/univariate_analysis.py data/manual_20260320/readings.csv +# ... add --img to also publish its three summary charts into analysis/img/ (linked from README) + +# Cluster overlaid PCA/t-SNE/UMAP plots (after multivariate analysis) +python3 scripts/visualize_clusters.py data/manual_20260320/readings.csv + +# Compare two runs, including across bicorder versions (renamed gradients are aligned) +python3 scripts/compare_analyses.py data/synthetic_1.4.0/readings.csv data/synthetic_1.2.6/readings.csv ``` -Use `--min-coverage` (0.0–1.0) to drop dimension columns below the given coverage fraction before analysis. This is important for datasets with many shortform readings where most dimensions are sparsely filled. +Use `--min-coverage` (0.0–1.0) to drop dimension columns below the given coverage fraction before analysis. This is important for datasets with many shortform readings where most dimensions are sparsely filled. `classify_readings.py` still accepts an explicit `--training` CSV to override the automatic selection. ## Converting JSON reading files to CSV