- README: new 'Command flow' section — dependency-ordered pipeline (setup → batch generation → univariate → multivariate → LDA/cluster visualizations → classify → cross-version compare → JSON/model export) with actual output paths per step, using data/synthetic_1.4.0/ as the example - README: document chart production — univariate_analysis --img publishes the img/ summary charts; visualize_clusters added to the flow - README: fix -b path in 'Adding a new run' (bicorder.json lives at repo root) - WORKFLOW: script list and any-readings-CSV command block updated to match
376 lines
21 KiB
Markdown
376 lines
21 KiB
Markdown
# Bicorder data analysis
|
||
|
||
This directory concerns analyses conducted with the Protocol Bicorder across multiple datasets.
|
||
|
||
Scripts were created with the assistance of various AI tools. Data processing was done largely with either local models or the Ollama cloud service, which does not retain user data. Thanks to [Seth Frey (UC Davis)](https://enfascination.com/) for guidance, but all mistakes are the responsibility of the author, [Nathan Schneider](https://nathanschneider.info).
|
||
|
||
## Datasets
|
||
|
||
Readings are organized under `data/<type>_<version>/`, each self-contained with its own `readings.csv`, `analysis/`, and `json/` subdirectories. The shared protocol inputs (the list of protocols themselves, before any diagnostic is run) live directly under `data/`:
|
||
|
||
- **`data/protocols_edited.csv`** — the 411 cleaned protocol descriptors+descriptions (the input shared by every synthetic run)
|
||
- **`data/protocols_raw.csv`** — the 774 un-cleaned entries the chunking stage produced
|
||
|
||
Runs:
|
||
|
||
- **`data/synthetic_1.2.6/`** — 411 synthetic LLM-generated readings scored against bicorder v1.2.6 (see detailed procedure below). Renamed from `synthetic_20251116` to make the bicorder version it was generated against explicit, since the gradient structure changes between versions.
|
||
- **`data/synthetic_1.4.0/`** - Same as above, but scored on v1.4.0, which has some changes to the gradients.
|
||
- **`data/manual_20260320/`** — manual readings collected at [git.medlab.host/ntnsndr/protocol-bicorder-data](https://git.medlab.host/ntnsndr/protocol-bicorder-data), continuously expanding
|
||
|
||
### Adding a new run (e.g. re-running on edited gradients)
|
||
|
||
1. Make a new run directory keyed by bicorder version, e.g. `data/synthetic_1.4.0/`.
|
||
2. Point `bicorder_batch.py` at the shared input and the new output (`-b` defaults to `../bicorder.json` in the repo root):
|
||
```bash
|
||
python3 scripts/bicorder_batch.py data/protocols_edited.csv \
|
||
-o data/synthetic_1.4.0/readings.csv -b ../bicorder.json \
|
||
-m <model> -a "<analyst>" -s "<standpoint>"
|
||
```
|
||
3. The output `readings.csv` now records a `bicorder_version` column, so every run is self-describing about which gradient structure produced it. Run the downstream analysis (`univariate_analysis.py`, `multivariate_analysis.py`, `classify_readings.py`, etc.) against the new `readings.csv` with the new run directory as `--output`. The analysis scripts are version-agnostic: `classify_readings.py` auto-matches training data to the recorded version, and `compare_analyses.py` aligns renamed gradients across runs.
|
||
|
||
### Syncing the manual dataset
|
||
|
||
The manual dataset is kept current via a `.sync_source` config file and a one-command sync script:
|
||
|
||
```bash
|
||
scripts/sync_readings.sh data/manual_20260320
|
||
```
|
||
|
||
This clones the remote repository, copies JSON reading files, regenerates `readings.csv`, runs multivariate analysis (filtering to well-covered dimensions), generates an LDA visualization, and saves per-reading cluster classifications to `analysis/classifications.csv`.
|
||
|
||
Options:
|
||
```bash
|
||
scripts/sync_readings.sh data/manual_20260320 --min-coverage 0.8 # default
|
||
scripts/sync_readings.sh data/manual_20260320 --no-analysis # sync JSON only
|
||
scripts/sync_readings.sh data/manual_20260320 --training data/synthetic_1.2.6/readings.csv
|
||
```
|
||
|
||
### Handling shortform readings
|
||
|
||
Many manual readings use the shortform bicorder (9 key dimensions rather than all 23). Two analysis strategies handle this:
|
||
|
||
1. **Multivariate analysis with `--min-coverage`**: Drops dimension columns below the coverage threshold so analysis runs on the shared well-filled dimensions (e.g., 8 dimensions at 80% coverage for the current manual dataset).
|
||
2. **Classifier (`classify_readings.py`)**: Applies the synthetic-trained LDA model to all readings, filling any missing dimensions with a neutral value (5). The `completeness` column in the output flags readings where confidence is limited by sparse data.
|
||
|
||
---
|
||
|
||
## Command flow
|
||
|
||
The analysis pipeline for a run, using `data/synthetic_1.4.0/` as the example. Commands run from this directory with the virtualenv activated. The flow is version-agnostic: each `readings.csv` records its `bicorder_version`, and every script reads the gradient structure from the file itself (shared helpers in `scripts/bicorder_common.py`), so the same sequence applies to any run and any future version.
|
||
|
||
### 0. Setup (once)
|
||
|
||
```bash
|
||
python3 -m venv .venv
|
||
source .venv/bin/activate
|
||
pip install -r requirements.txt
|
||
```
|
||
|
||
### 1. Generate the readings (LLM batch diagnostic)
|
||
|
||
Queries the LLM for every gradient of every protocol, one chat each. Interrupted runs can be continued by re-running with `--resume`, which skips completed work. (For the manual dataset, `scripts/sync_readings.sh data/manual_20260320` replaces this step.)
|
||
|
||
```bash
|
||
python3 scripts/bicorder_batch.py data/protocols_edited.csv -o data/synthetic_1.4.0/readings.csv \
|
||
-m <model> -a "<analyst>" -s "<standpoint>"
|
||
```
|
||
|
||
→ `data/synthetic_1.4.0/readings.csv` — 411 protocols × 23 gradients, plus provenance columns (`bicorder_version`, `analyst`, `standpoint`)
|
||
|
||
### 2. Univariate analysis (averages and distributions)
|
||
|
||
```bash
|
||
python3 scripts/univariate_analysis.py data/synthetic_1.4.0/readings.csv
|
||
|
||
# add --img to also publish the three summary charts into img/ (the charts linked in this README's analysis narrative)
|
||
python3 scripts/univariate_analysis.py data/synthetic_1.2.6/readings.csv --img
|
||
```
|
||
|
||
→ `analysis/plots/` (protocol averages, distribution histogram, gradient averages), `analysis/data/` (per-protocol and per-gradient tables), and `analysis/reports/univariate_summary.txt` (mean, median, midpoint deviation, skewness)
|
||
|
||
### 3. Multivariate analysis (clustering, PCA, correlation, feature importance)
|
||
|
||
```bash
|
||
python3 scripts/multivariate_analysis.py data/synthetic_1.4.0/readings.csv
|
||
```
|
||
|
||
→ `analysis/plots/`, `analysis/data/` (including `kmeans_clusters.csv`, **required by steps 4–5**), and `analysis/reports/analysis_summary.txt`. Use `--analyses clustering pca` to run a subset, and `--min-coverage 0.8` for sparse (shortform) datasets.
|
||
|
||
### 4. Cluster visualizations (LDA + overlaid dimensional reductions)
|
||
|
||
```bash
|
||
# Cluster separation histogram and projection onto the discriminant axis
|
||
python3 scripts/lda_visualization.py data/synthetic_1.4.0/readings.csv
|
||
|
||
# Clusters overlaid on PCA / t-SNE / UMAP scatter plots
|
||
python3 scripts/visualize_clusters.py data/synthetic_1.4.0/readings.csv
|
||
```
|
||
|
||
→ `analysis/plots/lda_cluster_separation.png` (plus separation statistics on the console) and `pca_2d_clustered.png` / `tsne_2d_clustered.png` / `umap_2d_clustered.png`
|
||
|
||
### 5. Classify all readings
|
||
|
||
```bash
|
||
python3 scripts/classify_readings.py data/manual_20260320/readings.csv
|
||
```
|
||
|
||
Most typically applied to manual/shortform readings, with a synthetic run as training data. The training CSV is auto-selected to match the readings' recorded bicorder version (falls back to the most recent run and warns when versions differ, aligning renamed gradients); `--training` overrides. → `analysis/classifications.csv` (cluster, confidence, completeness, recommended form)
|
||
|
||
### 6. Compare runs (including across bicorder versions)
|
||
|
||
```bash
|
||
python3 scripts/compare_analyses.py data/synthetic_1.4.0/readings.csv data/synthetic_1.2.6/readings.csv
|
||
```
|
||
|
||
Canonicalizes old gradient names and compares only shared dimensions, so runs from different versions align. Console report: Euclidean distance, RMSE, correlation; with several comparison files it ranks them. With no arguments it executes the legacy model-audit comparison (see *Manual and alternate model audit* below).
|
||
|
||
### 7. Per-protocol JSONs and model export (optional)
|
||
|
||
```bash
|
||
python3 scripts/convert_csv_to_json.py data/synthetic_1.4.0/readings.csv
|
||
python3 scripts/export_model_for_js.py data/synthetic_1.4.0/readings.csv
|
||
```
|
||
|
||
→ `json/` (one bicorder.json-spec reading per protocol); `bicorder_model.json` in this directory, read by `bicorder-app` at build time (requires step 3 first; see `INTEGRATION_GUIDE.md`)
|
||
|
||
---
|
||
|
||
## Purpose
|
||
|
||
This analyses has several purposes:
|
||
|
||
* To test the usefulness and limitations of the Protocol Bicorder
|
||
* To identify potential improvements to the Protocol Bicorder
|
||
* To identify any patterns in a synthetic dataset derived from recent works on protocols
|
||
|
||
## Procedure
|
||
|
||
### Document chunking
|
||
|
||
This stage gathered raw data from recent protocol-focused texts.
|
||
|
||
The following prompt was applied to book chapter drafts and major protocol-related books, including the draft of the author's book, _The Protocol Reader_, _As for Protocols_, and _Das Protokoll_. The texts were pasted in plain text and then divided into 5000-word files, with the following prompt applied to each of them with the `chunk.sh` script:
|
||
|
||
```yaml
|
||
model: "gemma3:12b"
|
||
context: "model running on ollama locally, accessed with llm on the command line"
|
||
prompt: "Return csv-formatted data (with no markdown wrapper) that consists of a list of protocols discussed or referred to in the attached text. Protocols are defined extremely broadly as 'patterns of interaction,' and may be of a nontechnical nature. Protocols should be as specific as possible, such as 'Sacrament of Reconciliation' rather than 'Religious Protocols.' The first column should provide a brief descriptor of the protocol, and the second column should describe it in a substantial paragraph of 3-5 sentences, encapsulated in quotation marks to avoid breaking on commas. Be sure to paraphrase rather than quoting directly from the source text."
|
||
```
|
||
|
||
The result was a CSV-formatted list of protocols (`data/protocols_raw.csv`, n=774 total protocols listed).
|
||
|
||
### Dataset cleaning
|
||
|
||
The dataset was then manually reviewed. The review involved the following:
|
||
|
||
* Removal of repetitive formatting material introduced by the LLM
|
||
* Correction or removal of formatting errors
|
||
* Removal of rows whose contents met the following criteria:
|
||
- Repetition of entries---though some repetitions were simply merged into a single entry
|
||
- Overly broad entries that lacked meaningful context-specificity
|
||
- Overly narrow entries, e.g., referring to specific events
|
||
|
||
The cleaning process was carried out in a subjective manner, so some entries that meet the above criteria may remain in the dataset. The dataset also appears to include some LLM hallucinations---that is, protocols not in the texts---but the hallucinations are often acceptable examples and so some have been retained. Some degree of noise in the dataset was considered acceptable for the purposes of the study. Some degree of repetition, also, provides the dataset with a kind of control cases for evaluating the diagnostic process.
|
||
|
||
The result was a CSV-formatted list of protocols (`data/protocols_edited.csv`, n=411).
|
||
|
||
|
||
### Initial diagnostic
|
||
|
||
This diagnostic used the file now at `bicorder_analyzed.json`, though the scripts are set up to analyze `../bicorder.json`. That file has since been updated based on this analysis.
|
||
|
||
For each row in the dataset, and on each gradient, a series of scripts prompts the LLM to apply each gradient to the protocol. The outputs are then added to a CSV output file.
|
||
|
||
The result was a CSV-formatted list of protocols (`data/synthetic_1.2.6/readings.csv`, n=411).
|
||
|
||
See detailed documentation of the scripts at `WORKFLOW.md`.
|
||
|
||
### Manual and alternate model audit
|
||
|
||
To test the output, a manual review of the first 10 protocols in the `data/protocols_edited.csv` dataset was produced in the file `data/synthetic_1.2.6/readings_manual.csv`. (Alphabetization in this case seems a reasonable proxy for a random sample of protocols. It includes some partially overlapping protocols, as does the dataset as a whole.) Additionally, three models were tested on the same cases:
|
||
|
||
```bash
|
||
python3 scripts/bicorder_batch.py data/protocols_edited.csv -o data/synthetic_1.2.6/readings_mistral.csv -m mistral -a "Mistral" -s "A careful ethnographer and outsider aspiring to achieve a neutral stance and a high degree of precision" --start 1 --end 10
|
||
```
|
||
|
||
```bash
|
||
python3 scripts/bicorder_batch.py data/protocols_edited.csv -o data/synthetic_1.2.6/readings_gpt-oss.csv -m gpt-oss -a "GPT-OSS" -s "A careful ethnographer and outsider aspiring to achieve a neutral stance and a high degree of precision" --start 1 --end 10
|
||
```
|
||
|
||
```bash
|
||
python3 scripts/bicorder_batch.py data/protocols_edited.csv -o data/synthetic_1.2.6/readings_gemma3-12b.csv -m gemma3:12b -a "Gemma3:12b" -s "A careful ethnographer and outsider aspiring to achieve a neutral stance and a high degree of precision" --start 1 --end 10
|
||
```
|
||
|
||
A Euclidean distance analysis (`python3 scripts/compare_analyses.py`, with paths given as arguments: a reference CSV followed by comparison CSVs) found that the `gpt-oss` model was closer to the manual example than the others. It was therefore selected to be the model used for conducting the bicorder diagnostic on the dataset.
|
||
|
||
```
|
||
Average Euclidean Distance:
|
||
1. readings_gpt-oss.csv - Avg Distance: 11.68
|
||
2. readings_gemma3-12b.csv - Avg Distance: 13.06
|
||
3. readings_mistral.csv - Avg Distance: 13.33
|
||
```
|
||
|
||
Command used to produce `data/synthetic_1.2.6/readings.csv` (using the Ollama cloud service for the `gpt-oss` model):
|
||
|
||
```bash
|
||
python3 scripts/bicorder_batch.py data/protocols_edited.csv -o data/synthetic_1.2.6/readings.csv -m gpt-oss:20b-cloud -a "GPT-OSS" -s "A careful ethnographer and outsider aspiring to achieve a neutral stance and a high degree of precision"
|
||
```
|
||
|
||
The result was a CSV-formatted list of protocols (`data/synthetic_1.2.6/readings.csv`, n=411).
|
||
|
||
The same command, with changed version names, was used to produce the `synthetic_1.4.0` readings.
|
||
|
||
### Further analysis
|
||
|
||
#### Basic averages
|
||
|
||
For any run, `scripts/univariate_analysis.py` computes per-protocol and per-gradient averages, the summary statistics (mean, median, deviation from the midpoint, skewness), and the plots below, saving them to the run's `analysis/` directory:
|
||
|
||
```bash
|
||
python3 scripts/univariate_analysis.py data/synthetic_1.4.0/readings.csv
|
||
```
|
||
|
||
Per-protocol values are meaningful for the bicorder because, despite varying levels of appropriateness, all of the gradients are structured as ranging from "hardness" to "softness"---with lower values associated with greater rigidity. The average value for a given protocol, therefore, provides a rough sense of the protocol's hardness.
|
||
|
||
Basic averages appear in `data/synthetic_1.2.6/readings-analysis.ods`.
|
||
|
||
#### Univariate analysis
|
||
|
||
The charts linked below are the repo-level copies in `img/`, reproduced from a run with `python3 scripts/univariate_analysis.py <readings.csv> --img` (these were produced from the 1.2.6 dataset).
|
||
|
||
First, a plot of average values for each protocol:
|
||
|
||

|
||
|
||
This reveals a linear distribution of values among the protocols, aside from exponential curves only at the extremes. Perhaps the most interesting finding is a skew toward the higher end of the scale, associated with softness. Even relatively hard, technical protocols appear to have significant soft characteristics.
|
||
|
||
The protocol value averages have a mean of 5.45 and a median of 5.48. In comparison to the midpoint of 5, the normalized midpoint deviation is 0.11. In comparison, the Pearson coefficient measures the skew at just -0.07, which means that the relative skew of the data is actually slightly downward. So the distribution of protocol values is very balanced but has a consistent upward deviation from the scale's baseline. (These calculations are in `data/synthetic_1.2.6/readings-analysis.odt[averages]`.)
|
||
|
||
Second, a plot of average values for each gradient (with gaps to indicate the three groupings of gradients):
|
||
|
||

|
||
|
||
This indicates that a few of the gradients appear to have outsized responsibility for the high skew of the protocol averages.
|
||
|
||
* `Entanglement_exclusive_vs_non-exclusive` is the highest by nearly a full point
|
||
* Three others have averages over 7:
|
||
- `Design_precise_vs_interpretive`
|
||
- `Design_documenting_vs_enabling`
|
||
- `Design_technical_vs_social`
|
||
|
||
There is also an extreme at the bottom end: `Design_durable_vs_ephemeral` is the lowest by a full point.
|
||
|
||
It is not clear whether these extremes reveal anything other than a bias in the dataset or the LLM interpreter. Their descriptions may contribute to the extremes.
|
||
|
||
#### Multivariate analysis
|
||
|
||
Expectations:
|
||
|
||
* There are some gradients whose values are highly correlated. These might point to redundancies in the bicorder design.
|
||
* Some correlations might be revealing about connections in the characteristics of protocols, but these should be considered carefully as they may be the result of design or LLM interpretation.
|
||
|
||
I created a `multivariate_analysis.py` tool to conduct this analysis. Usage:
|
||
|
||
```bash
|
||
# Run all analyses (default)
|
||
python3 scripts/multivariate_analysis.py data/synthetic_1.2.6/readings.csv
|
||
|
||
# Run specific analyses only
|
||
python3 scripts/multivariate_analysis.py data/synthetic_1.2.6/readings.csv --analyses
|
||
clustering pca
|
||
```
|
||
|
||
Initial manual observations:
|
||
|
||
* The correlations generally seem predictable; for example, the strongest is between `Design_static_vs_malleable` and `Experience_predictable_vs_emergent`, which is not surprising
|
||
* The elite vs. vernacular distinction appears to be the most predictive gradient (`data/synthetic_1.2.6/analysis/plots/feature_importances.png`)
|
||
|
||

|
||
|
||

|
||
|
||
|
||
Comments:
|
||
|
||
* The distinction along the lines of "Vernacular/Emergent" and "Institutional/Standardized" tracks well with the structure of the book and the argument in chapter 3 that vernacular protocols have a logic distinct from institutional ones
|
||
* Strange that "Ethereum Proof of Work" appears in the "Vernacular/Emergent" family
|
||
|
||
|
||
## Conclusions
|
||
|
||
### Improvements to the bicorder
|
||
|
||
These findings indicate some options for improving on the current version of the bicorder.
|
||
|
||
Reflections on manual review:
|
||
|
||
* The exclusion/exclusive names could be improved to overlap less
|
||
* Should have separate values for "n/a" (0---but that could screw up averages) and both (5)
|
||
* Remove the analysis section, or use analyses here for what becomes most meaningful
|
||
|
||
|
||
### Future work: Description modification
|
||
|
||
Hypothesis: The phrasing of gradient value descriptions could have a significant impact on the salience of particular gradients.
|
||
|
||
Method: Modify the descriptions and run the same tests, see if the results are different.
|
||
|
||
|
||
### Future work: Persona elaboration
|
||
|
||
Hypothesis: Changing the analyst and their standpoint could result in interesting differences of outcomes.
|
||
|
||
Method: Alongside the dataset of protocols, generate diverse personas, such as a) personas used to evaluate every protocols, and b) protocol-specific personas that reflect different relationships to the protocol. Modify the test suite to include personas as an additional dimension of the analysis.
|
||
|
||
## Integration with Bicorder Tool (historical)
|
||
|
||
> **Update (v1.3.0):** The bureaucratic↔relational (formal/informal) LDA analysis
|
||
> has been **removed from the bicorder itself**. The cluster classification lives
|
||
> on as research in this directory only — see `INTEGRATION_GUIDE.md` for how to
|
||
> reproduce it and why it was removed from the tool.
|
||
|
||
The cluster analysis findings were previously integrated into the bicorder system as an automated analysis gradient:
|
||
|
||
**Bureaucratic ↔ Relational** - A new analysis field that automatically calculates where a protocol falls on the spectrum between the two protocol families identified through clustering analysis.
|
||
|
||
### Implementation
|
||
|
||
- **Model**: Linear Discriminant Analysis (LDA) trained on 406 protocols
|
||
- **Input**: The 23 diagnostic dimension values (read from bicorder.json in gradient order)
|
||
- **Output**: A value from 1-9 where:
|
||
- **1-3**: Strongly bureaucratic/institutional (formal, top-down, externally enforced)
|
||
- **4-6**: Mixed or boundary characteristics
|
||
- **7-9**: Strongly relational/cultural (emergent, voluntary, community-based)
|
||
|
||
**Design philosophy**: The model includes a `bicorder_version` field matching the bicorder.json version it was trained on. The implementation checks versions match before calculating. When bicorder.json structure changes (gradients added/removed/reordered), increment the version and retrain the model.
|
||
|
||
This simple version-matching approach ensures compatibility without complex structure mapping.
|
||
|
||
### Files
|
||
|
||
- `bicorder_model.json` (~5KB) - Trained LDA model with coefficients and scaler parameters; read by `bicorder-app` at build time
|
||
- `bicorder-app/src/bicorder-classifier.ts` - TypeScript classifier implementation in the web app
|
||
- `ascii_bicorder.py` (updated) - Python script now calculates automated analysis values
|
||
- `../bicorder.json` (updated) - Added bureaucratic ↔ relational gradient to analysis section
|
||
|
||
### Usage
|
||
|
||
The calculation happens automatically when generating bicorder output:
|
||
|
||
```bash
|
||
python3 ascii_bicorder.py bicorder.json bicorder.txt
|
||
```
|
||
|
||
For web integration, see `INTEGRATION_GUIDE.md`. The app (`bicorder-app/`) has its own classifier implementation and reads `bicorder_model.json` from this directory at build time.
|
||
|
||
### Key Features
|
||
|
||
- **Automated**: Calculated from diagnostic values, no manual assessment needed
|
||
- **Data-driven**: Based on multivariate analysis of 406 protocols
|
||
- **Single metric**: Distance to boundary determines classification confidence
|
||
- **Form recommendation**: Can suggest short vs. long form based on boundary distance
|
||
- **Lightweight**: 5KB model, no dependencies, runs client-side
|
||
|
||
The integration provides a data-backed way to understand where a protocol sits on the fundamental institutional/relational spectrum identified in the clustering analysis.
|
||
|