Compare commits

..
13 Commits
Author SHA1 Message Date
Nathan Schneider b6fc370ba7 docs: cite reproducible skew stats from univariate_summary in README 2026-10-02 08:48:39 -06:00
Nathan Schneider 7c2de2cefb docs: refresh img/ summary charts; add img index; cite reproducible stats
- Regenerated the three README charts from data/synthetic_1.2.6 via
  scripts/univariate_analysis.py --img (replaces the Feb 2026 ad-hoc versions)
- img/README.md: documents img/ as the published-docs chart directory
  (stable README link targets), with provenance and the regeneration command
- README: skew/mean narrative updated to the reproducible script output
  (mean 5.44, median 5.48; Fisher skew -0.03) with pointer to
  univariate_summary.txt, keeping the historical ODS reference
2026-10-02 08:47:30 -06:00
Nathan Schneider e100848a0d data: univariate averages outputs for the 1.2.6 run
Produced alongside the img/ chart refresh via scripts/univariate_analysis.py:
per-protocol and per-gradient averages tables, summary plots, and
univariate_summary.txt (mean 5.44, median 5.48, normalized midpoint
deviation 0.11).
2026-10-02 08:47:27 -06:00
Nathan Schneider ff19af5dd1 data: analysis outputs for the synthetic_1.4.0 run
Full pipeline per README command flow (steps 2-6): univariate averages and
distributions, multivariate analysis (2-cluster k-means: 209 relational vs
199 institutional; PCA/t-SNE/UMAP/factor/correlation/network/importance),
LDA separation and overlaid cluster plots, and LDA classifications (trained
on the 1.2.6 run via version-matched auto-selection). Cross-version summary
is on the console report: avg Euclidean distance 9.06 vs 1.2.6 (r=0.79).
2026-10-02 08:39:52 -06:00
Nathan Schneider 71c1977b77 docs: bicorder_model.json is research-only (app integration removed in v1.3.0)
Verified: bicorder-app has no classifier implementation and reads
bicorder_model.json nowhere (src, build config); ascii_bicorder.py doesn't
use it either. Corrected the stale 'read by bicorder-app at build time'
claims in README, WORKFLOW, and the export script docstring.
2026-10-02 08:39:26 -06:00
Nathan Schneider f136df46db Slight rewording in README for uniformity 2026-10-02 08:33:58 -06:00
Nathan Schneider 5a84408587 docs: command flow for analysis runs in README; keep WORKFLOW in sync
- README: new 'Command flow' section — dependency-ordered pipeline (setup →
  batch generation → univariate → multivariate → LDA/cluster visualizations →
  classify → cross-version compare → JSON/model export) with actual output
  paths per step, using data/synthetic_1.4.0/ as the example
- README: document chart production — univariate_analysis --img publishes the
  img/ summary charts; visualize_clusters added to the flow
- README: fix -b path in 'Adding a new run' (bicorder.json lives at repo root)
- WORKFLOW: script list and any-readings-CSV command block updated to match
2026-10-02 08:33:08 -06:00
Nathan Schneider 8c44064196 data: add synthetic_1.4.0 run readings
411 protocols × 23 gradients scored against bicorder.json v1.4.0 with
gpt-oss:20b-cloud (same analyst standpoint as the 1.2.6 run); the
bicorder_version column makes the run self-describing. Analysis outputs to
follow in data/synthetic_1.4.0/analysis/.
2026-10-02 08:33:03 -06:00
Nathan Schneider 55cbd6cd5d feat: version-agnostic analysis scripts with shared version helpers
- scripts/bicorder_common.py (new): single source of truth for historical
  gradient renames (COLUMN_RENAMES), version detection (bicorder_version col
  → version col → data/<type>_<version>/ dir convention), and training-CSV
  auto-selection (find_training_csv)
- classify_readings.py: auto-select training run by recorded bicorder version
  (excludes the input itself to avoid circular training), canonicalize old
  column names, version-mismatch warnings
- bicorder_classifier.py / export_model_for_js.py: share renames + dimension
  loader; instructive error when clustering results are missing;
  bicorder_version recorded in exported models
- compare_analyses.py: CLI (reference + comparison CSVs), rename
  canonicalization so runs of different versions align on shared gradients,
  Descriptor dedup (was silently skewing merges); legacy no-arg audit intact
- scripts/univariate_analysis.py (new): per-protocol/per-gradient averages,
  distributions, summary stats; --img publishes the three README summary
  charts to img/
- sync_readings.sh: defer classifier training to auto-matching; gitignore
  analysis/.venv and __pycache__
2026-10-02 08:32:59 -06:00
Nathan Schneider a4f16e8e4d Removed raw Claude outputs from README 2026-09-23 14:44:05 -06:00
Protocolbot 161ba3b136 fix: carry institutional→formal rename through analysis rename maps
Nathan renamed the Design gradient 'institutional' → 'formal' (v1.4.0).
The analysis scripts carry historical rename maps that route old data
(elite/institutional) into the current bicorder.json gradient names, and
those maps still stopped at 'institutional' — so old readings would have
silently misaligned against the v1.4.0 structure.

- bicorder_classifier.py / export_model_for_js.py: add
  institutional_vs_vernacular → formal_vs_vernacular, keep elite→formal
- json_to_csv.py TERM_RENAMES: same two-step route
- convert_csv_to_json.py GRADIENT_MAPPINGS: elite_vs_vernacular now maps
  to 'formal'
- bicorder_classifier.py demo ratings: use the current dimension name
- analysis/README.md: example run name 1.3.0 → 1.4.0
2026-09-23 14:27:51 -06:00
Protocolbot 6ae77a4f9b refactor: reorganize analysis data to support multiple runs
Strategy: key runs by bicorder version (not date), promote shared inputs
to analysis/data/, and stamp every output with its bicorder_version so
a re-run on edited gradients is self-describing.

Data layout:
- Promote the shared protocol inputs out of the run directory:
    analysis/data/protocols_edited.csv  (411 cleaned protocols)
    analysis/data/protocols_raw.csv     (774 un-cleaned entries)
- Rename the v1.2.6 synthetic run:
    data/synthetic_20251116/ -> data/synthetic_1.2.6/
  so the bicorder version it was scored against is explicit (gradient
  structure changes between versions make date-based names ambiguous)

Provenance:
- bicorder_analyze.py now writes a 'bicorder_version' column into every
  output readings.csv, recording which gradient structure produced it

Scripts:
- Update the real code defaults that pointed at the old run path
  (bicorder_classifier.py, classify_readings.py, sync_readings.sh,
  compare_analyses.py) and refresh docstring/help examples
- Remove a stray committed __pycache__/.pyc

Docs: analysis/README.md documents the new layout + how to add a run;
WORKFLOW.md, TEST_COMMANDS.md, INTEGRATION_GUIDE.md paths updated.
2026-09-23 14:23:17 -06:00
Nathan Schneider 459015fe17 Switched institutional > formal 2026-09-23 09:37:25 -06:00
537 changed files with 6803 additions and 422 deletions

No files matched your search

+2
View File
@@ -1,3 +1,5 @@
tmp
analysis/venv
analysis/.venv
__pycache__/
.aider*
+1 -1
View File
@@ -22,7 +22,7 @@ previously integrated into the Bicorder web application to provide:
**Version-based compatibility**: The model included a `bicorder_version` field.
The classifier checked that versions matched. When bicorder.json structure changed:
1. The version number in bicorder.json was incremented
2. The model was retrained with `python3 scripts/export_model_for_js.py data/synthetic_20251116/readings.csv`
2. The model was retrained with `python3 scripts/export_model_for_js.py data/synthetic_1.2.6/readings.csv`
3. The new model had the updated version
## Files (research-only now)
+133 -227
View File
@@ -2,15 +2,32 @@
This directory concerns analyses conducted with the Protocol Bicorder across multiple datasets.
Scripts were created with the assistance of Claude Code. Data processing was done largely with either local models or the Ollama cloud service, which does not retain user data. Thanks to [Seth Frey (UC Davis)](https://enfascination.com/) for guidance, but all mistakes are the responsibility of the author, [Nathan Schneider](https://nathanschneider.info). This is the work of a researcher working with AI outside their field of expertise and should be treated as a playful experiment, not a model of rigorous methodology.
Scripts were created with the assistance of various AI tools. Data processing was done largely with either local models or the Ollama cloud service, which does not retain user data. Thanks to [Seth Frey (UC Davis)](https://enfascination.com/) for guidance, but all mistakes are the responsibility of the author, [Nathan Schneider](https://nathanschneider.info).
## Datasets
Readings are organized under `data/<type>_<YYYYMMDD>/`, each self-contained with its own `readings.csv`, `analysis/`, and `json/` subdirectories:
Readings are organized under `data/<type>_<version>/`, each self-contained with its own `readings.csv`, `analysis/`, and `json/` subdirectories. The shared protocol inputs (the list of protocols themselves, before any diagnostic is run) live directly under `data/`:
- **`data/synthetic_20251116/`** — 411 protocols from synthetic LLM-generated readings (see detailed procedure below)
- **`data/protocols_edited.csv`** — the 411 cleaned protocol descriptors+descriptions (the input shared by every synthetic run)
- **`data/protocols_raw.csv`** — the 774 un-cleaned entries the chunking stage produced
Runs:
- **`data/synthetic_1.2.6/`** — 411 synthetic LLM-generated readings scored against bicorder v1.2.6 (see detailed procedure below). Renamed from `synthetic_20251116` to make the bicorder version it was generated against explicit, since the gradient structure changes between versions.
- **`data/synthetic_1.4.0/`** - Same as above, but scored on v1.4.0, which has some changes to the gradients.
- **`data/manual_20260320/`** — manual readings collected at [git.medlab.host/ntnsndr/protocol-bicorder-data](https://git.medlab.host/ntnsndr/protocol-bicorder-data), continuously expanding
### Adding a new run (e.g. re-running on edited gradients)
1. Make a new run directory keyed by bicorder version, e.g. `data/synthetic_1.4.0/`.
2. Point `bicorder_batch.py` at the shared input and the new output (`-b` defaults to `../bicorder.json` in the repo root):
```bash
python3 scripts/bicorder_batch.py data/protocols_edited.csv \
-o data/synthetic_1.4.0/readings.csv -b ../bicorder.json \
-m <model> -a "<analyst>" -s "<standpoint>"
```
3. The output `readings.csv` now records a `bicorder_version` column, so every run is self-describing about which gradient structure produced it. Run the downstream analysis (`univariate_analysis.py`, `multivariate_analysis.py`, `classify_readings.py`, etc.) against the new `readings.csv` with the new run directory as `--output`. The analysis scripts are version-agnostic: `classify_readings.py` auto-matches training data to the recorded version, and `compare_analyses.py` aligns renamed gradients across runs.
### Syncing the manual dataset
The manual dataset is kept current via a `.sync_source` config file and a one-command sync script:
@@ -25,7 +42,7 @@ Options:
```bash
scripts/sync_readings.sh data/manual_20260320 --min-coverage 0.8 # default
scripts/sync_readings.sh data/manual_20260320 --no-analysis # sync JSON only
scripts/sync_readings.sh data/manual_20260320 --training data/synthetic_20251116/readings.csv
scripts/sync_readings.sh data/manual_20260320 --training data/synthetic_1.2.6/readings.csv
```
### Handling shortform readings
@@ -37,6 +54,87 @@ Many manual readings use the shortform bicorder (9 key dimensions rather than al
---
## Command flow
The analysis pipeline for a run, using `data/synthetic_1.4.0/` as the example. Commands run from this directory with the virtualenv activated. The flow is version-agnostic: each `readings.csv` records its `bicorder_version`, and every script reads the gradient structure from the file itself (shared helpers in `scripts/bicorder_common.py`), so the same sequence applies to any run and any future version.
### 0. Setup (once)
```bash
python3 -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt
```
### 1. Generate the readings (LLM batch diagnostic)
Queries the LLM for every gradient of every protocol, one chat each. Interrupted runs can be continued by re-running with `--resume`, which skips completed work. (For the manual dataset, `scripts/sync_readings.sh data/manual_20260320` replaces this step.)
```bash
python3 scripts/bicorder_batch.py data/protocols_edited.csv -o data/synthetic_1.4.0/readings.csv \
-m <model> -a "<analyst>" -s "<standpoint>"
```
→ `data/synthetic_1.4.0/readings.csv` — 411 protocols × 23 gradients, plus provenance columns (`bicorder_version`, `analyst`, `standpoint`)
### 2. Univariate analysis (averages and distributions)
```bash
python3 scripts/univariate_analysis.py data/synthetic_1.4.0/readings.csv
# add --img to also publish the three summary charts into img/ (the charts linked in this README's analysis narrative)
python3 scripts/univariate_analysis.py data/synthetic_1.2.6/readings.csv --img
```
→ `analysis/plots/` (protocol averages, distribution histogram, gradient averages), `analysis/data/` (per-protocol and per-gradient tables), and `analysis/reports/univariate_summary.txt` (mean, median, midpoint deviation, skewness)
### 3. Multivariate analysis (clustering, PCA, correlation, feature importance)
```bash
python3 scripts/multivariate_analysis.py data/synthetic_1.4.0/readings.csv
```
→ `analysis/plots/`, `analysis/data/` (including `kmeans_clusters.csv`, **required by steps 4–5**), and `analysis/reports/analysis_summary.txt`. Use `--analyses clustering pca` to run a subset, and `--min-coverage 0.8` for sparse (shortform) datasets.
### 4. Cluster visualizations (LDA + overlaid dimensional reductions)
```bash
# Cluster separation histogram and projection onto the discriminant axis
python3 scripts/lda_visualization.py data/synthetic_1.4.0/readings.csv
# Clusters overlaid on PCA / t-SNE / UMAP scatter plots
python3 scripts/visualize_clusters.py data/synthetic_1.4.0/readings.csv
```
→ `analysis/plots/lda_cluster_separation.png` (plus separation statistics on the console) and `pca_2d_clustered.png` / `tsne_2d_clustered.png` / `umap_2d_clustered.png`
### 5. Classify all readings
```bash
python3 scripts/classify_readings.py data/manual_20260320/readings.csv
```
Most typically applied to manual/shortform readings, with a synthetic run as training data. The training CSV is auto-selected to match the readings' recorded bicorder version (falls back to the most recent run and warns when versions differ, aligning renamed gradients); `--training` overrides. → `analysis/classifications.csv` (cluster, confidence, completeness, recommended form)
### 6. Compare runs (including across bicorder versions)
```bash
python3 scripts/compare_analyses.py data/synthetic_1.4.0/readings.csv data/synthetic_1.2.6/readings.csv
```
Canonicalizes old gradient names and compares only shared dimensions, so runs from different versions align. Console report: Euclidean distance, RMSE, correlation; with several comparison files it ranks them. With no arguments it executes the legacy model-audit comparison (see *Manual and alternate model audit* below).
### 7. Per-protocol JSONs and model export (optional)
```bash
python3 scripts/convert_csv_to_json.py data/synthetic_1.4.0/readings.csv
python3 scripts/export_model_for_js.py data/synthetic_1.4.0/readings.csv
```
→ `json/` (one bicorder.json-spec reading per protocol). The optional `export_model_for_js.py` refreshes the research-only `bicorder_model.json` in this directory (app integration removed in bicorder v1.3.0; requires step 3 first)
---
## Purpose
This analyses has several purposes:
@@ -59,7 +157,7 @@ context: "model running on ollama locally, accessed with llm on the command line
prompt: "Return csv-formatted data (with no markdown wrapper) that consists of a list of protocols discussed or referred to in the attached text. Protocols are defined extremely broadly as 'patterns of interaction,' and may be of a nontechnical nature. Protocols should be as specific as possible, such as 'Sacrament of Reconciliation' rather than 'Religious Protocols.' The first column should provide a brief descriptor of the protocol, and the second column should describe it in a substantial paragraph of 3-5 sentences, encapsulated in quotation marks to avoid breaking on commas. Be sure to paraphrase rather than quoting directly from the source text."
```
The result was a CSV-formatted list of protocols (`data/synthetic_20251116/protocols_raw.csv`, n=774 total protocols listed).
The result was a CSV-formatted list of protocols (`data/protocols_raw.csv`, n=774 total protocols listed).
### Dataset cleaning
@@ -74,7 +172,7 @@ The dataset was then manually reviewed. The review involved the following:
The cleaning process was carried out in a subjective manner, so some entries that meet the above criteria may remain in the dataset. The dataset also appears to include some LLM hallucinations---that is, protocols not in the texts---but the hallucinations are often acceptable examples and so some have been retained. Some degree of noise in the dataset was considered acceptable for the purposes of the study. Some degree of repetition, also, provides the dataset with a kind of control cases for evaluating the diagnostic process.
The result was a CSV-formatted list of protocols (`data/synthetic_20251116/protocols_edited.csv`, n=411).
The result was a CSV-formatted list of protocols (`data/protocols_edited.csv`, n=411).
### Initial diagnostic
@@ -83,27 +181,27 @@ This diagnostic used the file now at `bicorder_analyzed.json`, though the script
For each row in the dataset, and on each gradient, a series of scripts prompts the LLM to apply each gradient to the protocol. The outputs are then added to a CSV output file.
The result was a CSV-formatted list of protocols (`data/synthetic_20251116/readings.csv`, n=411).
The result was a CSV-formatted list of protocols (`data/synthetic_1.2.6/readings.csv`, n=411).
See detailed documentation of the scripts at `WORKFLOW.md`.
### Manual and alternate model audit
To test the output, a manual review of the first 10 protocols in the `data/synthetic_20251116/protocols_edited.csv` dataset was produced in the file `data/synthetic_20251116/readings_manual.csv`. (Alphabetization in this case seems a reasonable proxy for a random sample of protocols. It includes some partially overlapping protocols, as does the dataset as a whole.) Additionally, three models were tested on the same cases:
To test the output, a manual review of the first 10 protocols in the `data/protocols_edited.csv` dataset was produced in the file `data/synthetic_1.2.6/readings_manual.csv`. (Alphabetization in this case seems a reasonable proxy for a random sample of protocols. It includes some partially overlapping protocols, as does the dataset as a whole.) Additionally, three models were tested on the same cases:
```bash
python3 scripts/bicorder_batch.py data/synthetic_20251116/protocols_edited.csv -o data/synthetic_20251116/readings_mistral.csv -m mistral -a "Mistral" -s "A careful ethnographer and outsider aspiring to achieve a neutral stance and a high degree of precision" --start 1 --end 10
python3 scripts/bicorder_batch.py data/protocols_edited.csv -o data/synthetic_1.2.6/readings_mistral.csv -m mistral -a "Mistral" -s "A careful ethnographer and outsider aspiring to achieve a neutral stance and a high degree of precision" --start 1 --end 10
```
```bash
python3 scripts/bicorder_batch.py data/synthetic_20251116/protocols_edited.csv -o data/synthetic_20251116/readings_gpt-oss.csv -m gpt-oss -a "GPT-OSS" -s "A careful ethnographer and outsider aspiring to achieve a neutral stance and a high degree of precision" --start 1 --end 10
python3 scripts/bicorder_batch.py data/protocols_edited.csv -o data/synthetic_1.2.6/readings_gpt-oss.csv -m gpt-oss -a "GPT-OSS" -s "A careful ethnographer and outsider aspiring to achieve a neutral stance and a high degree of precision" --start 1 --end 10
```
```bash
python3 scripts/bicorder_batch.py data/synthetic_20251116/protocols_edited.csv -o data/synthetic_20251116/readings_gemma3-12b.csv -m gemma3:12b -a "Gemma3:12b" -s "A careful ethnographer and outsider aspiring to achieve a neutral stance and a high degree of precision" --start 1 --end 10
python3 scripts/bicorder_batch.py data/protocols_edited.csv -o data/synthetic_1.2.6/readings_gemma3-12b.csv -m gemma3:12b -a "Gemma3:12b" -s "A careful ethnographer and outsider aspiring to achieve a neutral stance and a high degree of precision" --start 1 --end 10
```
A Euclidean distance analysis (`python3 scripts/compare_analyses.py`) found that the `gpt-oss` model was closer to the manual example than the others. It was therefore selected to be the model used for conducting the bicorder diagnostic on the dataset.
A Euclidean distance analysis (`python3 scripts/compare_analyses.py`, with paths given as arguments: a reference CSV followed by comparison CSVs) found that the `gpt-oss` model was closer to the manual example than the others. It was therefore selected to be the model used for conducting the bicorder diagnostic on the dataset.
```
Average Euclidean Distance:
@@ -112,31 +210,41 @@ Average Euclidean Distance:
3. readings_mistral.csv - Avg Distance: 13.33
```
Command used to produce `data/synthetic_20251116/readings.csv` (using the Ollama cloud service for the `gpt-oss` model):
Command used to produce `data/synthetic_1.2.6/readings.csv` (using the Ollama cloud service for the `gpt-oss` model):
```bash
python3 scripts/bicorder_batch.py data/synthetic_20251116/protocols_edited.csv -o data/synthetic_20251116/readings.csv -m gpt-oss:20b-cloud -a "GPT-OSS" -s "A careful ethnographer and outsider aspiring to achieve a neutral stance and a high degree of precision"
python3 scripts/bicorder_batch.py data/protocols_edited.csv -o data/synthetic_1.2.6/readings.csv -m gpt-oss:20b-cloud -a "GPT-OSS" -s "A careful ethnographer and outsider aspiring to achieve a neutral stance and a high degree of precision"
```
The result was a CSV-formatted list of protocols (`data/synthetic_20251116/readings.csv`, n=411).
The result was a CSV-formatted list of protocols (`data/synthetic_1.2.6/readings.csv`, n=411).
The same command, with changed version names, was used to produce the `synthetic_1.4.0` readings.
### Further analysis
#### Basic averages
For any run, `scripts/univariate_analysis.py` computes per-protocol and per-gradient averages, the summary statistics (mean, median, deviation from the midpoint, skewness), and the plots below, saving them to the run's `analysis/` directory:
```bash
python3 scripts/univariate_analysis.py data/synthetic_1.4.0/readings.csv
```
Per-protocol values are meaningful for the bicorder because, despite varying levels of appropriateness, all of the gradients are structured as ranging from "hardness" to "softness"---with lower values associated with greater rigidity. The average value for a given protocol, therefore, provides a rough sense of the protocol's hardness.
Basic averages appear in `data/synthetic_20251116/readings-analysis.ods`.
Basic averages appear in `data/synthetic_1.2.6/readings-analysis.ods`.
#### Univariate analysis
The charts linked below are the repo-level copies in `img/`, reproduced from a run with `python3 scripts/univariate_analysis.py <readings.csv> --img` (these were produced from the 1.2.6 dataset).
First, a plot of average values for each protocol:
![Protocol averages plot](img/protocol_averages.png)
This reveals a linear distribution of values among the protocols, aside from exponential curves only at the extremes. Perhaps the most interesting finding is a skew toward the higher end of the scale, associated with softness. Even relatively hard, technical protocols appear to have significant soft characteristics.
The protocol value averages have a mean of 5.45 and a median of 5.48. In comparison to the midpoint of 5, the normalized midpoint deviation is 0.11. In comparison, the Pearson coefficient measures the skew at just -0.07, which means that the relative skew of the data is actually slightly downward. So the distribution of protocol values is very balanced but has a consistent upward deviation from the scale's baseline. (These calculations are in `data/synthetic_20251116/readings-analysis.odt[averages]`.)
The protocol value averages have a mean of 5.44 and a median of 5.48. In comparison to the midpoint of 5, the normalized midpoint deviation is 0.11. Skewness is very slight and slightly negative (Fisher moment skew −0.03; Pearson −0.13), meaning the distribution of protocol values is very balanced but has a consistent upward deviation from the scale's baseline. (These calculations come from `scripts/univariate_analysis.py`; summary text in `data/synthetic_1.2.6/analysis/reports/univariate_summary.txt`. Earlier ad-hoc calculations are in `data/synthetic_1.2.6/readings-analysis.odt[averages]`.)
Second, a plot of average values for each gradient (with gaps to indicate the three groupings of gradients):
@@ -161,88 +269,26 @@ Expectations:
* There are some gradients whose values are highly correlated. These might point to redundancies in the bicorder design.
* Some correlations might be revealing about connections in the characteristics of protocols, but these should be considered carefully as they may be the result of design or LLM interpretation.
Claude Code created a `multivariate_analysis.py` tool to conduct this analysis. Usage:
The `multivariate_analysis.py` tool was used to conduct this analysis. Usage:
```bash
# Run all analyses (default)
python3 scripts/multivariate_analysis.py data/synthetic_20251116/readings.csv
python3 scripts/multivariate_analysis.py data/synthetic_1.2.6/readings.csv
# Run specific analyses only
python3 scripts/multivariate_analysis.py data/synthetic_20251116/readings.csv --analyses
python3 scripts/multivariate_analysis.py data/synthetic_1.2.6/readings.csv --analyses
clustering pca
```
Initial manual observations:
* The correlations generally seem predictable; for example, the strongest is between `Design_static_vs_malleable` and `Experience_predictable_vs_emergent`, which is not surprising
* The elite vs. vernacular distinction appears to be the most predictive gradient (`data/synthetic_20251116/analysis/plots/feature_importances.png`)
* The elite vs. vernacular distinction appears to be the most predictive gradient (`data/synthetic_1.2.6/analysis/plots/feature_importances.png`)
![Correlation heatmap](data/synthetic_20251116/analysis/plots/correlation_heatmap_full.png)
![Correlation heatmap](data/synthetic_1.2.6/analysis/plots/correlation_heatmap_full.png)
![Importance ranking](data/synthetic_20251116/analysis/plots/feature_importances.png)
![Importance ranking](data/synthetic_1.2.6/analysis/plots/feature_importances.png)
Claude's interpretation:
> 1. Two Fundamental Protocol Types (K-Means Clustering)
>
> The data reveals two distinct protocol families (216 vs 192 protocols):
>
> Cluster 1: "Vernacular/Emergent Protocols"
> - Examples: Marronage, Songlines, Access-Centered Practices, Ethereum Proof of Work, Sangoma Healing Practices
> - Characteristics:
> - HIGH: elite→vernacular (6.4), malleable (7.3), flocking→swarming (6.4)
> - LOW: self-enforcing→enforced (3.4), sovereign→subsidiary (2.9)
>
> Cluster 2: "Institutional/Standardized Protocols"
> - Examples: ISO standards, Greenwich Mean Time, Building Codes, German Bureaucratic Prose, Royal Access Protocol
> - Characteristics:
> - HIGH: self-enforcing (7.1), sovereign (6.0)
> - LOW: elite→vernacular (1.8), flocking→swarming (2.3), static (3.5)
>
> 2. Key Structural Dimensions (PCA)
>
> Three principal components explain 55% of variance:
>
> PC1 (main axis of variation): Elite/Static/Flocking ↔ Self-enforcing
> - Essentially captures the Vernacular vs. Institutional divide
>
> PC2: Sufficient/Crystallized ↔ Kafka-esque
> - Measures protocol "completeness" vs. bureaucratic nightmare quality
>
> PC3: Universal/Technical/Macro ↔ Particular/Embodied
> - Scale and abstraction level
>
> 3. Strong Correlations (Most Significant Relationships)
>
> 1. Static ↔ Predictable (r=0.61): Unchanging protocols create predictable experiences
> 2. Elite ↔ Self-enforcing (r=-0.58): Elite protocols need external enforcement; vernacular ones self-enforce
> 3. Self-enforcing ↔ Flocking (r=-0.56): Self-enforcing protocols resist swarming dynamics
> 4. Exclusion ↔ Kafka (r=0.52): Exclusionary protocols feel Kafka-esque
>
> 4. Most Discriminative Dimension (Feature Importance)
>
> Design_elite_vs_vernacular (22.7% importance) is by far the most powerful predictor of protocol type, followed by Entanglement_flocking_vs_swarming (13.8%).
>
> 5. Most "Central" Protocols (Network Analysis)
>
> These protocols share the most dimensional similarities with others:
> 1. VPN Usage (Circumvention Protocol) - bridges many protocol types
> 2. Access Check-in - connects accessibility and participation patterns
> 3. Quadratic Voting - spans governance dimensions
>
> 6. Outliers (DBSCAN found 281!)
>
> Most protocols are actually quite unique - DBSCAN identified 281 outliers, suggesting the dataset contains many distinctive protocol configurations that don't fit neat clusters. Only 10 tight sub-clusters exist.
>
> 7. Category Prediction Power
>
> - Design dimensions predict clustering with 90.4% accuracy
> - Entanglement dimensions: 89.2% accuracy
> - Experience dimensions: only 78.3% accuracy
>
> This suggests Design and Entanglement are more fundamental than Experience.
>
> The core insight: Protocols fundamentally divide between vernacular/emergent/malleable forms and institutional/standardized/static forms, with the elite↔vernacular dimension being the strongest predictor of all other characteristics.
Comments:
@@ -262,146 +308,6 @@ Reflections on manual review:
* Should have separate values for "n/a" (0---but that could screw up averages) and both (5)
* Remove the analysis section, or use analyses here for what becomes most meaningful
Claude report on possible improvements:
> 1. Prioritize High-Impact Dimensions ⭐
>
> The current 23 dimensions aren't equally informative. Reorder by importance:
>
> Tier 1 - Critical (>10% importance):
> - Design_elite_vs_vernacular (22.7%) - THE most discriminative dimension
> - Entanglement_flocking_vs_swarming (13.8%)
> - Design_static_vs_malleable (10.2%)
>
> Tier 2 - Important (5-10%):
> - Entanglement_self-enforcing_vs_enforced (9.2%)
> - Entanglement_obligatory_vs_voluntary (8.0%)
> - Experience_exclusion_vs_inclusion (5.9%)
>
> Tier 3 - Supplementary (<5%):
> - All remaining 17 dimensions
>
> Recommendation: Reorganize the tool to present Tier 1 dimensions first, or mark them as "core diagnostics" vs. "supplementary diagnostics."
>
> 2. Consider Reducing Low-Value Dimensions
>
> Several dimensions have low discriminative power AND low variance:
>
> Candidates for removal/consolidation:
> - Entanglement_exclusive_vs_non-exclusive (0.6% importance, σ=1.64, mean=8.5)
> - Responses heavily cluster at "non-exclusive" - not discriminating
> - Design_durable_vs_ephemeral (1.2% importance, σ=2.41)
> - Entanglement_defensible_vs_exposed (1.1% importance, σ=2.41)
>
> Recommendation: Either remove these or combine into composite measures. Going from 23→18 dimensions would reduce analyst burden by ~20% with minimal information loss.
>
> 3. Add Composite Scores 📊
>
> Since PC1 explains the main variance, create derived metrics:
>
> "Protocol Type Score" (based on PC1 loadings):
> Score = elite_vs_vernacular(0.36) + static_vs_malleable(0.33) + flocking_vs_swarming(0.31) - self-enforcing_vs_enforced(0.29)
> - High score = Institutional/Standardized
> - Low score = Vernacular/Emergent
>
> "Protocol Completeness Score" (based on PC2):
> Score = sufficient_vs_insufficient(0.43) + crystallized_vs_contested(0.38) - Kafka_vs_Whitehead(0.36)
> - Measures how "finished" vs. "kafkaesque" a protocol feels
>
> Recommendation: Display these composite scores alongside individual dimensions to provide quick high-level insights.
>
> 4. Highlight Key Correlations 🔗
>
> The tool should alert analysts to important relationships:
>
> Strong positive correlations:
> - Static ↔ Predictable (0.61)
> - Elite ↔ Static (0.53)
> - Exclusion ↔ Kafka (0.52)
>
> Strong negative correlations:
> - Elite ↔ Self-enforcing (-0.58)
> - Self-enforcing ↔ Flocking (-0.56)
>
> Recommendation: When an analyst rates a dimension, show a tooltip: "Protocols
> rated as 'elite' tend to also be 'static' and require 'enforcement'"
>
> 5. Flag Potential Redundancy
>
> Two dimension pairs show moderate correlation within the same category:
>
> 1. Entanglement_abstract_vs_embodied ↔ Entanglement_flocking_vs_swarming (r=-0.56)
> 2. Design_documenting_vs_enabling ↔ Design_static_vs_malleable (r=0.53)
>
> Recommendation: Consider merging these or making one primary and the other optional.
>
> 6. Rebalance Categories ⚖️
>
> Current split: Design (8), Entanglement (8), Experience (7)
>
> Performance by category:
> - Design: 90.4% predictive accuracy
> - Entanglement: 89.2% predictive accuracy
> - Experience: 78.3% predictive accuracy
>
> Recommendation:
> - Strengthen Design (add 1-2 high-variance dimensions)
> - Trim Experience (remove low-performers, down to 5)
> - Result: 9 Design, 8 Entanglement, 5 Experience = 22 dimensions (down from
> 23)
>
> 7. Add Diagnostic Quality Indicators
>
> Based on variance analysis, flag dimensions where responses are too clustered:
>
> - Entanglement_exclusive_vs_non-exclusive: 93% of protocols rate 7-9
> - Design_technical_vs_social: Mean=7.6, heavily skewed toward "social"
>
> Recommendation: Consider revising these gradient definitions or endpoints to achieve better distribution.
>
> 8. Create Shortened Version 🎯
>
> For rapid assessment, create a "Bicorder Core" with just the top 8 dimensions:
>
> 1. elite_vs_vernacular ⭐⭐⭐
> 2. flocking_vs_swarming
> 3. static_vs_malleable
> 4. self-enforcing_vs_enforced
> 5. obligatory_vs_voluntary
> 6. exclusion_vs_inclusion
> 7. universal_vs_particular (high variance)
> 8. explicit_vs_implicit (high variance)
>
> This captures ~65% of the discriminative power in 1/3 the time.
>
> 9. Add Comparison Features
>
> The network analysis shows some protocols are highly "central" (similar to many others):
> - VPN Usage (Circumvention Protocol)
> - Access Check-in
> - Quadratic Voting
>
> Recommendation: After rating a protocol, show: "This protocol is most similar to: [X, Y, Z]" based on dimensional proximity.
> Summary of Recommendations
> Immediate actions:
> 1. ✂️ Remove 3-5 low-value dimensions → 20 dimensions
> 2. 🔄 Reorder dimensions by importance (elite_vs_vernacular first)
> 3. ➕ Add 2 composite scores (Protocol Type, Completeness)
> Medium-term enhancements:
> 4. 🎯 Create "Bicorder Core" (8-dimension quick version)
> 5. 💡 Add contextual tooltips about correlations
> 6. 📊 Show similar protocols after assessment
> This would make the tool ~20% faster to use while maintaining 95%+ of its discriminative power.
Questions:
* What makes the elite/vernacular distinction so "important"? Is it because of some salience of the description?
* Why is the sovereign/subsidiary distinction, which is such a central part of the theorizing here, not more "important"? Would this change with a different description of the values?
### Future work: Description modification
@@ -442,7 +348,7 @@ This simple version-matching approach ensures compatibility without complex stru
### Files
- `bicorder_model.json` (~5KB) - Trained LDA model with coefficients and scaler parameters; read by `bicorder-app` at build time
- `bicorder_model.json` (~5KB) - Trained LDA model with coefficients and scaler parameters; research-only since v1.3.0 (the app no longer reads it)
- `bicorder-app/src/bicorder-classifier.ts` - TypeScript classifier implementation in the web app
- `ascii_bicorder.py` (updated) - Python script now calculates automated analysis values
- `../bicorder.json` (updated) - Added bureaucratic ↔ relational gradient to analysis section
@@ -455,7 +361,7 @@ The calculation happens automatically when generating bicorder output:
python3 ascii_bicorder.py bicorder.json bicorder.txt
```
For web integration, see `INTEGRATION_GUIDE.md`. The app (`bicorder-app/`) has its own classifier implementation and reads `bicorder_model.json` from this directory at build time.
For the history of the web integration, see `INTEGRATION_GUIDE.md`. The app (`bicorder-app/`) used to deploy its own TypeScript classifier and read `bicorder_model.json` from this directory at build time; that integration was removed in v1.3.0, so this model is now research-only.
### Key Features
+6 -6
View File
@@ -7,7 +7,7 @@ Run these tests in order to verify the refactored code works correctly.
Test that prompts are generated correctly with protocol context:
```bash
python3 scripts/bicorder_query.py data/synthetic_20251116/protocols_edited.csv 1 --dry-run | head -80
python3 scripts/bicorder_query.py data/protocols_edited.csv 1 --dry-run | head -80
```
**Expected result:**
@@ -21,7 +21,7 @@ python3 scripts/bicorder_query.py data/synthetic_20251116/protocols_edited.csv 1
Check that the analyze script still creates proper CSV structure:
```bash
python3 scripts/bicorder_analyze.py data/synthetic_20251116/protocols_edited.csv -o test_output.csv
python3 scripts/bicorder_analyze.py data/protocols_edited.csv -o test_output.csv
head -1 test_output.csv | tr ',' '\n' | grep -E "(explicit|precise|elite)" | head -5
```
@@ -76,7 +76,7 @@ llm logs list | grep -i bicorder
Test batch processing on rows 1-3:
```bash
python3 scripts/bicorder_batch.py data/synthetic_20251116/protocols_edited.csv -o test_batch_output.csv --start 1 --end 3 -m gpt-4o-mini
python3 scripts/bicorder_batch.py data/protocols_edited.csv -o test_batch_output.csv --start 1 --end 3 -m gpt-4o-mini
```
**Expected result:**
@@ -106,7 +106,7 @@ with open('test_batch_output.csv') as f:
Test that model parameter works in dry run:
```bash
python3 scripts/bicorder_query.py data/synthetic_20251116/protocols_edited.csv 5 --dry-run -m mistral | head -50
python3 scripts/bicorder_query.py data/protocols_edited.csv 5 --dry-run -m mistral | head -50
```
**Expected result:**
@@ -129,11 +129,11 @@ Compare the new standalone prompts vs old system prompt approach:
```bash
# New approach - protocol context in each prompt
python3 scripts/bicorder_query.py data/synthetic_20251116/protocols_edited.csv 1 --dry-run | grep -A 5 "Analyze this protocol"
python3 scripts/bicorder_query.py data/protocols_edited.csv 1 --dry-run | grep -A 5 "Analyze this protocol"
# Old approach would have had protocol in system prompt only (no longer used)
# Verify that protocol context appears in EVERY gradient prompt
python3 scripts/bicorder_query.py data/synthetic_20251116/protocols_edited.csv 1 --dry-run | grep -c "Analyze this protocol"
python3 scripts/bicorder_query.py data/protocols_edited.csv 1 --dry-run | grep -c "Analyze this protocol"
```
**Expected result:**
+25 -11
View File
@@ -21,9 +21,13 @@ The scripts automatically draw the gradients from the current state of the [bico
6. **scripts/multivariate_analysis.py** - Run clustering, PCA, correlation, and feature importance analysis on a readings CSV
7. **scripts/lda_visualization.py** - Generate LDA cluster separation plot and projection data
8. **scripts/classify_readings.py** - Apply the synthetic-trained LDA classifier to all readings; saves `analysis/classifications.csv`
9. **scripts/visualize_clusters.py** - Additional cluster visualizations
10. **scripts/export_model_for_js.py** - Export trained model to `bicorder_model.json` (read by `bicorder-app` at build time)
8. **scripts/classify_readings.py** - Apply the synthetic-trained LDA classifier to all readings; saves `analysis/classifications.csv` (training data is auto-matched to the input's recorded bicorder version)
9. **scripts/univariate_analysis.py** - Per-protocol and per-gradient averages, distributions, and summary stats (replaces the ad-hoc averages workflow)
10. **scripts/compare_analyses.py** - Compare readings CSVs to a reference (Euclidean distance, RMSE, correlation); canonicalizes renamed gradients so versions can be compared
11. **scripts/visualize_clusters.py** - Additional cluster visualizations
12. **scripts/export_model_for_js.py** - Export trained model to `bicorder_model.json` (research-only since v1.3.0 — the app integration was removed)
Version-agnostic helpers shared by these scripts live in **scripts/bicorder_common.py**: the historical gradient rename map, version detection (from the `bicorder_version`/`version` column, falling back to the `data/<type>_<version>/` directory convention), and training-CSV selection. When gradients are renamed in `../bicorder.json`, update `COLUMN_RENAMES` in that one module.
## Syncing a manual readings dataset
@@ -46,11 +50,21 @@ python3 scripts/multivariate_analysis.py data/manual_20260320/readings.csv \
# LDA visualization (cluster separation plot)
python3 scripts/lda_visualization.py data/manual_20260320/readings.csv
# Classify all readings (uses synthetic dataset as training data by default)
# Classify all readings (training data auto-matched to the dataset's bicorder version)
python3 scripts/classify_readings.py data/manual_20260320/readings.csv
# Univariate averages (protocol and gradient plots + summary stats)
python3 scripts/univariate_analysis.py data/manual_20260320/readings.csv
# ... add --img to also publish its three summary charts into analysis/img/ (linked from README)
# Cluster overlaid PCA/t-SNE/UMAP plots (after multivariate analysis)
python3 scripts/visualize_clusters.py data/manual_20260320/readings.csv
# Compare two runs, including across bicorder versions (renamed gradients are aligned)
python3 scripts/compare_analyses.py data/synthetic_1.4.0/readings.csv data/synthetic_1.2.6/readings.csv
```
Use `--min-coverage` (0.0–1.0) to drop dimension columns below the given coverage fraction before analysis. This is important for datasets with many shortform readings where most dimensions are sparsely filled.
Use `--min-coverage` (0.0–1.0) to drop dimension columns below the given coverage fraction before analysis. This is important for datasets with many shortform readings where most dimensions are sparsely filled. `classify_readings.py` still accepts an explicit `--training` CSV to override the automatic selection.
## Converting JSON reading files to CSV
@@ -68,7 +82,7 @@ python3 scripts/json_to_csv.py data/manual_20260320/json/ \
### Process All Protocols with One Command
```bash
python3 scripts/bicorder_batch.py data/synthetic_20251116/protocols_edited.csv -o analysis_output.csv
python3 scripts/bicorder_batch.py data/protocols_edited.csv -o analysis_output.csv
```
This will:
@@ -81,13 +95,13 @@ This will:
```bash
# Process only rows 1-5 (useful for testing)
python3 scripts/bicorder_batch.py data/synthetic_20251116/protocols_edited.csv -o analysis_output.csv --start 1 --end 5
python3 scripts/bicorder_batch.py data/protocols_edited.csv -o analysis_output.csv --start 1 --end 5
# Use specific LLM model
python3 scripts/bicorder_batch.py data/synthetic_20251116/protocols_edited.csv -o analysis_output.csv -m mistral
python3 scripts/bicorder_batch.py data/protocols_edited.csv -o analysis_output.csv -m mistral
# Add analyst metadata
python3 scripts/bicorder_batch.py data/synthetic_20251116/protocols_edited.csv -o analysis_output.csv \
python3 scripts/bicorder_batch.py data/protocols_edited.csv -o analysis_output.csv \
-a "Your Name" -s "Your analytical standpoint"
```
@@ -100,12 +114,12 @@ python3 scripts/bicorder_batch.py data/synthetic_20251116/protocols_edited.csv -
Create a CSV with empty gradient columns:
```bash
python3 scripts/bicorder_analyze.py data/synthetic_20251116/protocols_edited.csv -o analysis_output.csv
python3 scripts/bicorder_analyze.py data/protocols_edited.csv -o analysis_output.csv
```
Optional: Add analyst metadata:
```bash
python3 scripts/bicorder_analyze.py data/synthetic_20251116/protocols_edited.csv -o analysis_output.csv \
python3 scripts/bicorder_analyze.py data/protocols_edited.csv -o analysis_output.csv \
-a "Your Name" -s "Your analytical standpoint"
```
File renamed without changes.
Internal Server Error - Gitea
500 Internal Server Error

An error occurred:

An error occurred

Gitea Version: 28.0.0