Files
protocol-bicorder/analysis/INTEGRATION_GUIDE.md
T
Protocolbot 6ae77a4f9b refactor: reorganize analysis data to support multiple runs
Strategy: key runs by bicorder version (not date), promote shared inputs
to analysis/data/, and stamp every output with its bicorder_version so
a re-run on edited gradients is self-describing.

Data layout:
- Promote the shared protocol inputs out of the run directory:
    analysis/data/protocols_edited.csv  (411 cleaned protocols)
    analysis/data/protocols_raw.csv     (774 un-cleaned entries)
- Rename the v1.2.6 synthetic run:
    data/synthetic_20251116/ -> data/synthetic_1.2.6/
  so the bicorder version it was scored against is explicit (gradient
  structure changes between versions make date-based names ambiguous)

Provenance:
- bicorder_analyze.py now writes a 'bicorder_version' column into every
  output readings.csv, recording which gradient structure produced it

Scripts:
- Update the real code defaults that pointed at the old run path
  (bicorder_classifier.py, classify_readings.py, sync_readings.sh,
  compare_analyses.py) and refresh docstring/help examples
- Remove a stray committed __pycache__/.pyc

Docs: analysis/README.md documents the new layout + how to add a run;
WORKFLOW.md, TEST_COMMANDS.md, INTEGRATION_GUIDE.md paths updated.
2026-09-23 14:23:17 -06:00

81 lines
4.1 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Bicorder Classifier — Research Notes
> **Status: removed from the tool (v1.3.0).** The formal/informal (bureaucratic↔relational)
> LDA analysis was removed from the bicorder itself in v1.3.0. The cluster
> classification survives as **research only** — the scripts in this directory
> can still train and apply the model to datasets, but the web app and
> `ascii_bicorder.py` no longer consume it. This document is retained as a
> historical record of how the integration worked and how to reproduce the
> research analysis.
## Overview
The analysis directory contains a cluster classification system that was
previously integrated into the Bicorder web application to provide:
1. **Real-time cluster prediction** as users filled out diagnostics
2. **Smart form selection** (short vs. long form based on classification confidence)
3. **Visual feedback** showing protocol family positioning
## Original Design Philosophy
**Version-based compatibility**: The model included a `bicorder_version` field.
The classifier checked that versions matched. When bicorder.json structure changed:
1. The version number in bicorder.json was incremented
2. The model was retrained with `python3 scripts/export_model_for_js.py data/synthetic_1.2.6/readings.csv`
3. The new model had the updated version
## Files (research-only now)
- `bicorder_model.json` - Trained model parameters (~5KB), trained on the synthetic dataset (bicorder v1.2.6 structure — **stale** relative to v1.3.0; retrain before applying to new readings)
- `scripts/bicorder_classifier.py` - Python classifier (used by `classify_readings.py`)
- `scripts/export_model_for_js.py` - Retrain and export the model to JSON
- `scripts/classify_readings.py` - Apply the classifier to a readings CSV
## Reproducing the research analysis
```bash
# Retrain the model on a (new) synthetic dataset
python3 scripts/export_model_for_js.py data/<dataset>/readings.csv
# Classify a dataset's readings
python3 scripts/classify_readings.py data/<dataset>/readings.csv --training data/<dataset>/readings.csv
```
The classifier predicts which of two protocol families a reading belongs to:
- **Cluster 1: Relational/Cultural** — community-based, emergent, voluntary protocols
- **Cluster 2: Institutional/Bureaucratic** — formal, top-down, externally enforced protocols
See `analysis/README.md` for the full multivariate analysis these clusters came from.
## Historical integration patterns
The removed web-app integration supported progressive classification display,
smart form selection (suggesting the long form when classification confidence
was low), short-form optimization around the most discriminative dimensions,
and readiness checks. The Python classifier API remains:
- `predict(ratings, options)` → cluster, clusterName, confidence, completeness, recommendedForm (detailed mode adds ldaScore, distanceToBoundary, dimension counts)
- `explain_classification(ratings)` → human-readable explanation
- `get_key_dimensions()` → the shortform/key dimensions from bicorder.json
- `assess_short_form_readiness(ratings)` (TS only, removed) — the Python `recommended_form` field remains
The shortform gradients themselves are defined in `bicorder.json`
(`shortform: true`), derived from the original feature-importance analysis —
that part of the research lives on in the tool.
## Why it was removed
- The LDA sign convention was inverted in `ascii_bicorder.py` (never caught
there because a term-rename also silently disabled the calculation), while
the web app had been separately fixed — two divergent implementations.
- Compressing a two-family classification into a 1–9 gradient was semantically
awkward and produced recurring bugs (see commit `fd556d9`).
- The version-mismatch handling differed between implementations (Python
skipped; TypeScript continued with a stale model).
- The two-families finding is a research result, not a diagnostic — it belongs
in analysis, not in the instrument itself.
The form-recommendation feature (suggesting long form when classification
confidence was low) was also removed. Shortform/longform selection is now
entirely the analyst's choice.