Files
protocol-bicorder/analysis/INTEGRATION_GUIDE.md
T
Protocolbot 6ae77a4f9b refactor: reorganize analysis data to support multiple runs
Strategy: key runs by bicorder version (not date), promote shared inputs
to analysis/data/, and stamp every output with its bicorder_version so
a re-run on edited gradients is self-describing.

Data layout:
- Promote the shared protocol inputs out of the run directory:
    analysis/data/protocols_edited.csv  (411 cleaned protocols)
    analysis/data/protocols_raw.csv     (774 un-cleaned entries)
- Rename the v1.2.6 synthetic run:
    data/synthetic_20251116/ -> data/synthetic_1.2.6/
  so the bicorder version it was scored against is explicit (gradient
  structure changes between versions make date-based names ambiguous)

Provenance:
- bicorder_analyze.py now writes a 'bicorder_version' column into every
  output readings.csv, recording which gradient structure produced it

Scripts:
- Update the real code defaults that pointed at the old run path
  (bicorder_classifier.py, classify_readings.py, sync_readings.sh,
  compare_analyses.py) and refresh docstring/help examples
- Remove a stray committed __pycache__/.pyc

Docs: analysis/README.md documents the new layout + how to add a run;
WORKFLOW.md, TEST_COMMANDS.md, INTEGRATION_GUIDE.md paths updated.
2026-09-23 14:23:17 -06:00

4.1 KiB
Raw Blame History

Bicorder Classifier — Research Notes

Status: removed from the tool (v1.3.0). The formal/informal (bureaucratic↔relational) LDA analysis was removed from the bicorder itself in v1.3.0. The cluster classification survives as research only — the scripts in this directory can still train and apply the model to datasets, but the web app and ascii_bicorder.py no longer consume it. This document is retained as a historical record of how the integration worked and how to reproduce the research analysis.

Overview

The analysis directory contains a cluster classification system that was previously integrated into the Bicorder web application to provide:

  1. Real-time cluster prediction as users filled out diagnostics
  2. Smart form selection (short vs. long form based on classification confidence)
  3. Visual feedback showing protocol family positioning

Original Design Philosophy

Version-based compatibility: The model included a bicorder_version field. The classifier checked that versions matched. When bicorder.json structure changed:

  1. The version number in bicorder.json was incremented
  2. The model was retrained with python3 scripts/export_model_for_js.py data/synthetic_1.2.6/readings.csv
  3. The new model had the updated version

Files (research-only now)

  • bicorder_model.json - Trained model parameters (~5KB), trained on the synthetic dataset (bicorder v1.2.6 structure — stale relative to v1.3.0; retrain before applying to new readings)
  • scripts/bicorder_classifier.py - Python classifier (used by classify_readings.py)
  • scripts/export_model_for_js.py - Retrain and export the model to JSON
  • scripts/classify_readings.py - Apply the classifier to a readings CSV

Reproducing the research analysis

# Retrain the model on a (new) synthetic dataset
python3 scripts/export_model_for_js.py data/<dataset>/readings.csv

# Classify a dataset's readings
python3 scripts/classify_readings.py data/<dataset>/readings.csv --training data/<dataset>/readings.csv

The classifier predicts which of two protocol families a reading belongs to:

  • Cluster 1: Relational/Cultural — community-based, emergent, voluntary protocols
  • Cluster 2: Institutional/Bureaucratic — formal, top-down, externally enforced protocols

See analysis/README.md for the full multivariate analysis these clusters came from.

Historical integration patterns

The removed web-app integration supported progressive classification display, smart form selection (suggesting the long form when classification confidence was low), short-form optimization around the most discriminative dimensions, and readiness checks. The Python classifier API remains:

  • predict(ratings, options) → cluster, clusterName, confidence, completeness, recommendedForm (detailed mode adds ldaScore, distanceToBoundary, dimension counts)
  • explain_classification(ratings) → human-readable explanation
  • get_key_dimensions() → the shortform/key dimensions from bicorder.json
  • assess_short_form_readiness(ratings) (TS only, removed) — the Python recommended_form field remains

The shortform gradients themselves are defined in bicorder.json (shortform: true), derived from the original feature-importance analysis — that part of the research lives on in the tool.

Why it was removed

  • The LDA sign convention was inverted in ascii_bicorder.py (never caught there because a term-rename also silently disabled the calculation), while the web app had been separately fixed — two divergent implementations.
  • Compressing a two-family classification into a 1–9 gradient was semantically awkward and produced recurring bugs (see commit fd556d9).
  • The version-mismatch handling differed between implementations (Python skipped; TypeScript continued with a stale model).
  • The two-families finding is a research result, not a diagnostic — it belongs in analysis, not in the instrument itself.

The form-recommendation feature (suggesting long form when classification confidence was low) was also removed. Shortform/longform selection is now entirely the analyst's choice.