Files
protocol-bicorder/analysis/INTEGRATION_GUIDE.md
T
Protocolbot e7d6465ceb docs: mark classifier integration historical; fix shortform count
- INTEGRATION_GUIDE.md: rewritten as research notes — how to reproduce the
  cluster classification with the analysis scripts, and why it was removed
  from the tool in v1.3.0
- analysis/README.md: integration section marked historical
- bicorder-app/README.md: shortform is 9 gradients, not 10
2026-09-23 07:58:17 -06:00

4.1 KiB
Raw Blame History

Bicorder Classifier — Research Notes

Status: removed from the tool (v1.3.0). The formal/informal (bureaucratic↔relational) LDA analysis was removed from the bicorder itself in v1.3.0. The cluster classification survives as research only — the scripts in this directory can still train and apply the model to datasets, but the web app and ascii_bicorder.py no longer consume it. This document is retained as a historical record of how the integration worked and how to reproduce the research analysis.

Overview

The analysis directory contains a cluster classification system that was previously integrated into the Bicorder web application to provide:

  1. Real-time cluster prediction as users filled out diagnostics
  2. Smart form selection (short vs. long form based on classification confidence)
  3. Visual feedback showing protocol family positioning

Original Design Philosophy

Version-based compatibility: The model included a bicorder_version field. The classifier checked that versions matched. When bicorder.json structure changed:

  1. The version number in bicorder.json was incremented
  2. The model was retrained with python3 scripts/export_model_for_js.py data/synthetic_20251116/readings.csv
  3. The new model had the updated version

Files (research-only now)

  • bicorder_model.json - Trained model parameters (~5KB), trained on the synthetic dataset (bicorder v1.2.6 structure — stale relative to v1.3.0; retrain before applying to new readings)
  • scripts/bicorder_classifier.py - Python classifier (used by classify_readings.py)
  • scripts/export_model_for_js.py - Retrain and export the model to JSON
  • scripts/classify_readings.py - Apply the classifier to a readings CSV

Reproducing the research analysis

# Retrain the model on a (new) synthetic dataset
python3 scripts/export_model_for_js.py data/<dataset>/readings.csv

# Classify a dataset's readings
python3 scripts/classify_readings.py data/<dataset>/readings.csv --training data/<dataset>/readings.csv

The classifier predicts which of two protocol families a reading belongs to:

  • Cluster 1: Relational/Cultural — community-based, emergent, voluntary protocols
  • Cluster 2: Institutional/Bureaucratic — formal, top-down, externally enforced protocols

See analysis/README.md for the full multivariate analysis these clusters came from.

Historical integration patterns

The removed web-app integration supported progressive classification display, smart form selection (suggesting the long form when classification confidence was low), short-form optimization around the most discriminative dimensions, and readiness checks. The Python classifier API remains:

  • predict(ratings, options) → cluster, clusterName, confidence, completeness, recommendedForm (detailed mode adds ldaScore, distanceToBoundary, dimension counts)
  • explain_classification(ratings) → human-readable explanation
  • get_key_dimensions() → the shortform/key dimensions from bicorder.json
  • assess_short_form_readiness(ratings) (TS only, removed) — the Python recommended_form field remains

The shortform gradients themselves are defined in bicorder.json (shortform: true), derived from the original feature-importance analysis — that part of the research lives on in the tool.

Why it was removed

  • The LDA sign convention was inverted in ascii_bicorder.py (never caught there because a term-rename also silently disabled the calculation), while the web app had been separately fixed — two divergent implementations.
  • Compressing a two-family classification into a 1–9 gradient was semantically awkward and produced recurring bugs (see commit fd556d9).
  • The version-mismatch handling differed between implementations (Python skipped; TypeScript continued with a stale model).
  • The two-families finding is a research result, not a diagnostic — it belongs in analysis, not in the instrument itself.

The form-recommendation feature (suggesting long form when classification confidence was low) was also removed. Shortform/longform selection is now entirely the analyst's choice.