Files
protocol-bicorder/analysis/INTEGRATION_GUIDE.md
T
Protocolbot e7d6465ceb docs: mark classifier integration historical; fix shortform count
- INTEGRATION_GUIDE.md: rewritten as research notes — how to reproduce the
  cluster classification with the analysis scripts, and why it was removed
  from the tool in v1.3.0
- analysis/README.md: integration section marked historical
- bicorder-app/README.md: shortform is 9 gradients, not 10
2026-09-23 07:58:17 -06:00

81 lines
4.1 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Bicorder Classifier — Research Notes
> **Status: removed from the tool (v1.3.0).** The formal/informal (bureaucratic↔relational)
> LDA analysis was removed from the bicorder itself in v1.3.0. The cluster
> classification survives as **research only** — the scripts in this directory
> can still train and apply the model to datasets, but the web app and
> `ascii_bicorder.py` no longer consume it. This document is retained as a
> historical record of how the integration worked and how to reproduce the
> research analysis.
## Overview
The analysis directory contains a cluster classification system that was
previously integrated into the Bicorder web application to provide:
1. **Real-time cluster prediction** as users filled out diagnostics
2. **Smart form selection** (short vs. long form based on classification confidence)
3. **Visual feedback** showing protocol family positioning
## Original Design Philosophy
**Version-based compatibility**: The model included a `bicorder_version` field.
The classifier checked that versions matched. When bicorder.json structure changed:
1. The version number in bicorder.json was incremented
2. The model was retrained with `python3 scripts/export_model_for_js.py data/synthetic_20251116/readings.csv`
3. The new model had the updated version
## Files (research-only now)
- `bicorder_model.json` - Trained model parameters (~5KB), trained on the synthetic dataset (bicorder v1.2.6 structure — **stale** relative to v1.3.0; retrain before applying to new readings)
- `scripts/bicorder_classifier.py` - Python classifier (used by `classify_readings.py`)
- `scripts/export_model_for_js.py` - Retrain and export the model to JSON
- `scripts/classify_readings.py` - Apply the classifier to a readings CSV
## Reproducing the research analysis
```bash
# Retrain the model on a (new) synthetic dataset
python3 scripts/export_model_for_js.py data/<dataset>/readings.csv
# Classify a dataset's readings
python3 scripts/classify_readings.py data/<dataset>/readings.csv --training data/<dataset>/readings.csv
```
The classifier predicts which of two protocol families a reading belongs to:
- **Cluster 1: Relational/Cultural** — community-based, emergent, voluntary protocols
- **Cluster 2: Institutional/Bureaucratic** — formal, top-down, externally enforced protocols
See `analysis/README.md` for the full multivariate analysis these clusters came from.
## Historical integration patterns
The removed web-app integration supported progressive classification display,
smart form selection (suggesting the long form when classification confidence
was low), short-form optimization around the most discriminative dimensions,
and readiness checks. The Python classifier API remains:
- `predict(ratings, options)` → cluster, clusterName, confidence, completeness, recommendedForm (detailed mode adds ldaScore, distanceToBoundary, dimension counts)
- `explain_classification(ratings)` → human-readable explanation
- `get_key_dimensions()` → the shortform/key dimensions from bicorder.json
- `assess_short_form_readiness(ratings)` (TS only, removed) — the Python `recommended_form` field remains
The shortform gradients themselves are defined in `bicorder.json`
(`shortform: true`), derived from the original feature-importance analysis —
that part of the research lives on in the tool.
## Why it was removed
- The LDA sign convention was inverted in `ascii_bicorder.py` (never caught
there because a term-rename also silently disabled the calculation), while
the web app had been separately fixed — two divergent implementations.
- Compressing a two-family classification into a 1–9 gradient was semantically
awkward and produced recurring bugs (see commit `fd556d9`).
- The version-mismatch handling differed between implementations (Python
skipped; TypeScript continued with a stale model).
- The two-families finding is a research result, not a diagnostic — it belongs
in analysis, not in the instrument itself.
The form-recommendation feature (suggesting long form when classification
confidence was low) was also removed. Shortform/longform selection is now
entirely the analyst's choice.