Strategy: key runs by bicorder version (not date), promote shared inputs
to analysis/data/, and stamp every output with its bicorder_version so
a re-run on edited gradients is self-describing.
Data layout:
- Promote the shared protocol inputs out of the run directory:
analysis/data/protocols_edited.csv (411 cleaned protocols)
analysis/data/protocols_raw.csv (774 un-cleaned entries)
- Rename the v1.2.6 synthetic run:
data/synthetic_20251116/ -> data/synthetic_1.2.6/
so the bicorder version it was scored against is explicit (gradient
structure changes between versions make date-based names ambiguous)
Provenance:
- bicorder_analyze.py now writes a 'bicorder_version' column into every
output readings.csv, recording which gradient structure produced it
Scripts:
- Update the real code defaults that pointed at the old run path
(bicorder_classifier.py, classify_readings.py, sync_readings.sh,
compare_analyses.py) and refresh docstring/help examples
- Remove a stray committed __pycache__/.pyc
Docs: analysis/README.md documents the new layout + how to add a run;
WORKFLOW.md, TEST_COMMANDS.md, INTEGRATION_GUIDE.md paths updated.
81 lines
4.1 KiB
Markdown
81 lines
4.1 KiB
Markdown
# Bicorder Classifier — Research Notes
|
||
|
||
> **Status: removed from the tool (v1.3.0).** The formal/informal (bureaucratic↔relational)
|
||
> LDA analysis was removed from the bicorder itself in v1.3.0. The cluster
|
||
> classification survives as **research only** — the scripts in this directory
|
||
> can still train and apply the model to datasets, but the web app and
|
||
> `ascii_bicorder.py` no longer consume it. This document is retained as a
|
||
> historical record of how the integration worked and how to reproduce the
|
||
> research analysis.
|
||
|
||
## Overview
|
||
|
||
The analysis directory contains a cluster classification system that was
|
||
previously integrated into the Bicorder web application to provide:
|
||
|
||
1. **Real-time cluster prediction** as users filled out diagnostics
|
||
2. **Smart form selection** (short vs. long form based on classification confidence)
|
||
3. **Visual feedback** showing protocol family positioning
|
||
|
||
## Original Design Philosophy
|
||
|
||
**Version-based compatibility**: The model included a `bicorder_version` field.
|
||
The classifier checked that versions matched. When bicorder.json structure changed:
|
||
1. The version number in bicorder.json was incremented
|
||
2. The model was retrained with `python3 scripts/export_model_for_js.py data/synthetic_1.2.6/readings.csv`
|
||
3. The new model had the updated version
|
||
|
||
## Files (research-only now)
|
||
|
||
- `bicorder_model.json` - Trained model parameters (~5KB), trained on the synthetic dataset (bicorder v1.2.6 structure — **stale** relative to v1.3.0; retrain before applying to new readings)
|
||
- `scripts/bicorder_classifier.py` - Python classifier (used by `classify_readings.py`)
|
||
- `scripts/export_model_for_js.py` - Retrain and export the model to JSON
|
||
- `scripts/classify_readings.py` - Apply the classifier to a readings CSV
|
||
|
||
## Reproducing the research analysis
|
||
|
||
```bash
|
||
# Retrain the model on a (new) synthetic dataset
|
||
python3 scripts/export_model_for_js.py data/<dataset>/readings.csv
|
||
|
||
# Classify a dataset's readings
|
||
python3 scripts/classify_readings.py data/<dataset>/readings.csv --training data/<dataset>/readings.csv
|
||
```
|
||
|
||
The classifier predicts which of two protocol families a reading belongs to:
|
||
- **Cluster 1: Relational/Cultural** — community-based, emergent, voluntary protocols
|
||
- **Cluster 2: Institutional/Bureaucratic** — formal, top-down, externally enforced protocols
|
||
|
||
See `analysis/README.md` for the full multivariate analysis these clusters came from.
|
||
|
||
## Historical integration patterns
|
||
|
||
The removed web-app integration supported progressive classification display,
|
||
smart form selection (suggesting the long form when classification confidence
|
||
was low), short-form optimization around the most discriminative dimensions,
|
||
and readiness checks. The Python classifier API remains:
|
||
|
||
- `predict(ratings, options)` → cluster, clusterName, confidence, completeness, recommendedForm (detailed mode adds ldaScore, distanceToBoundary, dimension counts)
|
||
- `explain_classification(ratings)` → human-readable explanation
|
||
- `get_key_dimensions()` → the shortform/key dimensions from bicorder.json
|
||
- `assess_short_form_readiness(ratings)` (TS only, removed) — the Python `recommended_form` field remains
|
||
|
||
The shortform gradients themselves are defined in `bicorder.json`
|
||
(`shortform: true`), derived from the original feature-importance analysis —
|
||
that part of the research lives on in the tool.
|
||
|
||
## Why it was removed
|
||
|
||
- The LDA sign convention was inverted in `ascii_bicorder.py` (never caught
|
||
there because a term-rename also silently disabled the calculation), while
|
||
the web app had been separately fixed — two divergent implementations.
|
||
- Compressing a two-family classification into a 1–9 gradient was semantically
|
||
awkward and produced recurring bugs (see commit `fd556d9`).
|
||
- The version-mismatch handling differed between implementations (Python
|
||
skipped; TypeScript continued with a stale model).
|
||
- The two-families finding is a research result, not a diagnostic — it belongs
|
||
in analysis, not in the instrument itself.
|
||
|
||
The form-recommendation feature (suggesting long form when classification
|
||
confidence was low) was also removed. Shortform/longform selection is now
|
||
entirely the analyst's choice. |