docs: mark classifier integration historical; fix shortform count

- INTEGRATION_GUIDE.md: rewritten as research notes — how to reproduce the
  cluster classification with the analysis scripts, and why it was removed
  from the tool in v1.3.0
- analysis/README.md: integration section marked historical
- bicorder-app/README.md: shortform is 9 gradients, not 10
This commit is contained in:
Protocolbot committed 2026-09-23 07:58:17 -06:00
1 parent 8c40ca076b
commit e7d6465ceb
3 files changed
+74 -374

No files matched your search

+66 -371
View File
@@ -1,386 +1,81 @@
# Bicorder Classifier Integration Guide
# Bicorder Classifier — Research Notes
> **Status: removed from the tool (v1.3.0).** The formal/informal (bureaucratic↔relational)
> LDA analysis was removed from the bicorder itself in v1.3.0. The cluster
> classification survives as **research only** — the scripts in this directory
> can still train and apply the model to datasets, but the web app and
> `ascii_bicorder.py` no longer consume it. This document is retained as a
> historical record of how the integration worked and how to reproduce the
> research analysis.
## Overview
This guide explains how to integrate the cluster classification system into the Bicorder web application to provide:
The analysis directory contains a cluster classification system that was
previously integrated into the Bicorder web application to provide:
1. **Real-time cluster prediction** as users fill out diagnostics
1. **Real-time cluster prediction** as users filled out diagnostics
2. **Smart form selection** (short vs. long form based on classification confidence)
3. **Visual feedback** showing protocol family positioning
## Design Philosophy
## Original Design Philosophy
**Version-based compatibility**: The model includes a `bicorder_version` field. The classifier checks that versions match. When bicorder.json structure changes:
1. Increment the version number in bicorder.json
2. Retrain the model with `python3 scripts/export_model_for_js.py data/synthetic_20251116/readings.csv`
3. The new model will have the updated version
**Version-based compatibility**: The model included a `bicorder_version` field.
The classifier checked that versions matched. When bicorder.json structure changed:
1. The version number in bicorder.json was incremented
2. The model was retrained with `python3 scripts/export_model_for_js.py data/synthetic_20251116/readings.csv`
3. The new model had the updated version
This ensures the web app and model stay in sync without complex backward compatibility.
## Files (research-only now)
## Files
- `bicorder_model.json` - Trained model parameters (~5KB), trained on the synthetic dataset (bicorder v1.2.6 structure — **stale** relative to v1.3.0; retrain before applying to new readings)
- `scripts/bicorder_classifier.py` - Python classifier (used by `classify_readings.py`)
- `scripts/export_model_for_js.py` - Retrain and export the model to JSON
- `scripts/classify_readings.py` - Apply the classifier to a readings CSV
- `bicorder_model.json` - Trained model parameters (~5KB); read by `bicorder-app` at build time from `../analysis/bicorder_model.json`
- `bicorder-app/src/bicorder-classifier.ts` - TypeScript classifier implementation (lives in the app, not here)
The model is the only artifact produced by this analysis directory that the app consumes. Regenerate it after re-running analysis on the synthetic dataset:
## Reproducing the research analysis
```bash
python3 scripts/export_model_for_js.py data/synthetic_20251116/readings.csv
# Retrain the model on a (new) synthetic dataset
python3 scripts/export_model_for_js.py data/<dataset>/readings.csv
# Classify a dataset's readings
python3 scripts/classify_readings.py data/<dataset>/readings.csv --training data/<dataset>/readings.csv
```
## Quick Start
### Basic Usage
```javascript
import { loadClassifier } from './lib/bicorder-classifier.js';
// Load model once at app startup
const classifier = await loadClassifier('/bicorder_model.json');
// As user fills in diagnostic form
function onDimensionChange(dimensionName, value) {
const currentRatings = getCurrentFormValues(); // Your form state
const result = classifier.predict(currentRatings);
console.log(`Cluster: ${result.clusterName}`);
console.log(`Confidence: ${result.confidence}%`);
console.log(`Recommend: ${result.recommendedForm} form`);
updateUI(result);
}
```
## Integration Patterns
### Pattern 1: Progressive Classification Display
Show classification results as the user fills out the form:
```javascript
// React/Svelte component example
function DiagnosticForm() {
const [ratings, setRatings] = useState({});
const [classification, setClassification] = useState(null);
useEffect(() => {
if (Object.keys(ratings).length > 0) {
const result = classifier.predict(ratings);
setClassification(result);
}
}, [ratings]);
return (
<div>
<DiagnosticQuestions onChange={setRatings} />
{classification && (
<ClassificationIndicator
cluster={classification.clusterName}
confidence={classification.confidence}
completeness={classification.completeness}
/>
)}
</div>
);
}
```
### Pattern 2: Smart Form Selection
Automatically switch between short and long forms:
```javascript
function DiagnosticWizard() {
const [ratings, setRatings] = useState({});
function handleDimensionComplete(dimension, value) {
const newRatings = { ...ratings, [dimension]: value };
setRatings(newRatings);
// Check if we should switch forms
const result = classifier.predict(newRatings);
if (result.recommendedForm === 'long' && currentForm === 'short') {
showFormSwitchPrompt(
'Your protocol shows characteristics of both families. ' +
'Would you like to use the detailed form for better classification?'
);
}
}
return <Form onDimensionComplete={handleDimensionComplete} />;
}
```
### Pattern 3: Short Form Optimization
Only ask the 8 most discriminative dimensions for quick classification:
```javascript
const shortFormDimensions = classifier.getKeyDimensions();
// Returns:
// [
// 'Design_elite_vs_vernacular',
// 'Entanglement_flocking_vs_swarming',
// 'Design_static_vs_malleable',
// 'Entanglement_obligatory_vs_voluntary',
// 'Entanglement_self-enforcing_vs_enforced',
// 'Design_explicit_vs_implicit',
// 'Entanglement_sovereign_vs_subsidiary',
// 'Design_technical_vs_social',
// ]
function ShortForm() {
return (
<div>
<h2>Quick Classification (8 questions)</h2>
{shortFormDimensions.map(dim => (
<DimensionSlider key={dim} dimension={dim} />
))}
</div>
);
}
```
### Pattern 4: Readiness Check
Check if user has provided enough data for reliable classification:
```javascript
function ClassificationStatus() {
const assessment = classifier.assessShortFormReadiness(ratings);
if (!assessment.ready) {
return (
<div className="status-warning">
<p>
Need {assessment.keyDimensionsTotal - assessment.keyDimensionsProvided} more
key dimensions for reliable classification ({assessment.coverage}% complete)
</p>
<ul>
{assessment.missingKeyDimensions.slice(0, 3).map(dim => (
<li key={dim}>{formatDimensionName(dim)}</li>
))}
</ul>
</div>
);
}
return <ClassificationResult result={classifier.predict(ratings)} />;
}
```
## UI Components
### Classification Indicator
Visual indicator showing cluster and confidence:
```javascript
function ClassificationIndicator({ cluster, confidence, completeness }) {
const color = cluster === 1 ? '#2E86AB' : '#A23B72';
return (
<div className="classification-indicator" style={{ borderColor: color }}>
<div className="cluster-badge" style={{ backgroundColor: color }}>
{cluster === 1 ? 'Relational/Cultural' : 'Institutional/Bureaucratic'}
</div>
<div className="confidence-bar">
<div
className="confidence-fill"
style={{
width: `${confidence}%`,
backgroundColor: color,
opacity: 0.3 + (confidence / 100) * 0.7,
}}
/>
<span className="confidence-text">{confidence}% confidence</span>
</div>
<div className="completeness">
{completeness}% of dimensions provided
</div>
</div>
);
}
```
### Spectrum Visualization
Show protocol position on the relational ↔ institutional spectrum:
```javascript
function SpectrumVisualization({ ldaScore, distanceToBoundary }) {
// Scale LDA score to 0-100 for display
// Typical range is -4 to +4
const position = ((ldaScore + 4) / 8) * 100;
const boundaryZone = distanceToBoundary < 0.5;
return (
<div className="spectrum">
<div className="spectrum-bar">
<div className="spectrum-label left">Relational/Cultural</div>
<div className="spectrum-label right">Institutional/Bureaucratic</div>
<div className="spectrum-track">
{boundaryZone && (
<div className="boundary-zone" style={{ left: '45%', width: '10%' }}>
Boundary
</div>
)}
<div
className="protocol-marker"
style={{ left: `${position}%` }}
title={`LDA Score: ${ldaScore.toFixed(2)}`}
/>
</div>
</div>
</div>
);
}
```
## Form Selection Logic
### When to Use Short Form
- Initial protocol scan
- User wants quick classification
- Protocol clearly fits one family (confidence > 60%, distance > 0.5)
### When to Use Long Form
- Protocol near boundary (distance < 0.5)
- Low confidence (< 60%)
- User wants detailed analysis
- Research/documentation purposes
### Recommended Flow
```
User starts diagnostic
↓
Show short form (8 key dimensions)
↓
Calculate partial classification
↓
Is confidence > 60% AND completeness > 75%?
↓ YES ↓ NO
Show result Offer long form
"For better accuracy,
complete full diagnostic?"
```
## API Reference
### `predict(ratings, options)`
Main classification function.
**Parameters:**
- `ratings`: Object mapping dimension names to values (1-9)
- `options.detailed`: Return detailed information (default: true)
**Returns:**
```javascript
{
cluster: 1 | 2,
clusterName: "Relational/Cultural" | "Institutional/Bureaucratic",
confidence: 0-100,
completeness: 0-100,
recommendedForm: "short" | "long",
// If detailed: true
ldaScore: number,
distanceToBoundary: number,
dimensionsProvided: number,
dimensionsTotal: 23,
keyDimensionsProvided: number,
keyDimensionsTotal: 8
}
```
### `explainClassification(ratings)`
Generate human-readable explanation.
**Returns:** String with explanation text
### `getKeyDimensions()`
Get the 8 most discriminative dimensions for short form.
**Returns:** Array of dimension names
### `assessShortFormReadiness(ratings)`
Check if enough key dimensions are provided.
**Returns:**
```javascript
{
ready: boolean,
keyDimensionsProvided: number,
keyDimensionsTotal: 8,
coverage: 0-100,
missingKeyDimensions: string[]
}
```
## Testing
Test the classifier with example protocols (run from within `bicorder-app`):
```javascript
import { BicorderClassifier } from './bicorder-classifier';
import modelData from '../../analysis/bicorder_model.json';
const classifier = new BicorderClassifier(modelData);
// Test 1: Clearly institutional
const institutional = {
'Design_elite_vs_vernacular': 1,
'Entanglement_obligatory_vs_voluntary': 1,
'Entanglement_flocking_vs_swarming': 1,
};
console.log(classifier.predict(institutional));
// Expected: Cluster 2, high confidence
// Test 2: Clearly relational
const relational = {
'Design_elite_vs_vernacular': 9,
'Entanglement_obligatory_vs_voluntary': 9,
'Entanglement_flocking_vs_swarming': 9,
};
console.log(classifier.predict(relational));
// Expected: Cluster 1, high confidence
// Test 3: Boundary case
const boundary = {
'Design_elite_vs_vernacular': 5,
'Entanglement_obligatory_vs_voluntary': 5,
};
console.log(classifier.predict(boundary));
// Expected: Recommend long form
```
## Performance
- Model size: ~5KB (negligible)
- Classification time: < 1ms
- No network calls needed (runs entirely client-side)
- Works offline once model is loaded
## Next Steps
1. Integrate classifier into existing bicorder form
2. Design UI components for classification display
3. Add user preference for form selection
4. Consider adding classification to protocol browsing/search
5. Export classification data with completed diagnostics
## Questions?
See `bicorder-app/src/bicorder-classifier.ts` for the live implementation, and `bicorder-app/src/App.svelte` for how it's wired into the form.
The classifier predicts which of two protocol families a reading belongs to:
- **Cluster 1: Relational/Cultural** — community-based, emergent, voluntary protocols
- **Cluster 2: Institutional/Bureaucratic** — formal, top-down, externally enforced protocols
See `analysis/README.md` for the full multivariate analysis these clusters came from.
## Historical integration patterns
The removed web-app integration supported progressive classification display,
smart form selection (suggesting the long form when classification confidence
was low), short-form optimization around the most discriminative dimensions,
and readiness checks. The Python classifier API remains:
- `predict(ratings, options)` → cluster, clusterName, confidence, completeness, recommendedForm (detailed mode adds ldaScore, distanceToBoundary, dimension counts)
- `explain_classification(ratings)` → human-readable explanation
- `get_key_dimensions()` → the shortform/key dimensions from bicorder.json
- `assess_short_form_readiness(ratings)` (TS only, removed) — the Python `recommended_form` field remains
The shortform gradients themselves are defined in `bicorder.json`
(`shortform: true`), derived from the original feature-importance analysis —
that part of the research lives on in the tool.
## Why it was removed
- The LDA sign convention was inverted in `ascii_bicorder.py` (never caught
there because a term-rename also silently disabled the calculation), while
the web app had been separately fixed — two divergent implementations.
- Compressing a two-family classification into a 1–9 gradient was semantically
awkward and produced recurring bugs (see commit `fd556d9`).
- The version-mismatch handling differed between implementations (Python
skipped; TypeScript continued with a stale model).
- The two-families finding is a research result, not a diagnostic — it belongs
in analysis, not in the instrument itself.
The form-recommendation feature (suggesting long form when classification
confidence was low) was also removed. Shortform/longform selection is now
entirely the analyst's choice.
+7 -2
View File
@@ -416,9 +416,14 @@ Hypothesis: Changing the analyst and their standpoint could result in interestin
Method: Alongside the dataset of protocols, generate diverse personas, such as a) personas used to evaluate every protocols, and b) protocol-specific personas that reflect different relationships to the protocol. Modify the test suite to include personas as an additional dimension of the analysis.
## Integration with Bicorder Tool
## Integration with Bicorder Tool (historical)
The cluster analysis findings have been integrated into the bicorder system as an automated analysis gradient:
> **Update (v1.3.0):** The bureaucratic↔relational (formal/informal) LDA analysis
> has been **removed from the bicorder itself**. The cluster classification lives
> on as research in this directory only — see `INTEGRATION_GUIDE.md` for how to
> reproduce it and why it was removed from the tool.
The cluster analysis findings were previously integrated into the bicorder system as an automated analysis gradient:
**Bureaucratic ↔ Relational** - A new analysis field that automatically calculates where a protocol falls on the spectrum between the two protocol families identified through clustering analysis.