docs: mark classifier integration historical; fix shortform count

- INTEGRATION_GUIDE.md: rewritten as research notes — how to reproduce the
  cluster classification with the analysis scripts, and why it was removed
  from the tool in v1.3.0
- analysis/README.md: integration section marked historical
- bicorder-app/README.md: shortform is 9 gradients, not 10
This commit is contained in:
Protocolbot committed 2026-09-23 07:58:17 -06:00
1 parent 8c40ca076b
commit e7d6465ceb
3 files changed
+74 -374

No files matched your search

+66 -371
View File
@@ -1,386 +1,81 @@
# Bicorder Classifier Integration Guide # Bicorder Classifier — Research Notes
> **Status: removed from the tool (v1.3.0).** The formal/informal (bureaucratic↔relational)
> LDA analysis was removed from the bicorder itself in v1.3.0. The cluster
> classification survives as **research only** — the scripts in this directory
> can still train and apply the model to datasets, but the web app and
> `ascii_bicorder.py` no longer consume it. This document is retained as a
> historical record of how the integration worked and how to reproduce the
> research analysis.
## Overview ## Overview
This guide explains how to integrate the cluster classification system into the Bicorder web application to provide: The analysis directory contains a cluster classification system that was
previously integrated into the Bicorder web application to provide:
1. **Real-time cluster prediction** as users fill out diagnostics 1. **Real-time cluster prediction** as users filled out diagnostics
2. **Smart form selection** (short vs. long form based on classification confidence) 2. **Smart form selection** (short vs. long form based on classification confidence)
3. **Visual feedback** showing protocol family positioning 3. **Visual feedback** showing protocol family positioning
## Design Philosophy ## Original Design Philosophy
**Version-based compatibility**: The model includes a `bicorder_version` field. The classifier checks that versions match. When bicorder.json structure changes: **Version-based compatibility**: The model included a `bicorder_version` field.
1. Increment the version number in bicorder.json The classifier checked that versions matched. When bicorder.json structure changed:
2. Retrain the model with `python3 scripts/export_model_for_js.py data/synthetic_20251116/readings.csv` 1. The version number in bicorder.json was incremented
3. The new model will have the updated version 2. The model was retrained with `python3 scripts/export_model_for_js.py data/synthetic_20251116/readings.csv`
3. The new model had the updated version
This ensures the web app and model stay in sync without complex backward compatibility. ## Files (research-only now)
## Files - `bicorder_model.json` - Trained model parameters (~5KB), trained on the synthetic dataset (bicorder v1.2.6 structure — **stale** relative to v1.3.0; retrain before applying to new readings)
- `scripts/bicorder_classifier.py` - Python classifier (used by `classify_readings.py`)
- `scripts/export_model_for_js.py` - Retrain and export the model to JSON
- `scripts/classify_readings.py` - Apply the classifier to a readings CSV
- `bicorder_model.json` - Trained model parameters (~5KB); read by `bicorder-app` at build time from `../analysis/bicorder_model.json` ## Reproducing the research analysis
- `bicorder-app/src/bicorder-classifier.ts` - TypeScript classifier implementation (lives in the app, not here)
The model is the only artifact produced by this analysis directory that the app consumes. Regenerate it after re-running analysis on the synthetic dataset:
```bash ```bash
python3 scripts/export_model_for_js.py data/synthetic_20251116/readings.csv # Retrain the model on a (new) synthetic dataset
python3 scripts/export_model_for_js.py data/<dataset>/readings.csv
# Classify a dataset's readings
python3 scripts/classify_readings.py data/<dataset>/readings.csv --training data/<dataset>/readings.csv
``` ```
## Quick Start The classifier predicts which of two protocol families a reading belongs to:
- **Cluster 1: Relational/Cultural** — community-based, emergent, voluntary protocols
### Basic Usage - **Cluster 2: Institutional/Bureaucratic** — formal, top-down, externally enforced protocols
```javascript See `analysis/README.md` for the full multivariate analysis these clusters came from.
import { loadClassifier } from './lib/bicorder-classifier.js';
## Historical integration patterns
// Load model once at app startup
const classifier = await loadClassifier('/bicorder_model.json'); The removed web-app integration supported progressive classification display,
smart form selection (suggesting the long form when classification confidence
// As user fills in diagnostic form was low), short-form optimization around the most discriminative dimensions,
function onDimensionChange(dimensionName, value) { and readiness checks. The Python classifier API remains:
const currentRatings = getCurrentFormValues(); // Your form state
- `predict(ratings, options)` → cluster, clusterName, confidence, completeness, recommendedForm (detailed mode adds ldaScore, distanceToBoundary, dimension counts)
const result = classifier.predict(currentRatings); - `explain_classification(ratings)` → human-readable explanation
- `get_key_dimensions()` → the shortform/key dimensions from bicorder.json
console.log(`Cluster: ${result.clusterName}`); - `assess_short_form_readiness(ratings)` (TS only, removed) — the Python `recommended_form` field remains
console.log(`Confidence: ${result.confidence}%`);
console.log(`Recommend: ${result.recommendedForm} form`); The shortform gradients themselves are defined in `bicorder.json`
(`shortform: true`), derived from the original feature-importance analysis —
updateUI(result); that part of the research lives on in the tool.
}
``` ## Why it was removed
## Integration Patterns - The LDA sign convention was inverted in `ascii_bicorder.py` (never caught
there because a term-rename also silently disabled the calculation), while
### Pattern 1: Progressive Classification Display the web app had been separately fixed — two divergent implementations.
- Compressing a two-family classification into a 1–9 gradient was semantically
Show classification results as the user fills out the form: awkward and produced recurring bugs (see commit `fd556d9`).
- The version-mismatch handling differed between implementations (Python
```javascript skipped; TypeScript continued with a stale model).
// React/Svelte component example - The two-families finding is a research result, not a diagnostic — it belongs
function DiagnosticForm() { in analysis, not in the instrument itself.
const [ratings, setRatings] = useState({});
const [classification, setClassification] = useState(null); The form-recommendation feature (suggesting long form when classification
confidence was low) was also removed. Shortform/longform selection is now
useEffect(() => { entirely the analyst's choice.
if (Object.keys(ratings).length > 0) {
const result = classifier.predict(ratings);
setClassification(result);
}
}, [ratings]);
return (
<div>
<DiagnosticQuestions onChange={setRatings} />
{classification && (
<ClassificationIndicator
cluster={classification.clusterName}
confidence={classification.confidence}
completeness={classification.completeness}
/>
)}
</div>
);
}
```
### Pattern 2: Smart Form Selection
Automatically switch between short and long forms:
```javascript
function DiagnosticWizard() {
const [ratings, setRatings] = useState({});
function handleDimensionComplete(dimension, value) {
const newRatings = { ...ratings, [dimension]: value };
setRatings(newRatings);
// Check if we should switch forms
const result = classifier.predict(newRatings);
if (result.recommendedForm === 'long' && currentForm === 'short') {
showFormSwitchPrompt(
'Your protocol shows characteristics of both families. ' +
'Would you like to use the detailed form for better classification?'
);
}
}
return <Form onDimensionComplete={handleDimensionComplete} />;
}
```
### Pattern 3: Short Form Optimization
Only ask the 8 most discriminative dimensions for quick classification:
```javascript
const shortFormDimensions = classifier.getKeyDimensions();
// Returns:
// [
// 'Design_elite_vs_vernacular',
// 'Entanglement_flocking_vs_swarming',
// 'Design_static_vs_malleable',
// 'Entanglement_obligatory_vs_voluntary',
// 'Entanglement_self-enforcing_vs_enforced',
// 'Design_explicit_vs_implicit',
// 'Entanglement_sovereign_vs_subsidiary',
// 'Design_technical_vs_social',
// ]
function ShortForm() {
return (
<div>
<h2>Quick Classification (8 questions)</h2>
{shortFormDimensions.map(dim => (
<DimensionSlider key={dim} dimension={dim} />
))}
</div>
);
}
```
### Pattern 4: Readiness Check
Check if user has provided enough data for reliable classification:
```javascript
function ClassificationStatus() {
const assessment = classifier.assessShortFormReadiness(ratings);
if (!assessment.ready) {
return (
<div className="status-warning">
<p>
Need {assessment.keyDimensionsTotal - assessment.keyDimensionsProvided} more
key dimensions for reliable classification ({assessment.coverage}% complete)
</p>
<ul>
{assessment.missingKeyDimensions.slice(0, 3).map(dim => (
<li key={dim}>{formatDimensionName(dim)}</li>
))}
</ul>
</div>
);
}
return <ClassificationResult result={classifier.predict(ratings)} />;
}
```
## UI Components
### Classification Indicator
Visual indicator showing cluster and confidence:
```javascript
function ClassificationIndicator({ cluster, confidence, completeness }) {
const color = cluster === 1 ? '#2E86AB' : '#A23B72';
return (
<div className="classification-indicator" style={{ borderColor: color }}>
<div className="cluster-badge" style={{ backgroundColor: color }}>
{cluster === 1 ? 'Relational/Cultural' : 'Institutional/Bureaucratic'}
</div>
<div className="confidence-bar">
<div
className="confidence-fill"
style={{
width: `${confidence}%`,
backgroundColor: color,
opacity: 0.3 + (confidence / 100) * 0.7,
}}
/>
<span className="confidence-text">{confidence}% confidence</span>
</div>
<div className="completeness">
{completeness}% of dimensions provided
</div>
</div>
);
}
```
### Spectrum Visualization
Show protocol position on the relational ↔ institutional spectrum:
```javascript
function SpectrumVisualization({ ldaScore, distanceToBoundary }) {
// Scale LDA score to 0-100 for display
// Typical range is -4 to +4
const position = ((ldaScore + 4) / 8) * 100;
const boundaryZone = distanceToBoundary < 0.5;
return (
<div className="spectrum">
<div className="spectrum-bar">
<div className="spectrum-label left">Relational/Cultural</div>
<div className="spectrum-label right">Institutional/Bureaucratic</div>
<div className="spectrum-track">
{boundaryZone && (
<div className="boundary-zone" style={{ left: '45%', width: '10%' }}>
Boundary
</div>
)}
<div
className="protocol-marker"
style={{ left: `${position}%` }}
title={`LDA Score: ${ldaScore.toFixed(2)}`}
/>
</div>
</div>
</div>
);
}
```
## Form Selection Logic
### When to Use Short Form
- Initial protocol scan
- User wants quick classification
- Protocol clearly fits one family (confidence > 60%, distance > 0.5)
### When to Use Long Form
- Protocol near boundary (distance < 0.5)
- Low confidence (< 60%)
- User wants detailed analysis
- Research/documentation purposes
### Recommended Flow
```
User starts diagnostic
↓
Show short form (8 key dimensions)
↓
Calculate partial classification
↓
Is confidence > 60% AND completeness > 75%?
↓ YES ↓ NO
Show result Offer long form
"For better accuracy,
complete full diagnostic?"
```
## API Reference
### `predict(ratings, options)`
Main classification function.
**Parameters:**
- `ratings`: Object mapping dimension names to values (1-9)
- `options.detailed`: Return detailed information (default: true)
**Returns:**
```javascript
{
cluster: 1 | 2,
clusterName: "Relational/Cultural" | "Institutional/Bureaucratic",
confidence: 0-100,
completeness: 0-100,
recommendedForm: "short" | "long",
// If detailed: true
ldaScore: number,
distanceToBoundary: number,
dimensionsProvided: number,
dimensionsTotal: 23,
keyDimensionsProvided: number,
keyDimensionsTotal: 8
}
```
### `explainClassification(ratings)`
Generate human-readable explanation.
**Returns:** String with explanation text
### `getKeyDimensions()`
Get the 8 most discriminative dimensions for short form.
**Returns:** Array of dimension names
### `assessShortFormReadiness(ratings)`
Check if enough key dimensions are provided.
**Returns:**
```javascript
{
ready: boolean,
keyDimensionsProvided: number,
keyDimensionsTotal: 8,
coverage: 0-100,
missingKeyDimensions: string[]
}
```
## Testing
Test the classifier with example protocols (run from within `bicorder-app`):
```javascript
import { BicorderClassifier } from './bicorder-classifier';
import modelData from '../../analysis/bicorder_model.json';
const classifier = new BicorderClassifier(modelData);
// Test 1: Clearly institutional
const institutional = {
'Design_elite_vs_vernacular': 1,
'Entanglement_obligatory_vs_voluntary': 1,
'Entanglement_flocking_vs_swarming': 1,
};
console.log(classifier.predict(institutional));
// Expected: Cluster 2, high confidence
// Test 2: Clearly relational
const relational = {
'Design_elite_vs_vernacular': 9,
'Entanglement_obligatory_vs_voluntary': 9,
'Entanglement_flocking_vs_swarming': 9,
};
console.log(classifier.predict(relational));
// Expected: Cluster 1, high confidence
// Test 3: Boundary case
const boundary = {
'Design_elite_vs_vernacular': 5,
'Entanglement_obligatory_vs_voluntary': 5,
};
console.log(classifier.predict(boundary));
// Expected: Recommend long form
```
## Performance
- Model size: ~5KB (negligible)
- Classification time: < 1ms
- No network calls needed (runs entirely client-side)
- Works offline once model is loaded
## Next Steps
1. Integrate classifier into existing bicorder form
2. Design UI components for classification display
3. Add user preference for form selection
4. Consider adding classification to protocol browsing/search
5. Export classification data with completed diagnostics
## Questions?
See `bicorder-app/src/bicorder-classifier.ts` for the live implementation, and `bicorder-app/src/App.svelte` for how it's wired into the form.
+7 -2
View File
@@ -416,9 +416,14 @@ Hypothesis: Changing the analyst and their standpoint could result in interestin
Method: Alongside the dataset of protocols, generate diverse personas, such as a) personas used to evaluate every protocols, and b) protocol-specific personas that reflect different relationships to the protocol. Modify the test suite to include personas as an additional dimension of the analysis. Method: Alongside the dataset of protocols, generate diverse personas, such as a) personas used to evaluate every protocols, and b) protocol-specific personas that reflect different relationships to the protocol. Modify the test suite to include personas as an additional dimension of the analysis.
## Integration with Bicorder Tool ## Integration with Bicorder Tool (historical)
The cluster analysis findings have been integrated into the bicorder system as an automated analysis gradient: > **Update (v1.3.0):** The bureaucratic↔relational (formal/informal) LDA analysis
> has been **removed from the bicorder itself**. The cluster classification lives
> on as research in this directory only — see `INTEGRATION_GUIDE.md` for how to
> reproduce it and why it was removed from the tool.
The cluster analysis findings were previously integrated into the bicorder system as an automated analysis gradient:
**Bureaucratic ↔ Relational** - A new analysis field that automatically calculates where a protocol falls on the spectrum between the two protocol families identified through clustering analysis. **Bureaucratic ↔ Relational** - A new analysis field that automatically calculates where a protocol falls on the spectrum between the two protocol families identified through clustering analysis.
+1 -1
View File
@@ -6,7 +6,7 @@ A Svelte Progressive Web App (PWA) for carrying out protocol diagnostics as defi
- **Single-page diagnostic tool** with ASCII-styled interface - **Single-page diagnostic tool** with ASCII-styled interface
- **Touch-friendly controls** optimized for mobile devices - **Touch-friendly controls** optimized for mobile devices
- **Shortform toggle** - switch between full (23 gradients) and short (10 gradients) versions - **Shortform toggle** - switch between full (23 gradients) and short (9 gradients) versions
- **Tooltips** on all gradient terms (long-press on mobile, hover on desktop) - **Tooltips** on all gradient terms (long-press on mobile, hover on desktop)
- **Editable metadata** fields with auto-generated timestamps - **Editable metadata** fields with auto-generated timestamps
- **Auto-calculated analysis** section (hardness/softness, polarized/centrist) - **Auto-calculated analysis** section (hardness/softness, polarized/centrist)