docs: mark classifier integration historical; fix shortform count
- INTEGRATION_GUIDE.md: rewritten as research notes — how to reproduce the cluster classification with the analysis scripts, and why it was removed from the tool in v1.3.0 - analysis/README.md: integration section marked historical - bicorder-app/README.md: shortform is 9 gradients, not 10
This commit is contained in:
1 parent
8c40ca076b
commit
e7d6465ceb
3 files changed
+74
-374
No files matched your search
+66
-371
@@ -1,386 +1,81 @@
|
||||
# Bicorder Classifier Integration Guide
|
||||
# Bicorder Classifier — Research Notes
|
||||
|
||||
> **Status: removed from the tool (v1.3.0).** The formal/informal (bureaucratic↔relational)
|
||||
> LDA analysis was removed from the bicorder itself in v1.3.0. The cluster
|
||||
> classification survives as **research only** — the scripts in this directory
|
||||
> can still train and apply the model to datasets, but the web app and
|
||||
> `ascii_bicorder.py` no longer consume it. This document is retained as a
|
||||
> historical record of how the integration worked and how to reproduce the
|
||||
> research analysis.
|
||||
|
||||
## Overview
|
||||
|
||||
This guide explains how to integrate the cluster classification system into the Bicorder web application to provide:
|
||||
The analysis directory contains a cluster classification system that was
|
||||
previously integrated into the Bicorder web application to provide:
|
||||
|
||||
1. **Real-time cluster prediction** as users fill out diagnostics
|
||||
1. **Real-time cluster prediction** as users filled out diagnostics
|
||||
2. **Smart form selection** (short vs. long form based on classification confidence)
|
||||
3. **Visual feedback** showing protocol family positioning
|
||||
|
||||
## Design Philosophy
|
||||
## Original Design Philosophy
|
||||
|
||||
**Version-based compatibility**: The model includes a `bicorder_version` field. The classifier checks that versions match. When bicorder.json structure changes:
|
||||
1. Increment the version number in bicorder.json
|
||||
2. Retrain the model with `python3 scripts/export_model_for_js.py data/synthetic_20251116/readings.csv`
|
||||
3. The new model will have the updated version
|
||||
**Version-based compatibility**: The model included a `bicorder_version` field.
|
||||
The classifier checked that versions matched. When bicorder.json structure changed:
|
||||
1. The version number in bicorder.json was incremented
|
||||
2. The model was retrained with `python3 scripts/export_model_for_js.py data/synthetic_20251116/readings.csv`
|
||||
3. The new model had the updated version
|
||||
|
||||
This ensures the web app and model stay in sync without complex backward compatibility.
|
||||
## Files (research-only now)
|
||||
|
||||
## Files
|
||||
- `bicorder_model.json` - Trained model parameters (~5KB), trained on the synthetic dataset (bicorder v1.2.6 structure — **stale** relative to v1.3.0; retrain before applying to new readings)
|
||||
- `scripts/bicorder_classifier.py` - Python classifier (used by `classify_readings.py`)
|
||||
- `scripts/export_model_for_js.py` - Retrain and export the model to JSON
|
||||
- `scripts/classify_readings.py` - Apply the classifier to a readings CSV
|
||||
|
||||
- `bicorder_model.json` - Trained model parameters (~5KB); read by `bicorder-app` at build time from `../analysis/bicorder_model.json`
|
||||
- `bicorder-app/src/bicorder-classifier.ts` - TypeScript classifier implementation (lives in the app, not here)
|
||||
|
||||
The model is the only artifact produced by this analysis directory that the app consumes. Regenerate it after re-running analysis on the synthetic dataset:
|
||||
## Reproducing the research analysis
|
||||
|
||||
```bash
|
||||
python3 scripts/export_model_for_js.py data/synthetic_20251116/readings.csv
|
||||
# Retrain the model on a (new) synthetic dataset
|
||||
python3 scripts/export_model_for_js.py data/<dataset>/readings.csv
|
||||
|
||||
# Classify a dataset's readings
|
||||
python3 scripts/classify_readings.py data/<dataset>/readings.csv --training data/<dataset>/readings.csv
|
||||
```
|
||||
|
||||
## Quick Start
|
||||
|
||||
### Basic Usage
|
||||
|
||||
```javascript
|
||||
import { loadClassifier } from './lib/bicorder-classifier.js';
|
||||
|
||||
// Load model once at app startup
|
||||
const classifier = await loadClassifier('/bicorder_model.json');
|
||||
|
||||
// As user fills in diagnostic form
|
||||
function onDimensionChange(dimensionName, value) {
|
||||
const currentRatings = getCurrentFormValues(); // Your form state
|
||||
|
||||
const result = classifier.predict(currentRatings);
|
||||
|
||||
console.log(`Cluster: ${result.clusterName}`);
|
||||
console.log(`Confidence: ${result.confidence}%`);
|
||||
console.log(`Recommend: ${result.recommendedForm} form`);
|
||||
|
||||
updateUI(result);
|
||||
}
|
||||
```
|
||||
|
||||
## Integration Patterns
|
||||
|
||||
### Pattern 1: Progressive Classification Display
|
||||
|
||||
Show classification results as the user fills out the form:
|
||||
|
||||
```javascript
|
||||
// React/Svelte component example
|
||||
function DiagnosticForm() {
|
||||
const [ratings, setRatings] = useState({});
|
||||
const [classification, setClassification] = useState(null);
|
||||
|
||||
useEffect(() => {
|
||||
if (Object.keys(ratings).length > 0) {
|
||||
const result = classifier.predict(ratings);
|
||||
setClassification(result);
|
||||
}
|
||||
}, [ratings]);
|
||||
|
||||
return (
|
||||
<div>
|
||||
<DiagnosticQuestions onChange={setRatings} />
|
||||
|
||||
{classification && (
|
||||
<ClassificationIndicator
|
||||
cluster={classification.clusterName}
|
||||
confidence={classification.confidence}
|
||||
completeness={classification.completeness}
|
||||
/>
|
||||
)}
|
||||
</div>
|
||||
);
|
||||
}
|
||||
```
|
||||
|
||||
### Pattern 2: Smart Form Selection
|
||||
|
||||
Automatically switch between short and long forms:
|
||||
|
||||
```javascript
|
||||
function DiagnosticWizard() {
|
||||
const [ratings, setRatings] = useState({});
|
||||
|
||||
function handleDimensionComplete(dimension, value) {
|
||||
const newRatings = { ...ratings, [dimension]: value };
|
||||
setRatings(newRatings);
|
||||
|
||||
// Check if we should switch forms
|
||||
const result = classifier.predict(newRatings);
|
||||
|
||||
if (result.recommendedForm === 'long' && currentForm === 'short') {
|
||||
showFormSwitchPrompt(
|
||||
'Your protocol shows characteristics of both families. ' +
|
||||
'Would you like to use the detailed form for better classification?'
|
||||
);
|
||||
}
|
||||
}
|
||||
|
||||
return <Form onDimensionComplete={handleDimensionComplete} />;
|
||||
}
|
||||
```
|
||||
|
||||
### Pattern 3: Short Form Optimization
|
||||
|
||||
Only ask the 8 most discriminative dimensions for quick classification:
|
||||
|
||||
```javascript
|
||||
const shortFormDimensions = classifier.getKeyDimensions();
|
||||
// Returns:
|
||||
// [
|
||||
// 'Design_elite_vs_vernacular',
|
||||
// 'Entanglement_flocking_vs_swarming',
|
||||
// 'Design_static_vs_malleable',
|
||||
// 'Entanglement_obligatory_vs_voluntary',
|
||||
// 'Entanglement_self-enforcing_vs_enforced',
|
||||
// 'Design_explicit_vs_implicit',
|
||||
// 'Entanglement_sovereign_vs_subsidiary',
|
||||
// 'Design_technical_vs_social',
|
||||
// ]
|
||||
|
||||
function ShortForm() {
|
||||
return (
|
||||
<div>
|
||||
<h2>Quick Classification (8 questions)</h2>
|
||||
{shortFormDimensions.map(dim => (
|
||||
<DimensionSlider key={dim} dimension={dim} />
|
||||
))}
|
||||
</div>
|
||||
);
|
||||
}
|
||||
```
|
||||
|
||||
### Pattern 4: Readiness Check
|
||||
|
||||
Check if user has provided enough data for reliable classification:
|
||||
|
||||
```javascript
|
||||
function ClassificationStatus() {
|
||||
const assessment = classifier.assessShortFormReadiness(ratings);
|
||||
|
||||
if (!assessment.ready) {
|
||||
return (
|
||||
<div className="status-warning">
|
||||
<p>
|
||||
Need {assessment.keyDimensionsTotal - assessment.keyDimensionsProvided} more
|
||||
key dimensions for reliable classification ({assessment.coverage}% complete)
|
||||
</p>
|
||||
<ul>
|
||||
{assessment.missingKeyDimensions.slice(0, 3).map(dim => (
|
||||
<li key={dim}>{formatDimensionName(dim)}</li>
|
||||
))}
|
||||
</ul>
|
||||
</div>
|
||||
);
|
||||
}
|
||||
|
||||
return <ClassificationResult result={classifier.predict(ratings)} />;
|
||||
}
|
||||
```
|
||||
|
||||
## UI Components
|
||||
|
||||
### Classification Indicator
|
||||
|
||||
Visual indicator showing cluster and confidence:
|
||||
|
||||
```javascript
|
||||
function ClassificationIndicator({ cluster, confidence, completeness }) {
|
||||
const color = cluster === 1 ? '#2E86AB' : '#A23B72';
|
||||
|
||||
return (
|
||||
<div className="classification-indicator" style={{ borderColor: color }}>
|
||||
<div className="cluster-badge" style={{ backgroundColor: color }}>
|
||||
{cluster === 1 ? 'Relational/Cultural' : 'Institutional/Bureaucratic'}
|
||||
</div>
|
||||
|
||||
<div className="confidence-bar">
|
||||
<div
|
||||
className="confidence-fill"
|
||||
style={{
|
||||
width: `${confidence}%`,
|
||||
backgroundColor: color,
|
||||
opacity: 0.3 + (confidence / 100) * 0.7,
|
||||
}}
|
||||
/>
|
||||
<span className="confidence-text">{confidence}% confidence</span>
|
||||
</div>
|
||||
|
||||
<div className="completeness">
|
||||
{completeness}% of dimensions provided
|
||||
</div>
|
||||
</div>
|
||||
);
|
||||
}
|
||||
```
|
||||
|
||||
### Spectrum Visualization
|
||||
|
||||
Show protocol position on the relational ↔ institutional spectrum:
|
||||
|
||||
```javascript
|
||||
function SpectrumVisualization({ ldaScore, distanceToBoundary }) {
|
||||
// Scale LDA score to 0-100 for display
|
||||
// Typical range is -4 to +4
|
||||
const position = ((ldaScore + 4) / 8) * 100;
|
||||
const boundaryZone = distanceToBoundary < 0.5;
|
||||
|
||||
return (
|
||||
<div className="spectrum">
|
||||
<div className="spectrum-bar">
|
||||
<div className="spectrum-label left">Relational/Cultural</div>
|
||||
<div className="spectrum-label right">Institutional/Bureaucratic</div>
|
||||
|
||||
<div className="spectrum-track">
|
||||
{boundaryZone && (
|
||||
<div className="boundary-zone" style={{ left: '45%', width: '10%' }}>
|
||||
Boundary
|
||||
</div>
|
||||
)}
|
||||
<div
|
||||
className="protocol-marker"
|
||||
style={{ left: `${position}%` }}
|
||||
title={`LDA Score: ${ldaScore.toFixed(2)}`}
|
||||
/>
|
||||
</div>
|
||||
</div>
|
||||
</div>
|
||||
);
|
||||
}
|
||||
```
|
||||
|
||||
## Form Selection Logic
|
||||
|
||||
### When to Use Short Form
|
||||
|
||||
- Initial protocol scan
|
||||
- User wants quick classification
|
||||
- Protocol clearly fits one family (confidence > 60%, distance > 0.5)
|
||||
|
||||
### When to Use Long Form
|
||||
|
||||
- Protocol near boundary (distance < 0.5)
|
||||
- Low confidence (< 60%)
|
||||
- User wants detailed analysis
|
||||
- Research/documentation purposes
|
||||
|
||||
### Recommended Flow
|
||||
|
||||
```
|
||||
User starts diagnostic
|
||||
↓
|
||||
Show short form (8 key dimensions)
|
||||
↓
|
||||
Calculate partial classification
|
||||
↓
|
||||
Is confidence > 60% AND completeness > 75%?
|
||||
↓ YES ↓ NO
|
||||
Show result Offer long form
|
||||
"For better accuracy,
|
||||
complete full diagnostic?"
|
||||
```
|
||||
|
||||
## API Reference
|
||||
|
||||
### `predict(ratings, options)`
|
||||
|
||||
Main classification function.
|
||||
|
||||
**Parameters:**
|
||||
- `ratings`: Object mapping dimension names to values (1-9)
|
||||
- `options.detailed`: Return detailed information (default: true)
|
||||
|
||||
**Returns:**
|
||||
```javascript
|
||||
{
|
||||
cluster: 1 | 2,
|
||||
clusterName: "Relational/Cultural" | "Institutional/Bureaucratic",
|
||||
confidence: 0-100,
|
||||
completeness: 0-100,
|
||||
recommendedForm: "short" | "long",
|
||||
// If detailed: true
|
||||
ldaScore: number,
|
||||
distanceToBoundary: number,
|
||||
dimensionsProvided: number,
|
||||
dimensionsTotal: 23,
|
||||
keyDimensionsProvided: number,
|
||||
keyDimensionsTotal: 8
|
||||
}
|
||||
```
|
||||
|
||||
### `explainClassification(ratings)`
|
||||
|
||||
Generate human-readable explanation.
|
||||
|
||||
**Returns:** String with explanation text
|
||||
|
||||
### `getKeyDimensions()`
|
||||
|
||||
Get the 8 most discriminative dimensions for short form.
|
||||
|
||||
**Returns:** Array of dimension names
|
||||
|
||||
### `assessShortFormReadiness(ratings)`
|
||||
|
||||
Check if enough key dimensions are provided.
|
||||
|
||||
**Returns:**
|
||||
```javascript
|
||||
{
|
||||
ready: boolean,
|
||||
keyDimensionsProvided: number,
|
||||
keyDimensionsTotal: 8,
|
||||
coverage: 0-100,
|
||||
missingKeyDimensions: string[]
|
||||
}
|
||||
```
|
||||
|
||||
## Testing
|
||||
|
||||
Test the classifier with example protocols (run from within `bicorder-app`):
|
||||
|
||||
```javascript
|
||||
import { BicorderClassifier } from './bicorder-classifier';
|
||||
import modelData from '../../analysis/bicorder_model.json';
|
||||
|
||||
const classifier = new BicorderClassifier(modelData);
|
||||
|
||||
// Test 1: Clearly institutional
|
||||
const institutional = {
|
||||
'Design_elite_vs_vernacular': 1,
|
||||
'Entanglement_obligatory_vs_voluntary': 1,
|
||||
'Entanglement_flocking_vs_swarming': 1,
|
||||
};
|
||||
console.log(classifier.predict(institutional));
|
||||
// Expected: Cluster 2, high confidence
|
||||
|
||||
// Test 2: Clearly relational
|
||||
const relational = {
|
||||
'Design_elite_vs_vernacular': 9,
|
||||
'Entanglement_obligatory_vs_voluntary': 9,
|
||||
'Entanglement_flocking_vs_swarming': 9,
|
||||
};
|
||||
console.log(classifier.predict(relational));
|
||||
// Expected: Cluster 1, high confidence
|
||||
|
||||
// Test 3: Boundary case
|
||||
const boundary = {
|
||||
'Design_elite_vs_vernacular': 5,
|
||||
'Entanglement_obligatory_vs_voluntary': 5,
|
||||
};
|
||||
console.log(classifier.predict(boundary));
|
||||
// Expected: Recommend long form
|
||||
```
|
||||
|
||||
## Performance
|
||||
|
||||
- Model size: ~5KB (negligible)
|
||||
- Classification time: < 1ms
|
||||
- No network calls needed (runs entirely client-side)
|
||||
- Works offline once model is loaded
|
||||
|
||||
## Next Steps
|
||||
|
||||
1. Integrate classifier into existing bicorder form
|
||||
2. Design UI components for classification display
|
||||
3. Add user preference for form selection
|
||||
4. Consider adding classification to protocol browsing/search
|
||||
5. Export classification data with completed diagnostics
|
||||
|
||||
## Questions?
|
||||
|
||||
See `bicorder-app/src/bicorder-classifier.ts` for the live implementation, and `bicorder-app/src/App.svelte` for how it's wired into the form.
|
||||
The classifier predicts which of two protocol families a reading belongs to:
|
||||
- **Cluster 1: Relational/Cultural** — community-based, emergent, voluntary protocols
|
||||
- **Cluster 2: Institutional/Bureaucratic** — formal, top-down, externally enforced protocols
|
||||
|
||||
See `analysis/README.md` for the full multivariate analysis these clusters came from.
|
||||
|
||||
## Historical integration patterns
|
||||
|
||||
The removed web-app integration supported progressive classification display,
|
||||
smart form selection (suggesting the long form when classification confidence
|
||||
was low), short-form optimization around the most discriminative dimensions,
|
||||
and readiness checks. The Python classifier API remains:
|
||||
|
||||
- `predict(ratings, options)` → cluster, clusterName, confidence, completeness, recommendedForm (detailed mode adds ldaScore, distanceToBoundary, dimension counts)
|
||||
- `explain_classification(ratings)` → human-readable explanation
|
||||
- `get_key_dimensions()` → the shortform/key dimensions from bicorder.json
|
||||
- `assess_short_form_readiness(ratings)` (TS only, removed) — the Python `recommended_form` field remains
|
||||
|
||||
The shortform gradients themselves are defined in `bicorder.json`
|
||||
(`shortform: true`), derived from the original feature-importance analysis —
|
||||
that part of the research lives on in the tool.
|
||||
|
||||
## Why it was removed
|
||||
|
||||
- The LDA sign convention was inverted in `ascii_bicorder.py` (never caught
|
||||
there because a term-rename also silently disabled the calculation), while
|
||||
the web app had been separately fixed — two divergent implementations.
|
||||
- Compressing a two-family classification into a 1–9 gradient was semantically
|
||||
awkward and produced recurring bugs (see commit `fd556d9`).
|
||||
- The version-mismatch handling differed between implementations (Python
|
||||
skipped; TypeScript continued with a stale model).
|
||||
- The two-families finding is a research result, not a diagnostic — it belongs
|
||||
in analysis, not in the instrument itself.
|
||||
|
||||
The form-recommendation feature (suggesting long form when classification
|
||||
confidence was low) was also removed. Shortform/longform selection is now
|
||||
entirely the analyst's choice.
|
||||
+7
-2
@@ -416,9 +416,14 @@ Hypothesis: Changing the analyst and their standpoint could result in interestin
|
||||
|
||||
Method: Alongside the dataset of protocols, generate diverse personas, such as a) personas used to evaluate every protocols, and b) protocol-specific personas that reflect different relationships to the protocol. Modify the test suite to include personas as an additional dimension of the analysis.
|
||||
|
||||
## Integration with Bicorder Tool
|
||||
## Integration with Bicorder Tool (historical)
|
||||
|
||||
The cluster analysis findings have been integrated into the bicorder system as an automated analysis gradient:
|
||||
> **Update (v1.3.0):** The bureaucratic↔relational (formal/informal) LDA analysis
|
||||
> has been **removed from the bicorder itself**. The cluster classification lives
|
||||
> on as research in this directory only — see `INTEGRATION_GUIDE.md` for how to
|
||||
> reproduce it and why it was removed from the tool.
|
||||
|
||||
The cluster analysis findings were previously integrated into the bicorder system as an automated analysis gradient:
|
||||
|
||||
**Bureaucratic ↔ Relational** - A new analysis field that automatically calculates where a protocol falls on the spectrum between the two protocol families identified through clustering analysis.
|
||||
|
||||
|
||||
@@ -6,7 +6,7 @@ A Svelte Progressive Web App (PWA) for carrying out protocol diagnostics as defi
|
||||
|
||||
- **Single-page diagnostic tool** with ASCII-styled interface
|
||||
- **Touch-friendly controls** optimized for mobile devices
|
||||
- **Shortform toggle** - switch between full (23 gradients) and short (10 gradients) versions
|
||||
- **Shortform toggle** - switch between full (23 gradients) and short (9 gradients) versions
|
||||
- **Tooltips** on all gradient terms (long-press on mobile, hover on desktop)
|
||||
- **Editable metadata** fields with auto-generated timestamps
|
||||
- **Auto-calculated analysis** section (hardness/softness, polarized/centrist)
|
||||
|
||||
Reference in new issue
Block a user