Removed raw Claude outputs from README
This commit is contained in:
1 parent
161ba3b136
commit
a4f16e8e4d
1 file changed
+1
-203
+1
-203
@@ -2,7 +2,7 @@
|
|||||||
|
|
||||||
This directory concerns analyses conducted with the Protocol Bicorder across multiple datasets.
|
This directory concerns analyses conducted with the Protocol Bicorder across multiple datasets.
|
||||||
|
|
||||||
Scripts were created with the assistance of Claude Code. Data processing was done largely with either local models or the Ollama cloud service, which does not retain user data. Thanks to [Seth Frey (UC Davis)](https://enfascination.com/) for guidance, but all mistakes are the responsibility of the author, [Nathan Schneider](https://nathanschneider.info). This is the work of a researcher working with AI outside their field of expertise and should be treated as a playful experiment, not a model of rigorous methodology.
|
Scripts were created with the assistance of various AI tools. Data processing was done largely with either local models or the Ollama cloud service, which does not retain user data. Thanks to [Seth Frey (UC Davis)](https://enfascination.com/) for guidance, but all mistakes are the responsibility of the author, [Nathan Schneider](https://nathanschneider.info).
|
||||||
|
|
||||||
## Datasets
|
## Datasets
|
||||||
|
|
||||||
@@ -198,68 +198,6 @@ Initial manual observations:
|
|||||||
|
|
||||||

|

|
||||||
|
|
||||||
Claude's interpretation:
|
|
||||||
|
|
||||||
> 1. Two Fundamental Protocol Types (K-Means Clustering)
|
|
||||||
>
|
|
||||||
> The data reveals two distinct protocol families (216 vs 192 protocols):
|
|
||||||
>
|
|
||||||
> Cluster 1: "Vernacular/Emergent Protocols"
|
|
||||||
> - Examples: Marronage, Songlines, Access-Centered Practices, Ethereum Proof of Work, Sangoma Healing Practices
|
|
||||||
> - Characteristics:
|
|
||||||
> - HIGH: elite→vernacular (6.4), malleable (7.3), flocking→swarming (6.4)
|
|
||||||
> - LOW: self-enforcing→enforced (3.4), sovereign→subsidiary (2.9)
|
|
||||||
>
|
|
||||||
> Cluster 2: "Institutional/Standardized Protocols"
|
|
||||||
> - Examples: ISO standards, Greenwich Mean Time, Building Codes, German Bureaucratic Prose, Royal Access Protocol
|
|
||||||
> - Characteristics:
|
|
||||||
> - HIGH: self-enforcing (7.1), sovereign (6.0)
|
|
||||||
> - LOW: elite→vernacular (1.8), flocking→swarming (2.3), static (3.5)
|
|
||||||
>
|
|
||||||
> 2. Key Structural Dimensions (PCA)
|
|
||||||
>
|
|
||||||
> Three principal components explain 55% of variance:
|
|
||||||
>
|
|
||||||
> PC1 (main axis of variation): Elite/Static/Flocking ↔ Self-enforcing
|
|
||||||
> - Essentially captures the Vernacular vs. Institutional divide
|
|
||||||
>
|
|
||||||
> PC2: Sufficient/Crystallized ↔ Kafka-esque
|
|
||||||
> - Measures protocol "completeness" vs. bureaucratic nightmare quality
|
|
||||||
>
|
|
||||||
> PC3: Universal/Technical/Macro ↔ Particular/Embodied
|
|
||||||
> - Scale and abstraction level
|
|
||||||
>
|
|
||||||
> 3. Strong Correlations (Most Significant Relationships)
|
|
||||||
>
|
|
||||||
> 1. Static ↔ Predictable (r=0.61): Unchanging protocols create predictable experiences
|
|
||||||
> 2. Elite ↔ Self-enforcing (r=-0.58): Elite protocols need external enforcement; vernacular ones self-enforce
|
|
||||||
> 3. Self-enforcing ↔ Flocking (r=-0.56): Self-enforcing protocols resist swarming dynamics
|
|
||||||
> 4. Exclusion ↔ Kafka (r=0.52): Exclusionary protocols feel Kafka-esque
|
|
||||||
>
|
|
||||||
> 4. Most Discriminative Dimension (Feature Importance)
|
|
||||||
>
|
|
||||||
> Design_elite_vs_vernacular (22.7% importance) is by far the most powerful predictor of protocol type, followed by Entanglement_flocking_vs_swarming (13.8%).
|
|
||||||
>
|
|
||||||
> 5. Most "Central" Protocols (Network Analysis)
|
|
||||||
>
|
|
||||||
> These protocols share the most dimensional similarities with others:
|
|
||||||
> 1. VPN Usage (Circumvention Protocol) - bridges many protocol types
|
|
||||||
> 2. Access Check-in - connects accessibility and participation patterns
|
|
||||||
> 3. Quadratic Voting - spans governance dimensions
|
|
||||||
>
|
|
||||||
> 6. Outliers (DBSCAN found 281!)
|
|
||||||
>
|
|
||||||
> Most protocols are actually quite unique - DBSCAN identified 281 outliers, suggesting the dataset contains many distinctive protocol configurations that don't fit neat clusters. Only 10 tight sub-clusters exist.
|
|
||||||
>
|
|
||||||
> 7. Category Prediction Power
|
|
||||||
>
|
|
||||||
> - Design dimensions predict clustering with 90.4% accuracy
|
|
||||||
> - Entanglement dimensions: 89.2% accuracy
|
|
||||||
> - Experience dimensions: only 78.3% accuracy
|
|
||||||
>
|
|
||||||
> This suggests Design and Entanglement are more fundamental than Experience.
|
|
||||||
>
|
|
||||||
> The core insight: Protocols fundamentally divide between vernacular/emergent/malleable forms and institutional/standardized/static forms, with the elite↔vernacular dimension being the strongest predictor of all other characteristics.
|
|
||||||
|
|
||||||
Comments:
|
Comments:
|
||||||
|
|
||||||
@@ -279,146 +217,6 @@ Reflections on manual review:
|
|||||||
* Should have separate values for "n/a" (0---but that could screw up averages) and both (5)
|
* Should have separate values for "n/a" (0---but that could screw up averages) and both (5)
|
||||||
* Remove the analysis section, or use analyses here for what becomes most meaningful
|
* Remove the analysis section, or use analyses here for what becomes most meaningful
|
||||||
|
|
||||||
Claude report on possible improvements:
|
|
||||||
|
|
||||||
> 1. Prioritize High-Impact Dimensions ⭐
|
|
||||||
>
|
|
||||||
> The current 23 dimensions aren't equally informative. Reorder by importance:
|
|
||||||
>
|
|
||||||
> Tier 1 - Critical (>10% importance):
|
|
||||||
> - Design_elite_vs_vernacular (22.7%) - THE most discriminative dimension
|
|
||||||
> - Entanglement_flocking_vs_swarming (13.8%)
|
|
||||||
> - Design_static_vs_malleable (10.2%)
|
|
||||||
>
|
|
||||||
> Tier 2 - Important (5-10%):
|
|
||||||
> - Entanglement_self-enforcing_vs_enforced (9.2%)
|
|
||||||
> - Entanglement_obligatory_vs_voluntary (8.0%)
|
|
||||||
> - Experience_exclusion_vs_inclusion (5.9%)
|
|
||||||
>
|
|
||||||
> Tier 3 - Supplementary (<5%):
|
|
||||||
> - All remaining 17 dimensions
|
|
||||||
>
|
|
||||||
> Recommendation: Reorganize the tool to present Tier 1 dimensions first, or mark them as "core diagnostics" vs. "supplementary diagnostics."
|
|
||||||
>
|
|
||||||
> 2. Consider Reducing Low-Value Dimensions
|
|
||||||
>
|
|
||||||
> Several dimensions have low discriminative power AND low variance:
|
|
||||||
>
|
|
||||||
> Candidates for removal/consolidation:
|
|
||||||
> - Entanglement_exclusive_vs_non-exclusive (0.6% importance, σ=1.64, mean=8.5)
|
|
||||||
> - Responses heavily cluster at "non-exclusive" - not discriminating
|
|
||||||
> - Design_durable_vs_ephemeral (1.2% importance, σ=2.41)
|
|
||||||
> - Entanglement_defensible_vs_exposed (1.1% importance, σ=2.41)
|
|
||||||
>
|
|
||||||
> Recommendation: Either remove these or combine into composite measures. Going from 23→18 dimensions would reduce analyst burden by ~20% with minimal information loss.
|
|
||||||
>
|
|
||||||
> 3. Add Composite Scores 📊
|
|
||||||
>
|
|
||||||
> Since PC1 explains the main variance, create derived metrics:
|
|
||||||
>
|
|
||||||
> "Protocol Type Score" (based on PC1 loadings):
|
|
||||||
> Score = elite_vs_vernacular(0.36) + static_vs_malleable(0.33) + flocking_vs_swarming(0.31) - self-enforcing_vs_enforced(0.29)
|
|
||||||
> - High score = Institutional/Standardized
|
|
||||||
> - Low score = Vernacular/Emergent
|
|
||||||
>
|
|
||||||
> "Protocol Completeness Score" (based on PC2):
|
|
||||||
> Score = sufficient_vs_insufficient(0.43) + crystallized_vs_contested(0.38) - Kafka_vs_Whitehead(0.36)
|
|
||||||
> - Measures how "finished" vs. "kafkaesque" a protocol feels
|
|
||||||
>
|
|
||||||
> Recommendation: Display these composite scores alongside individual dimensions to provide quick high-level insights.
|
|
||||||
>
|
|
||||||
> 4. Highlight Key Correlations 🔗
|
|
||||||
>
|
|
||||||
> The tool should alert analysts to important relationships:
|
|
||||||
>
|
|
||||||
> Strong positive correlations:
|
|
||||||
> - Static ↔ Predictable (0.61)
|
|
||||||
> - Elite ↔ Static (0.53)
|
|
||||||
> - Exclusion ↔ Kafka (0.52)
|
|
||||||
>
|
|
||||||
> Strong negative correlations:
|
|
||||||
> - Elite ↔ Self-enforcing (-0.58)
|
|
||||||
> - Self-enforcing ↔ Flocking (-0.56)
|
|
||||||
>
|
|
||||||
> Recommendation: When an analyst rates a dimension, show a tooltip: "Protocols
|
|
||||||
> rated as 'elite' tend to also be 'static' and require 'enforcement'"
|
|
||||||
>
|
|
||||||
> 5. Flag Potential Redundancy
|
|
||||||
>
|
|
||||||
> Two dimension pairs show moderate correlation within the same category:
|
|
||||||
>
|
|
||||||
> 1. Entanglement_abstract_vs_embodied ↔ Entanglement_flocking_vs_swarming (r=-0.56)
|
|
||||||
> 2. Design_documenting_vs_enabling ↔ Design_static_vs_malleable (r=0.53)
|
|
||||||
>
|
|
||||||
> Recommendation: Consider merging these or making one primary and the other optional.
|
|
||||||
>
|
|
||||||
> 6. Rebalance Categories ⚖️
|
|
||||||
>
|
|
||||||
> Current split: Design (8), Entanglement (8), Experience (7)
|
|
||||||
>
|
|
||||||
> Performance by category:
|
|
||||||
> - Design: 90.4% predictive accuracy
|
|
||||||
> - Entanglement: 89.2% predictive accuracy
|
|
||||||
> - Experience: 78.3% predictive accuracy
|
|
||||||
>
|
|
||||||
> Recommendation:
|
|
||||||
> - Strengthen Design (add 1-2 high-variance dimensions)
|
|
||||||
> - Trim Experience (remove low-performers, down to 5)
|
|
||||||
> - Result: 9 Design, 8 Entanglement, 5 Experience = 22 dimensions (down from
|
|
||||||
> 23)
|
|
||||||
>
|
|
||||||
> 7. Add Diagnostic Quality Indicators
|
|
||||||
>
|
|
||||||
> Based on variance analysis, flag dimensions where responses are too clustered:
|
|
||||||
>
|
|
||||||
> - Entanglement_exclusive_vs_non-exclusive: 93% of protocols rate 7-9
|
|
||||||
> - Design_technical_vs_social: Mean=7.6, heavily skewed toward "social"
|
|
||||||
>
|
|
||||||
> Recommendation: Consider revising these gradient definitions or endpoints to achieve better distribution.
|
|
||||||
>
|
|
||||||
> 8. Create Shortened Version 🎯
|
|
||||||
>
|
|
||||||
> For rapid assessment, create a "Bicorder Core" with just the top 8 dimensions:
|
|
||||||
>
|
|
||||||
> 1. elite_vs_vernacular ⭐⭐⭐
|
|
||||||
> 2. flocking_vs_swarming
|
|
||||||
> 3. static_vs_malleable
|
|
||||||
> 4. self-enforcing_vs_enforced
|
|
||||||
> 5. obligatory_vs_voluntary
|
|
||||||
> 6. exclusion_vs_inclusion
|
|
||||||
> 7. universal_vs_particular (high variance)
|
|
||||||
> 8. explicit_vs_implicit (high variance)
|
|
||||||
>
|
|
||||||
> This captures ~65% of the discriminative power in 1/3 the time.
|
|
||||||
>
|
|
||||||
> 9. Add Comparison Features
|
|
||||||
>
|
|
||||||
> The network analysis shows some protocols are highly "central" (similar to many others):
|
|
||||||
> - VPN Usage (Circumvention Protocol)
|
|
||||||
> - Access Check-in
|
|
||||||
> - Quadratic Voting
|
|
||||||
>
|
|
||||||
> Recommendation: After rating a protocol, show: "This protocol is most similar to: [X, Y, Z]" based on dimensional proximity.
|
|
||||||
|
|
||||||
> Summary of Recommendations
|
|
||||||
|
|
||||||
> Immediate actions:
|
|
||||||
> 1. ✂️ Remove 3-5 low-value dimensions → 20 dimensions
|
|
||||||
> 2. 🔄 Reorder dimensions by importance (elite_vs_vernacular first)
|
|
||||||
> 3. ➕ Add 2 composite scores (Protocol Type, Completeness)
|
|
||||||
|
|
||||||
> Medium-term enhancements:
|
|
||||||
> 4. 🎯 Create "Bicorder Core" (8-dimension quick version)
|
|
||||||
> 5. 💡 Add contextual tooltips about correlations
|
|
||||||
> 6. 📊 Show similar protocols after assessment
|
|
||||||
|
|
||||||
> This would make the tool ~20% faster to use while maintaining 95%+ of its discriminative power.
|
|
||||||
|
|
||||||
Questions:
|
|
||||||
|
|
||||||
* What makes the elite/vernacular distinction so "important"? Is it because of some salience of the description?
|
|
||||||
* Why is the sovereign/subsidiary distinction, which is such a central part of the theorizing here, not more "important"? Would this change with a different description of the values?
|
|
||||||
|
|
||||||
|
|
||||||
### Future work: Description modification
|
### Future work: Description modification
|
||||||
|
|
||||||
|
|||||||
Reference in new issue
Block a user