Commit Graph
28 Commits
Author SHA1 Message Date
Nathan Schneider 8c44064196 data: add synthetic_1.4.0 run readings
411 protocols × 23 gradients scored against bicorder.json v1.4.0 with
gpt-oss:20b-cloud (same analyst standpoint as the 1.2.6 run); the
bicorder_version column makes the run self-describing. Analysis outputs to
follow in data/synthetic_1.4.0/analysis/.
2026-10-02 08:33:03 -06:00
Nathan Schneider 55cbd6cd5d feat: version-agnostic analysis scripts with shared version helpers
- scripts/bicorder_common.py (new): single source of truth for historical
  gradient renames (COLUMN_RENAMES), version detection (bicorder_version col
  → version col → data/<type>_<version>/ dir convention), and training-CSV
  auto-selection (find_training_csv)
- classify_readings.py: auto-select training run by recorded bicorder version
  (excludes the input itself to avoid circular training), canonicalize old
  column names, version-mismatch warnings
- bicorder_classifier.py / export_model_for_js.py: share renames + dimension
  loader; instructive error when clustering results are missing;
  bicorder_version recorded in exported models
- compare_analyses.py: CLI (reference + comparison CSVs), rename
  canonicalization so runs of different versions align on shared gradients,
  Descriptor dedup (was silently skewing merges); legacy no-arg audit intact
- scripts/univariate_analysis.py (new): per-protocol/per-gradient averages,
  distributions, summary stats; --img publishes the three README summary
  charts to img/
- sync_readings.sh: defer classifier training to auto-matching; gitignore
  analysis/.venv and __pycache__
2026-10-02 08:32:59 -06:00
Nathan Schneider a4f16e8e4d Removed raw Claude outputs from README 2026-09-23 14:44:05 -06:00
Protocolbot 161ba3b136 fix: carry institutional→formal rename through analysis rename maps
Nathan renamed the Design gradient 'institutional' → 'formal' (v1.4.0).
The analysis scripts carry historical rename maps that route old data
(elite/institutional) into the current bicorder.json gradient names, and
those maps still stopped at 'institutional' — so old readings would have
silently misaligned against the v1.4.0 structure.

- bicorder_classifier.py / export_model_for_js.py: add
  institutional_vs_vernacular → formal_vs_vernacular, keep elite→formal
- json_to_csv.py TERM_RENAMES: same two-step route
- convert_csv_to_json.py GRADIENT_MAPPINGS: elite_vs_vernacular now maps
  to 'formal'
- bicorder_classifier.py demo ratings: use the current dimension name
- analysis/README.md: example run name 1.3.0 → 1.4.0
2026-09-23 14:27:51 -06:00
Protocolbot 6ae77a4f9b refactor: reorganize analysis data to support multiple runs
Strategy: key runs by bicorder version (not date), promote shared inputs
to analysis/data/, and stamp every output with its bicorder_version so
a re-run on edited gradients is self-describing.

Data layout:
- Promote the shared protocol inputs out of the run directory:
    analysis/data/protocols_edited.csv  (411 cleaned protocols)
    analysis/data/protocols_raw.csv     (774 un-cleaned entries)
- Rename the v1.2.6 synthetic run:
    data/synthetic_20251116/ -> data/synthetic_1.2.6/
  so the bicorder version it was scored against is explicit (gradient
  structure changes between versions make date-based names ambiguous)

Provenance:
- bicorder_analyze.py now writes a 'bicorder_version' column into every
  output readings.csv, recording which gradient structure produced it

Scripts:
- Update the real code defaults that pointed at the old run path
  (bicorder_classifier.py, classify_readings.py, sync_readings.sh,
  compare_analyses.py) and refresh docstring/help examples
- Remove a stray committed __pycache__/.pyc

Docs: analysis/README.md documents the new layout + how to add a run;
WORKFLOW.md, TEST_COMMANDS.md, INTEGRATION_GUIDE.md paths updated.
2026-09-23 14:23:17 -06:00
Protocolbot e7d6465ceb docs: mark classifier integration historical; fix shortform count
- INTEGRATION_GUIDE.md: rewritten as research notes — how to reproduce the
  cluster classification with the analysis scripts, and why it was removed
  from the tool in v1.3.0
- analysis/README.md: integration section marked historical
- bicorder-app/README.md: shortform is 9 gradients, not 10
2026-09-23 07:58:17 -06:00
Protocolbot fb3bebcea0 fix: make --resume in batch pipeline actually skip completed work
Previously --resume re-queried every gradient in every row, overwriting
existing values — an interrupted run could not be resumed cheaply.

- bicorder_query.py: add --resume flag; skip gradients whose cells already
  have values, and report how many were skipped
- bicorder_batch.py: pass --resume through to query; skip fully-complete
  rows before invoking the query script; report partial rows
- bicorder_batch.py: import row/config helpers from bicorder_query instead
  of calling undefined names (would have crashed on --resume)
2026-09-23 07:58:04 -06:00
Nathan SchneiderandClaude Sonnet 4.6 60e83783ec Flatten data/readings/ → data/
Remove the intermediate readings/ subdirectory level — dataset naming
(synthetic_YYYYMMDD, manual_YYYYMMDD) already encodes what the data is.
Update all path references across scripts and docs accordingly.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-03-20 17:46:23 -06:00
Nathan SchneiderandClaude Sonnet 4.6 1a80219a25 Remove web/ prototype; update docs to reflect app integration
The web/ directory (bicorder-classifier.js, .d.ts, test-classifier.mjs)
was a prototype superseded by bicorder-app/src/bicorder-classifier.ts.
The only integration point between this analysis directory and the app is
bicorder_model.json, which Vite reads at build time.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-03-20 17:39:25 -06:00
Nathan SchneiderandClaude Sonnet 4.6 897c30406b Reorganize directory, add manual dataset and sync tooling
- Move all scripts to scripts/, web assets to web/, analysis results
  into self-contained data/readings/<type>_<YYYYMMDD>/ directories
- Add data/readings/manual_20260320/ with 32 JSON readings from
  git.medlab.host/ntnsndr/protocol-bicorder-data
- Add scripts/json_to_csv.py to convert bicorder JSON files to CSV
- Add scripts/sync_readings.sh for one-command sync + re-analysis of
  any dataset backed by a .sync_source config file
- Add scripts/classify_readings.py to apply the LDA classifier to all
  readings and save per-reading cluster assignments
- Add --min-coverage flag to multivariate_analysis.py for sparse/shortform
  datasets; also applies in lda_visualization.py
- Fix lda_visualization.py NaN handling and 0-d array annotation bug
- Update README.md and WORKFLOW.md to document datasets, sync workflow,
  shortform handling, and new scripts

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-03-20 17:35:13 -06:00
Nathan SchneiderandClaude Sonnet 4.6 f1ae9cac1f Derive classifier dimensions from bicorder.json automatically
Both export_model_for_js.py and bicorder_classifier.py now read
DIMENSIONS and KEY_DIMENSIONS directly from bicorder.json at runtime,
so the model stays in sync whenever gradient terms are renamed or
added. A COLUMN_RENAMES dict handles historical CSV column name
changes. The model now includes bicorder_version so the app's version
check works correctly.

Regenerated bicorder_model.json against bicorder.json v1.2.6 with
correct dimension names, 9 key dimensions from shortform flags, and
updated thresholds.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-03-20 15:13:54 -06:00
Nathan Schneider 730075f757 Tweaks to analysis readme 2026-03-03 20:42:07 -07:00
Nathan Schneider 62a6b0ba2c analysis updates 2026-02-21 17:13:15 -07:00
Nathan Schneider b004e9fb19 Analysis changes 2026-01-25 22:55:30 -07:00
Nathan Schneider 57b780fe95 Additional analysis in app and tweaked json descriptions 2026-01-25 22:48:54 -07:00
Nathan Schneider d1f288cda8 Adjusted the classifier to reduce false flags 2026-01-14 22:44:32 -07:00
Nathan Schneider 6230e6da7b Added arrows to focused view for clarity 2026-01-14 22:35:27 -07:00
Nathan Schneider 3355ed9f02 Added synthetic_readings JSON files to analysis 2026-01-14 21:59:28 -07:00
Nathan Schneider 7829141756 Analysis updates and json tweaks 2026-01-14 21:42:23 -07:00
Nathan Schneider 1b508b911f Added classifer analysis to bicorder ascii and web app 2025-12-21 21:38:39 -07:00
Nathan Schneider b541f85553 Additional analysis: examining clustering via LDA 2025-12-19 11:40:45 -07:00
Nathan Schneider 7be67c9eb5 Adjustments to analysis and AI disclaimer 2025-11-23 09:27:51 -05:00
Nathan Schneider 998c93b3f9 Legacy bicorder version added to analysis folder 2025-11-21 19:35:24 -07:00
Nathan Schneider fa527bd1f1 Initial analysis 2025-11-21 19:34:33 -07:00
Nathan Schneider 04eee1360f Formatting and clarity improvements 2025-11-21 11:43:13 -07:00
Nathan Schneider dcfd37fa4c Initial analysis complete 2025-11-16 23:47:10 -07:00
Nathan Schneider 815ed9d6f4 Set up analysis scripts 2025-10-30 10:56:21 -06:00
Nathan Schneider d2da0425c6 Restructured data/ into analysis/ 2025-10-28 11:04:06 -06:00