Commit Graph
2 Commits
Author SHA1 Message Date
Protocolbot fb3bebcea0 fix: make --resume in batch pipeline actually skip completed work
Previously --resume re-queried every gradient in every row, overwriting
existing values — an interrupted run could not be resumed cheaply.

- bicorder_query.py: add --resume flag; skip gradients whose cells already
  have values, and report how many were skipped
- bicorder_batch.py: pass --resume through to query; skip fully-complete
  rows before invoking the query script; report partial rows
- bicorder_batch.py: import row/config helpers from bicorder_query instead
  of calling undefined names (would have crashed on --resume)
2026-09-23 07:58:04 -06:00
Nathan SchneiderandClaude Sonnet 4.6 897c30406b Reorganize directory, add manual dataset and sync tooling
- Move all scripts to scripts/, web assets to web/, analysis results
  into self-contained data/readings/<type>_<YYYYMMDD>/ directories
- Add data/readings/manual_20260320/ with 32 JSON readings from
  git.medlab.host/ntnsndr/protocol-bicorder-data
- Add scripts/json_to_csv.py to convert bicorder JSON files to CSV
- Add scripts/sync_readings.sh for one-command sync + re-analysis of
  any dataset backed by a .sync_source config file
- Add scripts/classify_readings.py to apply the LDA classifier to all
  readings and save per-reading cluster assignments
- Add --min-coverage flag to multivariate_analysis.py for sparse/shortform
  datasets; also applies in lda_visualization.py
- Fix lda_visualization.py NaN handling and 0-d array annotation bug
- Update README.md and WORKFLOW.md to document datasets, sync workflow,
  shortform handling, and new scripts

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-03-20 17:35:13 -06:00