refactor: reorganize analysis data to support multiple runs
Strategy: key runs by bicorder version (not date), promote shared inputs
to analysis/data/, and stamp every output with its bicorder_version so
a re-run on edited gradients is self-describing.
Data layout:
- Promote the shared protocol inputs out of the run directory:
analysis/data/protocols_edited.csv (411 cleaned protocols)
analysis/data/protocols_raw.csv (774 un-cleaned entries)
- Rename the v1.2.6 synthetic run:
data/synthetic_20251116/ -> data/synthetic_1.2.6/
so the bicorder version it was scored against is explicit (gradient
structure changes between versions make date-based names ambiguous)
Provenance:
- bicorder_analyze.py now writes a 'bicorder_version' column into every
output readings.csv, recording which gradient structure produced it
Scripts:
- Update the real code defaults that pointed at the old run path
(bicorder_classifier.py, classify_readings.py, sync_readings.sh,
compare_analyses.py) and refresh docstring/help examples
- Remove a stray committed __pycache__/.pyc
Docs: analysis/README.md documents the new layout + how to add a run;
WORKFLOW.md, TEST_COMMANDS.md, INTEGRATION_GUIDE.md paths updated.
This commit is contained in:
1 parent
459015fe17
commit
6ae77a4f9b
474 files changed
+90
-66
No files matched your search
@@ -55,6 +55,7 @@ def process_csv(input_csv, output_csv, bicorder_path, analyst=None, standpoint=N
|
||||
# Load bicorder configuration
|
||||
bicorder_data = load_bicorder_config(bicorder_path)
|
||||
gradients = extract_gradients(bicorder_data)
|
||||
bicorder_version = bicorder_data.get('version', '')
|
||||
|
||||
with open(input_csv, 'r', encoding='utf-8') as infile, \
|
||||
open(output_csv, 'w', newline='', encoding='utf-8') as outfile:
|
||||
@@ -68,6 +69,9 @@ def process_csv(input_csv, output_csv, bicorder_path, analyst=None, standpoint=N
|
||||
gradient_columns = [g['column_name'] for g in gradients]
|
||||
output_fields = list(original_fields) + gradient_columns
|
||||
|
||||
# Add the bicorder version as a provenance column
|
||||
output_fields.append('bicorder_version')
|
||||
|
||||
# Add metadata columns if provided
|
||||
if analyst is not None:
|
||||
output_fields.append('analyst')
|
||||
@@ -87,6 +91,9 @@ def process_csv(input_csv, output_csv, bicorder_path, analyst=None, standpoint=N
|
||||
for gradient in gradients:
|
||||
output_row[gradient['column_name']] = ''
|
||||
|
||||
# Record which bicorder version these gradients came from
|
||||
output_row['bicorder_version'] = bicorder_version
|
||||
|
||||
# Add metadata if provided
|
||||
if analyst is not None:
|
||||
output_row['analyst'] = analyst
|
||||
|
||||
Reference in new issue
Block a user