refactor: reorganize analysis data to support multiple runs
Strategy: key runs by bicorder version (not date), promote shared inputs
to analysis/data/, and stamp every output with its bicorder_version so
a re-run on edited gradients is self-describing.
Data layout:
- Promote the shared protocol inputs out of the run directory:
analysis/data/protocols_edited.csv (411 cleaned protocols)
analysis/data/protocols_raw.csv (774 un-cleaned entries)
- Rename the v1.2.6 synthetic run:
data/synthetic_20251116/ -> data/synthetic_1.2.6/
so the bicorder version it was scored against is explicit (gradient
structure changes between versions make date-based names ambiguous)
Provenance:
- bicorder_analyze.py now writes a 'bicorder_version' column into every
output readings.csv, recording which gradient structure produced it
Scripts:
- Update the real code defaults that pointed at the old run path
(bicorder_classifier.py, classify_readings.py, sync_readings.sh,
compare_analyses.py) and refresh docstring/help examples
- Remove a stray committed __pycache__/.pyc
Docs: analysis/README.md documents the new layout + how to add a run;
WORKFLOW.md, TEST_COMMANDS.md, INTEGRATION_GUIDE.md paths updated.
This commit is contained in:
1 parent
459015fe17
commit
6ae77a4f9b
474 files changed
+90
-66
No files matched your search
@@ -7,7 +7,7 @@ Run these tests in order to verify the refactored code works correctly.
|
||||
Test that prompts are generated correctly with protocol context:
|
||||
|
||||
```bash
|
||||
python3 scripts/bicorder_query.py data/synthetic_20251116/protocols_edited.csv 1 --dry-run | head -80
|
||||
python3 scripts/bicorder_query.py data/protocols_edited.csv 1 --dry-run | head -80
|
||||
```
|
||||
|
||||
**Expected result:**
|
||||
@@ -21,7 +21,7 @@ python3 scripts/bicorder_query.py data/synthetic_20251116/protocols_edited.csv 1
|
||||
Check that the analyze script still creates proper CSV structure:
|
||||
|
||||
```bash
|
||||
python3 scripts/bicorder_analyze.py data/synthetic_20251116/protocols_edited.csv -o test_output.csv
|
||||
python3 scripts/bicorder_analyze.py data/protocols_edited.csv -o test_output.csv
|
||||
head -1 test_output.csv | tr ',' '\n' | grep -E "(explicit|precise|elite)" | head -5
|
||||
```
|
||||
|
||||
@@ -76,7 +76,7 @@ llm logs list | grep -i bicorder
|
||||
Test batch processing on rows 1-3:
|
||||
|
||||
```bash
|
||||
python3 scripts/bicorder_batch.py data/synthetic_20251116/protocols_edited.csv -o test_batch_output.csv --start 1 --end 3 -m gpt-4o-mini
|
||||
python3 scripts/bicorder_batch.py data/protocols_edited.csv -o test_batch_output.csv --start 1 --end 3 -m gpt-4o-mini
|
||||
```
|
||||
|
||||
**Expected result:**
|
||||
@@ -106,7 +106,7 @@ with open('test_batch_output.csv') as f:
|
||||
Test that model parameter works in dry run:
|
||||
|
||||
```bash
|
||||
python3 scripts/bicorder_query.py data/synthetic_20251116/protocols_edited.csv 5 --dry-run -m mistral | head -50
|
||||
python3 scripts/bicorder_query.py data/protocols_edited.csv 5 --dry-run -m mistral | head -50
|
||||
```
|
||||
|
||||
**Expected result:**
|
||||
@@ -129,11 +129,11 @@ Compare the new standalone prompts vs old system prompt approach:
|
||||
|
||||
```bash
|
||||
# New approach - protocol context in each prompt
|
||||
python3 scripts/bicorder_query.py data/synthetic_20251116/protocols_edited.csv 1 --dry-run | grep -A 5 "Analyze this protocol"
|
||||
python3 scripts/bicorder_query.py data/protocols_edited.csv 1 --dry-run | grep -A 5 "Analyze this protocol"
|
||||
|
||||
# Old approach would have had protocol in system prompt only (no longer used)
|
||||
# Verify that protocol context appears in EVERY gradient prompt
|
||||
python3 scripts/bicorder_query.py data/synthetic_20251116/protocols_edited.csv 1 --dry-run | grep -c "Analyze this protocol"
|
||||
python3 scripts/bicorder_query.py data/protocols_edited.csv 1 --dry-run | grep -c "Analyze this protocol"
|
||||
```
|
||||
|
||||
**Expected result:**
|
||||
|
||||
Reference in new issue
Block a user