upload: implementation (1).md
Browse files- implementation (1).md +1030 -0
implementation (1).md
ADDED
|
@@ -0,0 +1,1030 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# Cross-Session Continuity Env β Implementation Plan (v2)
|
| 2 |
+
|
| 3 |
+
> **Changelog from v1:** Addressed 20 potential failure modes identified in review.
|
| 4 |
+
> Each section marked [UPDATED], [NEW], or [UNCHANGED] for traceability.
|
| 5 |
+
|
| 6 |
+
---
|
| 7 |
+
|
| 8 |
+
## 1. Problem Statement [UNCHANGED]
|
| 9 |
+
|
| 10 |
+
**Capability Gap:** LLMs have no persistent memory across sessions. When a session ends,
|
| 11 |
+
everything is gone. In real-world usage this is a critical failure mode β long tasks
|
| 12 |
+
(codebases, research, planning) rarely fit in a single context window.
|
| 13 |
+
|
| 14 |
+
**What we train:** Can RL teach an LLM to write surgical, information-dense handoff notes
|
| 15 |
+
to its future self, such that a cold-start agent in session 2 can complete the task
|
| 16 |
+
successfully using only those notes?
|
| 17 |
+
|
| 18 |
+
**Why it's novel:** No existing RL environment specifically trains or benchmarks
|
| 19 |
+
cross-session state transfer behavior. This is underexplored and publishable.
|
| 20 |
+
|
| 21 |
+
**Theme:** Primarily Theme 2 (Long-Horizon Planning). Secondary fit with Theme 3.1 β
|
| 22 |
+
agent uses real tools (file I/O, test runner) in a dynamic coding environment.
|
| 23 |
+
|
| 24 |
+
---
|
| 25 |
+
|
| 26 |
+
## 2. High-Level Architecture [UPDATED]
|
| 27 |
+
|
| 28 |
+
```
|
| 29 |
+
Episode = Session 1 + Session 2 (ONE training episode, ONE reward signal)
|
| 30 |
+
|
| 31 |
+
Session 1:
|
| 32 |
+
Agent receives β task description + starter code + tool access
|
| 33 |
+
Agent works β reads files, writes code, runs tests
|
| 34 |
+
[Auxiliary rewards fire here β see Section 8]
|
| 35 |
+
Agent ends β calls write_handoff(structured_note) β session 1 terminates
|
| 36 |
+
|
| 37 |
+
β [handoff.md is the ONLY bridge]
|
| 38 |
+
β [filesystem wiped β no code persists]
|
| 39 |
+
β [function/variable names randomized per episode]
|
| 40 |
+
|
| 41 |
+
Session 2:
|
| 42 |
+
Agent receives β ONLY handoff.md + same tool access
|
| 43 |
+
Agent must call parse_handoff() before file access (enforced)
|
| 44 |
+
Agent works β picks up, finishes implementation
|
| 45 |
+
Agent ends β calls submit() β visible + hidden tests run β reward computed
|
| 46 |
+
|
| 47 |
+
Reward flows back through both sessions via GRPO (with normalization)
|
| 48 |
+
PPO run in parallel as stability baseline
|
| 49 |
+
```
|
| 50 |
+
|
| 51 |
+
---
|
| 52 |
+
|
| 53 |
+
## 3. Repository Structure [UPDATED]
|
| 54 |
+
|
| 55 |
+
```
|
| 56 |
+
cross-session-continuity-env/
|
| 57 |
+
β
|
| 58 |
+
βββ openenv.yaml
|
| 59 |
+
βββ README.md
|
| 60 |
+
βββ requirements.txt # pinned: openenv==x.y.z
|
| 61 |
+
β
|
| 62 |
+
βββ server/
|
| 63 |
+
β βββ env.py # MCPEnvironment subclass
|
| 64 |
+
β βββ task_generator.py # task + test generation with name randomization
|
| 65 |
+
β βββ session_manager.py # session 1 β 2 transition, filesystem wipe
|
| 66 |
+
β βββ sandbox.py # safe execution, strict ulimits
|
| 67 |
+
β βββ handoff_validator.py # NEW: validates handoff structure
|
| 68 |
+
β βββ rewards/
|
| 69 |
+
β βββ rubric.py # composable rubrics (UPDATED)
|
| 70 |
+
β βββ auxiliary.py # NEW: session 1 auxiliary rewards
|
| 71 |
+
β
|
| 72 |
+
βββ client/
|
| 73 |
+
β βββ agent.py # agent loop β no server imports, with retry logic
|
| 74 |
+
β
|
| 75 |
+
βββ tasks/
|
| 76 |
+
β βββ easy/ # single file, 3 visible + 1 hidden test
|
| 77 |
+
β βββ medium/ # 2-3 files, 5 visible + 2 hidden tests
|
| 78 |
+
β βββ hard/ # 5 files, 8 visible + 3 hidden tests
|
| 79 |
+
β βββ eval_holdout/ # NEW: unseen tasks for evaluation only
|
| 80 |
+
β
|
| 81 |
+
βββ training/
|
| 82 |
+
β βββ train_grpo.ipynb # primary training (GRPO)
|
| 83 |
+
β βββ train_ppo.ipynb # NEW: PPO baseline for stability comparison
|
| 84 |
+
β βββ grpo_config.yaml
|
| 85 |
+
β
|
| 86 |
+
βββ evals/
|
| 87 |
+
β βββ baselines/
|
| 88 |
+
β β βββ no_handoff.py # NEW: session 2 with no note at all
|
| 89 |
+
β β βββ random_handoff.py # NEW: random text as handoff
|
| 90 |
+
β β βββ full_transcript.py # NEW: upper bound β full S1 transcript
|
| 91 |
+
β βββ ablations/
|
| 92 |
+
β β βββ no_compression_reward.py # NEW: ablation
|
| 93 |
+
β β βββ no_linearity_reward.py # NEW: ablation
|
| 94 |
+
β β βββ no_auxiliary_reward.py # NEW: ablation
|
| 95 |
+
β βββ trained_run.py
|
| 96 |
+
β
|
| 97 |
+
βββ plots/ # all committed as PNG with captions
|
| 98 |
+
β βββ reward_curve.png
|
| 99 |
+
β βββ handoff_length_curve.png
|
| 100 |
+
β βββ baseline_vs_trained.png # all 4 baselines on same axes
|
| 101 |
+
β βββ ablation_comparison.png # NEW
|
| 102 |
+
β βββ difficulty_breakdown.png # NEW: easy/medium/hard separately
|
| 103 |
+
β βββ handoff_diff_over_epochs.png # NEW: interpretability
|
| 104 |
+
β
|
| 105 |
+
βββ demos/
|
| 106 |
+
βββ recorded_run_seed42.url # URL only β no large files in repo
|
| 107 |
+
```
|
| 108 |
+
|
| 109 |
+
---
|
| 110 |
+
|
| 111 |
+
## 4. OpenEnv Compliance [UNCHANGED]
|
| 112 |
+
|
| 113 |
+
### 4.1 openenv.yaml
|
| 114 |
+
|
| 115 |
+
```yaml
|
| 116 |
+
name: cross-session-continuity-env
|
| 117 |
+
version: 0.1.0
|
| 118 |
+
theme: long-horizon-planning
|
| 119 |
+
description: >
|
| 120 |
+
An RL environment where an LLM agent must complete a coding task across two
|
| 121 |
+
sessions with zero shared memory. The agent writes a structured handoff note
|
| 122 |
+
at the end of session 1; session 2 receives only that note. Reward depends
|
| 123 |
+
entirely on session 2 success.
|
| 124 |
+
entry: server/env.py
|
| 125 |
+
tools:
|
| 126 |
+
- read_file
|
| 127 |
+
- write_file
|
| 128 |
+
- run_tests
|
| 129 |
+
- write_handoff
|
| 130 |
+
- parse_handoff
|
| 131 |
+
- submit
|
| 132 |
+
sessions: 2
|
| 133 |
+
difficulty_levels:
|
| 134 |
+
- easy
|
| 135 |
+
- medium
|
| 136 |
+
- hard
|
| 137 |
+
```
|
| 138 |
+
|
| 139 |
+
### 4.2 Reserved Tool Names β Avoided
|
| 140 |
+
|
| 141 |
+
`reset`, `step`, `state`, `close` are OpenEnv reserved β none used.
|
| 142 |
+
Our tools: `read_file`, `write_file`, `run_tests`, `write_handoff`, `parse_handoff`, `submit` β all clear.
|
| 143 |
+
|
| 144 |
+
### 4.3 Client/Server Separation
|
| 145 |
+
|
| 146 |
+
- `client/agent.py` talks to env via MCP protocol only
|
| 147 |
+
- Client never imports from `server/`
|
| 148 |
+
- All state lives server-side
|
| 149 |
+
|
| 150 |
+
### 4.4 Gym-style API
|
| 151 |
+
|
| 152 |
+
```python
|
| 153 |
+
env.reset() # starts episode, returns session 1 observation
|
| 154 |
+
env.step() # action β (obs, reward, done, info)
|
| 155 |
+
env.state() # current env state dict
|
| 156 |
+
```
|
| 157 |
+
|
| 158 |
+
---
|
| 159 |
+
|
| 160 |
+
## 5. Environment Implementation [UPDATED]
|
| 161 |
+
|
| 162 |
+
Key changes from v1:
|
| 163 |
+
- Dynamic step limits by difficulty
|
| 164 |
+
- Auxiliary reward hooks in session 1
|
| 165 |
+
- Handoff structure validation before session 2 starts
|
| 166 |
+
- Invalid action handling with retry budget
|
| 167 |
+
- Agent must call `parse_handoff()` before file access in session 2
|
| 168 |
+
- Filesystem wiped on session transition
|
| 169 |
+
|
| 170 |
+
```python
|
| 171 |
+
# server/env.py
|
| 172 |
+
from openenv import MCPEnvironment
|
| 173 |
+
from .task_generator import TaskGenerator
|
| 174 |
+
from .session_manager import SessionManager
|
| 175 |
+
from .sandbox import Sandbox
|
| 176 |
+
from .rewards.rubric import ContinuityRubric
|
| 177 |
+
from .rewards.auxiliary import AuxiliaryRewarder
|
| 178 |
+
from .handoff_validator import HandoffValidator
|
| 179 |
+
|
| 180 |
+
STEP_LIMITS = {"easy": 20, "medium": 35, "hard": 55}
|
| 181 |
+
|
| 182 |
+
class CrossSessionContinuityEnv(MCPEnvironment):
|
| 183 |
+
|
| 184 |
+
def __init__(self, difficulty="medium"):
|
| 185 |
+
self.task_gen = TaskGenerator(difficulty)
|
| 186 |
+
self.session_mgr = SessionManager()
|
| 187 |
+
self.sandbox = Sandbox(timeout=10)
|
| 188 |
+
self.rubric = ContinuityRubric()
|
| 189 |
+
self.aux = AuxiliaryRewarder()
|
| 190 |
+
self.validator = HandoffValidator()
|
| 191 |
+
self.difficulty = difficulty
|
| 192 |
+
self.step_limit = STEP_LIMITS[difficulty]
|
| 193 |
+
|
| 194 |
+
def reset(self, task_id=None, seed=None):
|
| 195 |
+
self.task = self.task_gen.sample(task_id, seed=seed) # names randomized
|
| 196 |
+
self.session = 1
|
| 197 |
+
self.handoff = None
|
| 198 |
+
self.step_count = 0
|
| 199 |
+
self.invalid_action_count = 0
|
| 200 |
+
self.retry_budget = 3
|
| 201 |
+
self.s1_test_history = []
|
| 202 |
+
self.s2_edit_history = []
|
| 203 |
+
self.handoff_parsed = False
|
| 204 |
+
self.s2_failed_runs = 0
|
| 205 |
+
|
| 206 |
+
return {
|
| 207 |
+
"session": 1,
|
| 208 |
+
"task": self.task.description,
|
| 209 |
+
"starter_code": self.task.starter_code,
|
| 210 |
+
"message": "Session 1 started. Complete what you can, then call write_handoff().",
|
| 211 |
+
"step_limit": self.step_limit
|
| 212 |
+
}
|
| 213 |
+
|
| 214 |
+
def step(self, action):
|
| 215 |
+
self.step_count += 1
|
| 216 |
+
|
| 217 |
+
# Step limit enforcement
|
| 218 |
+
if self.step_count > self.step_limit and self.session == 1:
|
| 219 |
+
return {
|
| 220 |
+
"warning": "Step limit reached. Call write_handoff() now or episode terminates.",
|
| 221 |
+
"penalty": -0.1
|
| 222 |
+
}
|
| 223 |
+
|
| 224 |
+
# Invalid action guard
|
| 225 |
+
if not self._is_valid_action(action):
|
| 226 |
+
self.invalid_action_count += 1
|
| 227 |
+
self.retry_budget -= 1
|
| 228 |
+
if self.retry_budget <= 0:
|
| 229 |
+
return {"done": True, "reward": 0.0, "error": "Retry budget exhausted"}
|
| 230 |
+
return {"error": f"Invalid action '{action.tool}'. Retries left: {self.retry_budget}"}
|
| 231 |
+
|
| 232 |
+
if action.tool == "read_file":
|
| 233 |
+
if self.session == 2 and not self.handoff_parsed:
|
| 234 |
+
return {"error": "Call parse_handoff() before accessing files in session 2."}
|
| 235 |
+
content = self.task.files.get(action.path, "File not found.")
|
| 236 |
+
return {"output": content, "session": self.session}
|
| 237 |
+
|
| 238 |
+
if action.tool == "parse_handoff":
|
| 239 |
+
if self.session != 2:
|
| 240 |
+
return {"error": "parse_handoff only available in session 2"}
|
| 241 |
+
self.handoff_parsed = True
|
| 242 |
+
return {"output": self.handoff, "session": 2}
|
| 243 |
+
|
| 244 |
+
if action.tool == "write_file":
|
| 245 |
+
prev = self.task.files.get(action.path, "")
|
| 246 |
+
self.task.files[action.path] = action.content
|
| 247 |
+
if self.session == 2:
|
| 248 |
+
self.s2_edit_history.append({"path": action.path,
|
| 249 |
+
"prev": prev, "new": action.content})
|
| 250 |
+
return {"output": f"Written to {action.path}", "session": self.session}
|
| 251 |
+
|
| 252 |
+
if action.tool == "run_tests":
|
| 253 |
+
result = self.sandbox.run_tests(self.task.files, self.task.test_code)
|
| 254 |
+
if self.session == 1:
|
| 255 |
+
self.s1_test_history.append(result.passed)
|
| 256 |
+
aux = self.aux.s1_reward(result, self.task)
|
| 257 |
+
return {"output": result.summary, "passed": result.passed,
|
| 258 |
+
"auxiliary_reward": aux, "session": 1}
|
| 259 |
+
else:
|
| 260 |
+
if result.passed == 0:
|
| 261 |
+
self.s2_failed_runs += 1
|
| 262 |
+
return {"output": result.summary, "passed": result.passed, "session": 2}
|
| 263 |
+
|
| 264 |
+
if action.tool == "write_handoff":
|
| 265 |
+
if self.session != 1:
|
| 266 |
+
return {"error": "write_handoff only available in session 1"}
|
| 267 |
+
validation = self.validator.validate(action.content)
|
| 268 |
+
if not validation.valid:
|
| 269 |
+
return {"error": f"Handoff rejected: {validation.reason}. "
|
| 270 |
+
f"Required sections: {self.validator.REQUIRED_SECTIONS}"}
|
| 271 |
+
self.handoff = action.content
|
| 272 |
+
self.session = 2
|
| 273 |
+
self.handoff_parsed = False
|
| 274 |
+
self.task = self.session_mgr.transition(self.task) # wipe filesystem
|
| 275 |
+
self.retry_budget = 3
|
| 276 |
+
return {
|
| 277 |
+
"session": 2,
|
| 278 |
+
"message": "Session 2 started. Call parse_handoff() first."
|
| 279 |
+
}
|
| 280 |
+
|
| 281 |
+
if action.tool == "submit":
|
| 282 |
+
if self.session != 2:
|
| 283 |
+
return {"error": "submit only available in session 2"}
|
| 284 |
+
visible = self.sandbox.run_tests(self.task.files, self.task.test_code)
|
| 285 |
+
hidden = self.sandbox.run_tests(self.task.files, self.task.hidden_test_code)
|
| 286 |
+
reward = self.rubric.score(
|
| 287 |
+
visible_results=visible,
|
| 288 |
+
hidden_results=hidden,
|
| 289 |
+
handoff=self.handoff,
|
| 290 |
+
s2_edit_history=self.s2_edit_history,
|
| 291 |
+
s2_failed_runs=self.s2_failed_runs,
|
| 292 |
+
invalid_actions=self.invalid_action_count
|
| 293 |
+
)
|
| 294 |
+
return {"done": True, "reward": reward,
|
| 295 |
+
"visible": visible.summary, "hidden": hidden.summary}
|
| 296 |
+
|
| 297 |
+
def state(self):
|
| 298 |
+
return {
|
| 299 |
+
"session": self.session,
|
| 300 |
+
"step_count": self.step_count,
|
| 301 |
+
"step_limit": self.step_limit,
|
| 302 |
+
"handoff_written": self.handoff is not None,
|
| 303 |
+
"handoff_length": len(self.handoff.split()) if self.handoff else 0,
|
| 304 |
+
"difficulty": self.difficulty,
|
| 305 |
+
"invalid_actions": self.invalid_action_count
|
| 306 |
+
}
|
| 307 |
+
|
| 308 |
+
def _is_valid_action(self, action):
|
| 309 |
+
s1_tools = {"read_file", "write_file", "run_tests", "write_handoff"}
|
| 310 |
+
s2_tools = {"parse_handoff", "read_file", "write_file", "run_tests", "submit"}
|
| 311 |
+
return action.tool in (s1_tools if self.session == 1 else s2_tools)
|
| 312 |
+
```
|
| 313 |
+
|
| 314 |
+
---
|
| 315 |
+
|
| 316 |
+
## 6. Handoff Format β Standardized [NEW]
|
| 317 |
+
|
| 318 |
+
**Issue addressed (#19):** Free-form text leads to inconsistent quality and lets the agent
|
| 319 |
+
game the compression metric with dense-but-useless prose.
|
| 320 |
+
|
| 321 |
+
**Fix:** Enforce a required 6-section structure. `HandoffValidator` rejects the note and
|
| 322 |
+
returns an error (not a penalty) so the agent can retry within its retry budget.
|
| 323 |
+
|
| 324 |
+
### 6.1 Required handoff template
|
| 325 |
+
|
| 326 |
+
```
|
| 327 |
+
TASK:
|
| 328 |
+
[one sentence: what the overall task is]
|
| 329 |
+
|
| 330 |
+
COMPLETED:
|
| 331 |
+
[bullet list: what is fully implemented and verified by tests]
|
| 332 |
+
|
| 333 |
+
REMAINING:
|
| 334 |
+
[bullet list: what session 2 must still implement]
|
| 335 |
+
|
| 336 |
+
KEY FUNCTIONS:
|
| 337 |
+
[function/class names, signatures, and brief purpose]
|
| 338 |
+
|
| 339 |
+
EDGE CASES:
|
| 340 |
+
[constraints or tricky logic discovered in session 1]
|
| 341 |
+
|
| 342 |
+
NEXT STEPS:
|
| 343 |
+
[ordered list: what session 2 should do first]
|
| 344 |
+
```
|
| 345 |
+
|
| 346 |
+
### 6.2 HandoffValidator
|
| 347 |
+
|
| 348 |
+
```python
|
| 349 |
+
# server/handoff_validator.py
|
| 350 |
+
|
| 351 |
+
class HandoffValidator:
|
| 352 |
+
REQUIRED_SECTIONS = ["TASK:", "COMPLETED:", "REMAINING:",
|
| 353 |
+
"KEY FUNCTIONS:", "EDGE CASES:", "NEXT STEPS:"]
|
| 354 |
+
MAX_CODE_BLOCK_LINES = 5 # prevents code dumping
|
| 355 |
+
MAX_TOKENS = 400 # hard ceiling
|
| 356 |
+
|
| 357 |
+
def validate(self, content: str) -> ValidationResult:
|
| 358 |
+
for section in self.REQUIRED_SECTIONS:
|
| 359 |
+
if section not in content:
|
| 360 |
+
return ValidationResult(valid=False,
|
| 361 |
+
reason=f"Missing required section: '{section}'")
|
| 362 |
+
|
| 363 |
+
code_lines = self._count_code_block_lines(content)
|
| 364 |
+
if code_lines > self.MAX_CODE_BLOCK_LINES:
|
| 365 |
+
return ValidationResult(valid=False,
|
| 366 |
+
reason=f"Code block too long ({code_lines} lines, max {self.MAX_CODE_BLOCK_LINES}).")
|
| 367 |
+
|
| 368 |
+
token_count = len(content.split())
|
| 369 |
+
if token_count > self.MAX_TOKENS:
|
| 370 |
+
return ValidationResult(valid=False,
|
| 371 |
+
reason=f"Handoff too long ({token_count} tokens, max {self.MAX_TOKENS}).")
|
| 372 |
+
|
| 373 |
+
return ValidationResult(valid=True)
|
| 374 |
+
|
| 375 |
+
def _count_code_block_lines(self, content):
|
| 376 |
+
in_block, count = False, 0
|
| 377 |
+
for line in content.split("\n"):
|
| 378 |
+
if line.strip().startswith("```"):
|
| 379 |
+
in_block = not in_block
|
| 380 |
+
elif in_block:
|
| 381 |
+
count += 1
|
| 382 |
+
return count
|
| 383 |
+
```
|
| 384 |
+
|
| 385 |
+
**Why this prevents gaming:** Code dumps are blocked. The agent must write structured
|
| 386 |
+
prose. The reconstruction penalty in the rubric catches the remaining shortcut β
|
| 387 |
+
session 2 ignoring the note and reconstructing from pretrained priors.
|
| 388 |
+
|
| 389 |
+
---
|
| 390 |
+
|
| 391 |
+
## 7. Task Generator [UPDATED]
|
| 392 |
+
|
| 393 |
+
### 7.1 Name Randomization (addresses issue #5 β session separation)
|
| 394 |
+
|
| 395 |
+
Each episode, function and variable names are remapped so the agent cannot reconstruct
|
| 396 |
+
the solution from pretrained knowledge alone without reading the handoff.
|
| 397 |
+
|
| 398 |
+
```python
|
| 399 |
+
# server/task_generator.py
|
| 400 |
+
import random
|
| 401 |
+
|
| 402 |
+
NAME_BANK = {
|
| 403 |
+
"merge_intervals": ["combine_ranges", "fuse_spans", "join_segments"],
|
| 404 |
+
"RateLimiter": ["ThrottleGuard", "RequestBucket", "AccessGate"],
|
| 405 |
+
"process_data": ["transform_records", "handle_payload", "digest_input"],
|
| 406 |
+
# expanded for each task in the bank
|
| 407 |
+
}
|
| 408 |
+
|
| 409 |
+
class TaskGenerator:
|
| 410 |
+
def sample(self, task_id=None, seed=None):
|
| 411 |
+
if seed:
|
| 412 |
+
random.seed(seed)
|
| 413 |
+
task = self._load_template(task_id)
|
| 414 |
+
task = self._randomize_names(task)
|
| 415 |
+
task = self._inject_hidden_tests(task)
|
| 416 |
+
return task
|
| 417 |
+
|
| 418 |
+
def _randomize_names(self, task):
|
| 419 |
+
for canonical, variants in NAME_BANK.items():
|
| 420 |
+
replacement = random.choice(variants)
|
| 421 |
+
task.description = task.description.replace(canonical, replacement)
|
| 422 |
+
task.starter_code = {k: v.replace(canonical, replacement)
|
| 423 |
+
for k, v in task.starter_code.items()}
|
| 424 |
+
task.test_code = task.test_code.replace(canonical, replacement)
|
| 425 |
+
return task
|
| 426 |
+
```
|
| 427 |
+
|
| 428 |
+
### 7.2 Hidden Tests (addresses issue #4 β test suite exploitability)
|
| 429 |
+
|
| 430 |
+
Every task has visible tests (shown via `run_tests`) and hidden tests (only run at `submit`).
|
| 431 |
+
The agent cannot overfit to the visible test surface.
|
| 432 |
+
|
| 433 |
+
```
|
| 434 |
+
easy: 3 visible + 1 hidden adversarial
|
| 435 |
+
medium: 5 visible + 2 hidden adversarial
|
| 436 |
+
hard: 8 visible + 3 hidden adversarial
|
| 437 |
+
```
|
| 438 |
+
|
| 439 |
+
Hidden tests are hand-written: empty inputs, max-size inputs, concurrent calls, type
|
| 440 |
+
coercions β things a template-following agent won't naturally handle.
|
| 441 |
+
|
| 442 |
+
### 7.3 Handoff-Critical Task Design (addresses issue #7 β difficulty calibration)
|
| 443 |
+
|
| 444 |
+
All tasks are designed so session 1 **cannot** finish within the step limit. Verified
|
| 445 |
+
empirically: step limits allow ~60-70% task completion in session 1. Any task where
|
| 446 |
+
session 1 finishes fully is moved to a warmup set and excluded from training.
|
| 447 |
+
|
| 448 |
+
### 7.4 Eval Holdout Set (addresses issue #11 β template overfitting)
|
| 449 |
+
|
| 450 |
+
`tasks/eval_holdout/` β 10 tasks never seen during training. Used only for final
|
| 451 |
+
evaluation to check generalization. Never used in curriculum or hyperparameter tuning.
|
| 452 |
+
|
| 453 |
+
---
|
| 454 |
+
|
| 455 |
+
## 8. Reward Rubric [UPDATED]
|
| 456 |
+
|
| 457 |
+
### 8.1 Session 1 Auxiliary Rewards (addresses issue #1 β credit assignment)
|
| 458 |
+
|
| 459 |
+
Session 1 has no direct reward β credit assignment across two sessions is the core
|
| 460 |
+
RL challenge here. Pure GRPO on delayed reward causes early plateau.
|
| 461 |
+
|
| 462 |
+
**Fix:** Shaped auxiliary rewards during session 1, decaying over training.
|
| 463 |
+
|
| 464 |
+
```python
|
| 465 |
+
# server/rewards/auxiliary.py
|
| 466 |
+
|
| 467 |
+
class AuxiliaryRewarder:
|
| 468 |
+
|
| 469 |
+
def s1_reward(self, test_result, task):
|
| 470 |
+
reward = 0.0
|
| 471 |
+
if test_result.compiled:
|
| 472 |
+
reward += 0.05
|
| 473 |
+
reward += 0.02 * test_result.passed # small per-test bonus
|
| 474 |
+
return reward
|
| 475 |
+
|
| 476 |
+
def decay_factor(self, epoch, total_epochs):
|
| 477 |
+
# Fades out at 60% of training β agent transitions to final reward signal
|
| 478 |
+
return max(0.0, 1.0 - (epoch / (total_epochs * 0.6)))
|
| 479 |
+
```
|
| 480 |
+
|
| 481 |
+
These are multiplied by `decay_factor` so early training gets denser signal,
|
| 482 |
+
and late training relies on the real reward. This prevents the agent from
|
| 483 |
+
over-optimizing partial pass rates at the expense of handoff quality.
|
| 484 |
+
|
| 485 |
+
### 8.2 Main Rubric (addresses issues #3, #6, #2, #4)
|
| 486 |
+
|
| 487 |
+
```python
|
| 488 |
+
# server/rewards/rubric.py
|
| 489 |
+
from openenv import Rubric
|
| 490 |
+
|
| 491 |
+
HANDOFF_TOKEN_BUDGET = 300
|
| 492 |
+
|
| 493 |
+
class ContinuityRubric(Rubric):
|
| 494 |
+
|
| 495 |
+
def score(self, visible_results, hidden_results, handoff,
|
| 496 |
+
s2_edit_history, s2_failed_runs, invalid_actions):
|
| 497 |
+
|
| 498 |
+
# Component 1: Test score β visible + hidden weighted
|
| 499 |
+
v_score = visible_results.passed / max(visible_results.total, 1)
|
| 500 |
+
h_score = hidden_results.passed / max(hidden_results.total, 1)
|
| 501 |
+
test_score = 0.6 * v_score + 0.4 * h_score # hidden tests carry real weight
|
| 502 |
+
|
| 503 |
+
# Component 2: Handoff quality (replaces naive token count)
|
| 504 |
+
quality_score = self._handoff_quality(handoff)
|
| 505 |
+
|
| 506 |
+
# Component 3: Linearity (replaces re-read counting β see issue #3)
|
| 507 |
+
linearity_score = self._linearity(s2_edit_history, s2_failed_runs)
|
| 508 |
+
|
| 509 |
+
# Reconstruction penalty (addresses issue #2 shortcut)
|
| 510 |
+
rewrite_penalty = self._rewrite_penalty(s2_edit_history)
|
| 511 |
+
|
| 512 |
+
# Invalid action penalty
|
| 513 |
+
action_penalty = min(invalid_actions * 0.02, 0.1)
|
| 514 |
+
|
| 515 |
+
total = (
|
| 516 |
+
0.55 * test_score
|
| 517 |
+
+ 0.20 * quality_score
|
| 518 |
+
+ 0.15 * linearity_score
|
| 519 |
+
- rewrite_penalty
|
| 520 |
+
- action_penalty
|
| 521 |
+
)
|
| 522 |
+
|
| 523 |
+
return {
|
| 524 |
+
"total": round(max(0.0, total), 4),
|
| 525 |
+
"test_score": test_score,
|
| 526 |
+
"quality_score": quality_score,
|
| 527 |
+
"linearity_score": linearity_score,
|
| 528 |
+
"rewrite_penalty": rewrite_penalty,
|
| 529 |
+
"action_penalty": action_penalty
|
| 530 |
+
}
|
| 531 |
+
|
| 532 |
+
def _handoff_quality(self, handoff):
|
| 533 |
+
# Replaces naive token count β measures structure + density + compression
|
| 534 |
+
if not handoff:
|
| 535 |
+
return 0.0
|
| 536 |
+
score = 0.0
|
| 537 |
+
tokens = handoff.split()
|
| 538 |
+
token_count = len(tokens)
|
| 539 |
+
|
| 540 |
+
# Compression
|
| 541 |
+
if token_count <= HANDOFF_TOKEN_BUDGET:
|
| 542 |
+
score += 0.4
|
| 543 |
+
else:
|
| 544 |
+
overage = token_count - HANDOFF_TOKEN_BUDGET
|
| 545 |
+
score += max(0.0, 0.4 - (overage / HANDOFF_TOKEN_BUDGET) * 0.4)
|
| 546 |
+
|
| 547 |
+
# Structure: reward presence of all required sections
|
| 548 |
+
sections = ["COMPLETED:", "REMAINING:", "KEY FUNCTIONS:", "NEXT STEPS:"]
|
| 549 |
+
score += 0.3 * (sum(1 for s in sections if s in handoff) / len(sections))
|
| 550 |
+
|
| 551 |
+
# Information density: unique word ratio penalizes repetition
|
| 552 |
+
unique_ratio = len(set(tokens)) / max(token_count, 1)
|
| 553 |
+
score += 0.2 * min(unique_ratio * 2, 1.0)
|
| 554 |
+
|
| 555 |
+
# Structural formatting bonus
|
| 556 |
+
has_bullets = any(l.strip().startswith(("-", "*", "1.", "TODO"))
|
| 557 |
+
for l in handoff.split("\n"))
|
| 558 |
+
score += 0.1 if has_bullets else 0.0
|
| 559 |
+
|
| 560 |
+
return round(score, 4)
|
| 561 |
+
|
| 562 |
+
def _linearity(self, edit_history, failed_runs):
|
| 563 |
+
# Track thrashing (reverting writes) and failed test runs
|
| 564 |
+
# Better signal than counting re-reads (addresses issue #3)
|
| 565 |
+
if not edit_history:
|
| 566 |
+
return 0.5
|
| 567 |
+
|
| 568 |
+
thrash_count = sum(
|
| 569 |
+
1 for i in range(1, len(edit_history))
|
| 570 |
+
if edit_history[i]["new"] == edit_history[i-1]["prev"]
|
| 571 |
+
)
|
| 572 |
+
thrash_penalty = min(thrash_count * 0.1, 0.5)
|
| 573 |
+
run_penalty = min(failed_runs * 0.05, 0.3)
|
| 574 |
+
|
| 575 |
+
return round(max(0.0, 1.0 - thrash_penalty - run_penalty), 4)
|
| 576 |
+
|
| 577 |
+
def _rewrite_penalty(self, edit_history):
|
| 578 |
+
# If session 2 wrote large volumes to previously-empty files,
|
| 579 |
+
# it likely reconstructed from pretrained priors, not the handoff
|
| 580 |
+
if not edit_history:
|
| 581 |
+
return 0.0
|
| 582 |
+
total_written = sum(len(e["new"]) for e in edit_history)
|
| 583 |
+
total_previous = sum(len(e["prev"]) for e in edit_history)
|
| 584 |
+
if total_previous == 0 and total_written > 500:
|
| 585 |
+
return 0.15
|
| 586 |
+
return 0.0
|
| 587 |
+
```
|
| 588 |
+
|
| 589 |
+
### 8.3 Why the revised rubric is hard to game
|
| 590 |
+
|
| 591 |
+
| Game attempt | Why it fails |
|
| 592 |
+
|---|---|
|
| 593 |
+
| Dump code into handoff | HandoffValidator rejects code blocks > 5 lines |
|
| 594 |
+
| Write minimal/empty handoff | quality_score = 0, session 2 fails tests |
|
| 595 |
+
| Session 2 rewrites from pretrained priors | rewrite_penalty fires |
|
| 596 |
+
| Thrash writes in session 2 | linearity thrash detection penalizes |
|
| 597 |
+
| Pass visible tests, ignore edge cases | hidden tests weighted 40% of test_score |
|
| 598 |
+
| Rely on consistent tool patterns | name randomization breaks pattern reliance |
|
| 599 |
+
|
| 600 |
+
---
|
| 601 |
+
|
| 602 |
+
## 9. Sandbox [UPDATED β stricter ulimits]
|
| 603 |
+
|
| 604 |
+
```python
|
| 605 |
+
# server/sandbox.py
|
| 606 |
+
import subprocess, tempfile, os, resource
|
| 607 |
+
|
| 608 |
+
class Sandbox:
|
| 609 |
+
def __init__(self, timeout=10):
|
| 610 |
+
self.timeout = timeout
|
| 611 |
+
|
| 612 |
+
def run_tests(self, files, test_code):
|
| 613 |
+
with tempfile.TemporaryDirectory() as tmpdir:
|
| 614 |
+
self._write_files(tmpdir, files, test_code)
|
| 615 |
+
|
| 616 |
+
def set_limits():
|
| 617 |
+
resource.setrlimit(resource.RLIMIT_CPU, (8, 8))
|
| 618 |
+
resource.setrlimit(resource.RLIMIT_AS, (256*1024*1024,)*2) # 256MB RAM
|
| 619 |
+
resource.setrlimit(resource.RLIMIT_NOFILE, (20, 20)) # 20 file handles
|
| 620 |
+
resource.setrlimit(resource.RLIMIT_NPROC, (10, 10)) # no fork bombs
|
| 621 |
+
|
| 622 |
+
try:
|
| 623 |
+
result = subprocess.run(
|
| 624 |
+
["python", "-m", "pytest", "test_solution.py",
|
| 625 |
+
"--tb=short", "-q", "--no-header"],
|
| 626 |
+
capture_output=True, text=True,
|
| 627 |
+
timeout=self.timeout, cwd=tmpdir,
|
| 628 |
+
preexec_fn=set_limits,
|
| 629 |
+
env={"PATH": "/usr/bin:/bin"} # no network access
|
| 630 |
+
)
|
| 631 |
+
return self._parse_result(result.stdout, result.returncode)
|
| 632 |
+
except subprocess.TimeoutExpired:
|
| 633 |
+
return TestResult(passed=0, total=1, compiled=False,
|
| 634 |
+
summary="Timeout β likely infinite loop")
|
| 635 |
+
except Exception as e:
|
| 636 |
+
return TestResult(passed=0, total=1, compiled=False,
|
| 637 |
+
summary=f"Sandbox error: {e}")
|
| 638 |
+
```
|
| 639 |
+
|
| 640 |
+
Note: If on-site infrastructure permits, upgrade to Docker container isolation for
|
| 641 |
+
the full training run. Subprocess + ulimits is sufficient for dev and demo.
|
| 642 |
+
|
| 643 |
+
---
|
| 644 |
+
|
| 645 |
+
## 10. Training Pipeline [UPDATED]
|
| 646 |
+
|
| 647 |
+
### 10.1 Model
|
| 648 |
+
|
| 649 |
+
`unsloth/Qwen2.5-Coder-7B-Instruct` β coding-specialized, fits Colab T4 in 4-bit,
|
| 650 |
+
2x speedup from Unsloth over vanilla HF.
|
| 651 |
+
|
| 652 |
+
### 10.2 Algorithm: GRPO primary, PPO backup (addresses issue #15)
|
| 653 |
+
|
| 654 |
+
GRPO can be unstable with small batches and noisy rewards. Run PPO in parallel as
|
| 655 |
+
a sanity check. If GRPO diverges, PPO gives a usable training curve to show.
|
| 656 |
+
|
| 657 |
+
**Reward normalization β critical:**
|
| 658 |
+
```python
|
| 659 |
+
def normalize_rewards(rewards):
|
| 660 |
+
mean = sum(rewards) / len(rewards)
|
| 661 |
+
std = (sum((r-mean)**2 for r in rewards) / len(rewards)) ** 0.5
|
| 662 |
+
return [(r - mean) / (std + 1e-8) for r in rewards]
|
| 663 |
+
```
|
| 664 |
+
|
| 665 |
+
**GRPO config:**
|
| 666 |
+
```yaml
|
| 667 |
+
num_train_epochs: 6
|
| 668 |
+
per_device_train_batch_size: 2
|
| 669 |
+
gradient_accumulation_steps: 8
|
| 670 |
+
learning_rate: 2e-5
|
| 671 |
+
reward_normalization: true
|
| 672 |
+
clip_range: 0.2
|
| 673 |
+
kl_coeff: 0.05 # prevents reward hacking
|
| 674 |
+
warmup_steps: 50
|
| 675 |
+
```
|
| 676 |
+
|
| 677 |
+
### 10.3 Episode rollout (handles stuck agents and invalid actions)
|
| 678 |
+
|
| 679 |
+
```python
|
| 680 |
+
def rollout(env, agent, epoch, total_epochs):
|
| 681 |
+
obs = env.reset()
|
| 682 |
+
done = False
|
| 683 |
+
trajectory = []
|
| 684 |
+
total_aux = 0.0
|
| 685 |
+
decay = aux_rewarder.decay_factor(epoch, total_epochs)
|
| 686 |
+
|
| 687 |
+
# Session 1
|
| 688 |
+
for _ in range(env.step_limit + 2): # +2 buffer for late handoff warning
|
| 689 |
+
action = agent.act(obs)
|
| 690 |
+
obs, reward, done, info = env.step(action)
|
| 691 |
+
if "auxiliary_reward" in info:
|
| 692 |
+
total_aux += info["auxiliary_reward"] * decay
|
| 693 |
+
trajectory.append((obs, action, reward, info))
|
| 694 |
+
if done or info.get("session") == 2:
|
| 695 |
+
break
|
| 696 |
+
|
| 697 |
+
if env.state()["session"] == 1:
|
| 698 |
+
return trajectory, 0.0 # hit step limit without handoff
|
| 699 |
+
|
| 700 |
+
# Session 2
|
| 701 |
+
s2_obs = {"session": 2, "message": "Call parse_handoff() to retrieve your note."}
|
| 702 |
+
for _ in range(env.step_limit):
|
| 703 |
+
action = agent.act(s2_obs)
|
| 704 |
+
obs, reward, done, info = env.step(action)
|
| 705 |
+
trajectory.append((obs, action, reward, info))
|
| 706 |
+
if done:
|
| 707 |
+
break
|
| 708 |
+
|
| 709 |
+
final_reward = (reward or 0.0) + total_aux
|
| 710 |
+
return trajectory, normalize_reward(final_reward)
|
| 711 |
+
```
|
| 712 |
+
|
| 713 |
+
### 10.4 Curriculum (addresses issue #7)
|
| 714 |
+
|
| 715 |
+
```
|
| 716 |
+
Epochs 1-2: easy tasks only β learn basic handoff structure
|
| 717 |
+
Epochs 3-4: easy + medium β learn compression under step pressure
|
| 718 |
+
Epochs 5-6: medium + hard β learn surgical prioritization
|
| 719 |
+
Eval only: holdout set β generalization check, never in training
|
| 720 |
+
```
|
| 721 |
+
|
| 722 |
+
### 10.5 Colab notebook outline
|
| 723 |
+
|
| 724 |
+
```
|
| 725 |
+
Cell 1: Install: openenv unsloth trl transformers wandb pytest
|
| 726 |
+
Cell 2: Load env from HF Space
|
| 727 |
+
Cell 3: Load Qwen2.5-Coder-7B-Instruct (Unsloth 4-bit)
|
| 728 |
+
Cell 4: Run all 3 baselines β save baseline_results.json
|
| 729 |
+
Cell 5: GRPO training loop with rollout β log to wandb
|
| 730 |
+
Cell 6: Run PPO for comparison
|
| 731 |
+
Cell 7: Eval on holdout set (trained model vs baselines)
|
| 732 |
+
Cell 8: Save all plots as PNG to /plots/
|
| 733 |
+
Cell 9: Ablation runs (3 configs)
|
| 734 |
+
Cell 10: Print epoch 1 vs epoch 20 handoff notes side by side
|
| 735 |
+
```
|
| 736 |
+
|
| 737 |
+
---
|
| 738 |
+
|
| 739 |
+
## 11. Baselines [NEW β addresses issue #12]
|
| 740 |
+
|
| 741 |
+
All four on the same plot. Without this, reward improvement is meaningless.
|
| 742 |
+
|
| 743 |
+
| Baseline | Description | Expected S2 pass rate |
|
| 744 |
+
|---|---|---|
|
| 745 |
+
| No handoff | Session 2 starts with blank note | ~5-10% |
|
| 746 |
+
| Random handoff | Gibberish as the handoff note | ~8-12% |
|
| 747 |
+
| **Trained agent (ours)** | Our GRPO-trained model | Target: >60% |
|
| 748 |
+
| Full S1 transcript | Upper bound β all context given | ~75-85% |
|
| 749 |
+
|
| 750 |
+
The trained agent should be comfortably above random and approaching (not matching)
|
| 751 |
+
the full transcript upper bound. That gap tells the story clearly.
|
| 752 |
+
|
| 753 |
+
---
|
| 754 |
+
|
| 755 |
+
## 12. Ablation Studies [NEW β addresses issue #17]
|
| 756 |
+
|
| 757 |
+
Three ablations to justify each reward component to judges:
|
| 758 |
+
|
| 759 |
+
| Ablation | Removed component | Expected degradation |
|
| 760 |
+
|---|---|---|
|
| 761 |
+
| No compression reward | quality_score = 0 | Handoffs become bloated |
|
| 762 |
+
| No linearity reward | linearity_score = 0 | Session 2 thrashes more |
|
| 763 |
+
| No auxiliary S1 reward | AuxiliaryRewarder disabled | Slower convergence |
|
| 764 |
+
|
| 765 |
+
Plot all ablations vs full model on same axes in `plots/ablation_comparison.png`.
|
| 766 |
+
One-line caption per plot. Axes labeled: "Training Episode" (x) / "Total Reward" (y).
|
| 767 |
+
|
| 768 |
+
---
|
| 769 |
+
|
| 770 |
+
## 13. Evaluation Reporting [NEW β addresses issue #8]
|
| 771 |
+
|
| 772 |
+
Don't aggregate across difficulties β it hides where the agent struggles.
|
| 773 |
+
|
| 774 |
+
Report separately per difficulty and across seeds:
|
| 775 |
+
|
| 776 |
+
```
|
| 777 |
+
easy tasks: pass rate | avg handoff tokens | avg S2 steps
|
| 778 |
+
medium tasks: same
|
| 779 |
+
hard tasks: same
|
| 780 |
+
holdout tasks: same β generalization signal
|
| 781 |
+
|
| 782 |
+
Run 3 seeds minimum. Report mean Β± std.
|
| 783 |
+
```
|
| 784 |
+
|
| 785 |
+
---
|
| 786 |
+
|
| 787 |
+
## 14. Interpretability [NEW β addresses issue #16]
|
| 788 |
+
|
| 789 |
+
Show *what the agent learned to keep vs drop* across training epochs.
|
| 790 |
+
|
| 791 |
+
```python
|
| 792 |
+
# Track which handoff sections grow or shrink over training
|
| 793 |
+
def analyze_handoff_evolution(handoff_log):
|
| 794 |
+
section_lengths = {}
|
| 795 |
+
for epoch, handoffs in handoff_log.items():
|
| 796 |
+
section_lengths[epoch] = {}
|
| 797 |
+
for section in ["COMPLETED:", "REMAINING:", "KEY FUNCTIONS:", "NEXT STEPS:"]:
|
| 798 |
+
lengths = [len(extract_section(h, section)) for h in handoffs]
|
| 799 |
+
section_lengths[epoch][section] = sum(lengths) / len(lengths)
|
| 800 |
+
return section_lengths
|
| 801 |
+
```
|
| 802 |
+
|
| 803 |
+
Plot as stacked bar chart (`plots/handoff_diff_over_epochs.png`).
|
| 804 |
+
|
| 805 |
+
Expected learning signal visible in the chart:
|
| 806 |
+
- COMPLETED section shrinks (agent stops over-documenting finished work)
|
| 807 |
+
- REMAINING section gets more precise (specific function names, not vague prose)
|
| 808 |
+
- NEXT STEPS section grows and becomes the highest-value section for session 2
|
| 809 |
+
|
| 810 |
+
This is the interpretability story for the blog and pitch.
|
| 811 |
+
|
| 812 |
+
---
|
| 813 |
+
|
| 814 |
+
## 15. Agent Loop (Client) [UPDATED β addresses issue #13]
|
| 815 |
+
|
| 816 |
+
```python
|
| 817 |
+
# client/agent.py β no server imports
|
| 818 |
+
|
| 819 |
+
S1_SYSTEM_PROMPT = """You are working on a coding task in Session 1.
|
| 820 |
+
Complete as much as possible. When approaching your step limit, call write_handoff()
|
| 821 |
+
with a structured note following this format:
|
| 822 |
+
TASK: / COMPLETED: / REMAINING: / KEY FUNCTIONS: / EDGE CASES: / NEXT STEPS:
|
| 823 |
+
You have a retry budget for invalid actions. Use it wisely."""
|
| 824 |
+
|
| 825 |
+
S2_SYSTEM_PROMPT = """You are in Session 2. You have NO memory of Session 1.
|
| 826 |
+
Your ONLY information is the handoff note. Start by calling parse_handoff(),
|
| 827 |
+
then use the note to continue the task. Do not rewrite everything from scratch."""
|
| 828 |
+
|
| 829 |
+
class Agent:
|
| 830 |
+
def __init__(self, model, tokenizer, retry_budget=3):
|
| 831 |
+
self.model = model
|
| 832 |
+
self.tokenizer = tokenizer
|
| 833 |
+
self.retry_budget = retry_budget
|
| 834 |
+
self.context = []
|
| 835 |
+
|
| 836 |
+
def act(self, obs):
|
| 837 |
+
prompt = self._build_prompt(obs)
|
| 838 |
+
for attempt in range(self.retry_budget):
|
| 839 |
+
response = self._generate(prompt)
|
| 840 |
+
action = self._parse_action(response)
|
| 841 |
+
if action is not None:
|
| 842 |
+
self.context.append({"obs": obs, "action": action})
|
| 843 |
+
return action
|
| 844 |
+
prompt = self._build_retry_prompt(prompt, response, attempt)
|
| 845 |
+
return Action(tool="noop", content="") # graceful no-op on exhaustion
|
| 846 |
+
|
| 847 |
+
def _build_prompt(self, obs):
|
| 848 |
+
system = S1_SYSTEM_PROMPT if obs.get("session") == 1 else S2_SYSTEM_PROMPT
|
| 849 |
+
return system + "\n\n" + format_obs(obs)
|
| 850 |
+
```
|
| 851 |
+
|
| 852 |
+
---
|
| 853 |
+
|
| 854 |
+
## 16. Risk Register [UPDATED β full 20-issue resolution]
|
| 855 |
+
|
| 856 |
+
| # | Issue | Severity | Status | Resolution |
|
| 857 |
+
|---|---|---|---|---|
|
| 858 |
+
| 1 | Credit assignment β S1 no direct reward | HIGH | FIXED | Auxiliary shaped rewards + decay schedule |
|
| 859 |
+
| 2 | Handoff gaming β code dumps / hinting | HIGH | FIXED | HandoffValidator + code block limit + rewrite penalty |
|
| 860 |
+
| 3 | Linearity metric weak (re-read counting) | MEDIUM | FIXED | Thrash detection on edit history + failed run rate |
|
| 861 |
+
| 4 | Test suite exploitable | MEDIUM | FIXED | Hidden adversarial tests at submit |
|
| 862 |
+
| 5 | Session separation weak | MEDIUM | FIXED | Name randomization per episode seed |
|
| 863 |
+
| 6 | Compression metric naive | MEDIUM | FIXED | Multi-factor quality score: structure + density + ratio |
|
| 864 |
+
| 7 | Task difficulty miscalibrated | MEDIUM | FIXED | Step limits verified empirically, handoff-critical design |
|
| 865 |
+
| 8 | Evaluation hides per-difficulty gaps | MEDIUM | FIXED | Separate easy/medium/hard/holdout reporting |
|
| 866 |
+
| 9 | Sandbox not fully isolated | MEDIUM | FIXED | Strict ulimits: CPU, RAM, file handles, forks |
|
| 867 |
+
| 10 | Step limit too tight or too loose | LOW | FIXED | Dynamic by difficulty, late-handoff warning |
|
| 868 |
+
| 11 | Template overfitting | MEDIUM | FIXED | Name randomization + holdout eval set |
|
| 869 |
+
| 12 | No baselines | HIGH | FIXED | 3 baselines + upper bound, all on same plot |
|
| 870 |
+
| 13 | Agent gets stuck / invalid actions | LOW | FIXED | Retry budget, invalid action penalty, noop fallback |
|
| 871 |
+
| 14 | Tool pattern exploitation | LOW | ACCEPTED | Name randomization covers most of this; minor risk |
|
| 872 |
+
| 15 | GRPO instability | MEDIUM | FIXED | Reward normalization, KL coeff, PPO backup |
|
| 873 |
+
| 16 | No interpretability | MEDIUM | FIXED | Handoff section evolution tracking + diff plot |
|
| 874 |
+
| 17 | No ablation studies | MEDIUM | FIXED | 3 ablations with plots |
|
| 875 |
+
| 18 | Demo risk | LOW | FIXED | Deterministic seeds, pre-recorded run URL |
|
| 876 |
+
| 19 | Handoff format inconsistent | HIGH | FIXED | Mandatory 6-section structure enforced by validator |
|
| 877 |
+
| 20 | Tests don't capture understanding | LOW | PARTIALLY | Hidden adversarial tests cover this adequately for hackathon scope |
|
| 878 |
+
|
| 879 |
+
**Issue #14 accepted as low-risk** β name randomization already breaks most pattern
|
| 880 |
+
exploitation. Full tool response variation adds complexity with marginal gain.
|
| 881 |
+
|
| 882 |
+
**Issue #20 partial** β mutation testing is a research-grade addition, out of scope
|
| 883 |
+
for the hackathon timeline.
|
| 884 |
+
|
| 885 |
+
---
|
| 886 |
+
|
| 887 |
+
## 17. Demo Preparation [NEW β addresses issue #18]
|
| 888 |
+
|
| 889 |
+
- **Deterministic seed**: `env.reset(seed=42)` β same task, same names, reproducible
|
| 890 |
+
- **Pre-recorded run**: screen recording of a successful trained-agent episode, hosted
|
| 891 |
+
as URL (not committed to repo). Linked from README.
|
| 892 |
+
- **Fallback slide**: screenshot of epoch 1 vs epoch 20 handoff side by side β shows
|
| 893 |
+
the learning visually to a non-technical audience
|
| 894 |
+
|
| 895 |
+
**Never end the live demo on `submit()`** β too unpredictable. End on the handoff note
|
| 896 |
+
being written and displayed. That's the visual payoff.
|
| 897 |
+
|
| 898 |
+
---
|
| 899 |
+
|
| 900 |
+
## 18. Submission Checklist [UPDATED]
|
| 901 |
+
|
| 902 |
+
| Requirement | How satisfied | Status |
|
| 903 |
+
|---|---|---|
|
| 904 |
+
| OpenEnv latest release | `MCPEnvironment` subclass, `openenv.yaml`, pinned version in requirements.txt | [ ] |
|
| 905 |
+
| Training script (Unsloth/TRL) | `training/train_grpo.ipynb` β Colab T4, re-runnable in <30 min | [ ] |
|
| 906 |
+
| Training evidence | `plots/` β reward, length, 4-way baseline, ablations, interpretability β all PNG | [ ] |
|
| 907 |
+
| Mini blog OR video | HF blog post + <2 min YouTube video | [ ] |
|
| 908 |
+
| HF Space | `yourteam/cross-session-continuity-env` β live and runnable | [ ] |
|
| 909 |
+
| README with all links | Space, notebook, blog, video, WandB run | [ ] |
|
| 910 |
+
| No large files in repo | Videos as `.url` text files only | [ ] |
|
| 911 |
+
| Baselines | 3 baselines + upper bound documented and plotted | [ ] |
|
| 912 |
+
| Ablations | 3 ablations documented and plotted | [ ] |
|
| 913 |
+
| Holdout eval | Generalization results on 10 unseen tasks | [ ] |
|
| 914 |
+
| Per-difficulty breakdown | easy / medium / hard results reported separately | [ ] |
|
| 915 |
+
|
| 916 |
+
---
|
| 917 |
+
|
| 918 |
+
## 19. README Template [UPDATED]
|
| 919 |
+
|
| 920 |
+
```markdown
|
| 921 |
+
# Cross-Session Continuity Env
|
| 922 |
+
|
| 923 |
+
> Can RL teach an LLM to write better notes to its future self?
|
| 924 |
+
|
| 925 |
+
## Problem
|
| 926 |
+
LLMs forget everything when a session ends. For long coding tasks that span
|
| 927 |
+
multiple sessions this is critical. No existing RL environment trains for this.
|
| 928 |
+
|
| 929 |
+
## How It Works
|
| 930 |
+
[diagram: session1 β handoff.md β session2 β reward]
|
| 931 |
+
|
| 932 |
+
Session 1: agent gets task + starter code. Works until step limit.
|
| 933 |
+
Must write a structured 6-section handoff note before session ends.
|
| 934 |
+
|
| 935 |
+
Session 2: starts completely cold. Only the handoff note exists.
|
| 936 |
+
Must complete the task and pass tests.
|
| 937 |
+
|
| 938 |
+
Reward = test correctness (visible + hidden) + handoff quality + session 2 linearity.
|
| 939 |
+
|
| 940 |
+
## Reward Breakdown
|
| 941 |
+
| Component | Weight | What it measures |
|
| 942 |
+
|-------------------|--------|-------------------------------------|
|
| 943 |
+
| Tests (visible) | 33% | Session 2 correctness |
|
| 944 |
+
| Tests (hidden) | 22% | Generalization, no test overfitting |
|
| 945 |
+
| Handoff quality | 20% | Structure, density, compression |
|
| 946 |
+
| Linearity | 15% | Session 2 didn't thrash |
|
| 947 |
+
| Penalties | 10% | Invalid actions, reconstruction |
|
| 948 |
+
|
| 949 |
+
## Results
|
| 950 |
+
| Agent | S2 Test Pass Rate |
|
| 951 |
+
|------------------------|-------------------|
|
| 952 |
+
| No handoff (baseline) | ~8% |
|
| 953 |
+
| Random handoff | ~11% |
|
| 954 |
+
| Trained (ours) | ~65% |
|
| 955 |
+
| Full transcript (UB) | ~80% |
|
| 956 |
+
|
| 957 |
+

|
| 958 |
+
*Total reward over training episodes β all baselines on same axes*
|
| 959 |
+
|
| 960 |
+

|
| 961 |
+
*Each reward component contribution β ablation study*
|
| 962 |
+
|
| 963 |
+

|
| 964 |
+
*What the agent learned to keep vs drop over training*
|
| 965 |
+
|
| 966 |
+
## Before / After
|
| 967 |
+
**Epoch 1:** 900 tokens, rambling, full code blocks, no structure
|
| 968 |
+
**Epoch 20:** 180 tokens, 6 clear sections, precise function names, zero code
|
| 969 |
+
|
| 970 |
+
## Links
|
| 971 |
+
- HF Space: [url]
|
| 972 |
+
- Colab Notebook: [url]
|
| 973 |
+
- HF Blog Post: [url]
|
| 974 |
+
- YouTube Demo (<2 min): [url]
|
| 975 |
+
- WandB Training Run: [url]
|
| 976 |
+
```
|
| 977 |
+
|
| 978 |
+
---
|
| 979 |
+
|
| 980 |
+
## 20. Pitch Story [UPDATED]
|
| 981 |
+
|
| 982 |
+
> "Every developer has hit this wall. You're deep into a coding task with an AI
|
| 983 |
+
> assistant. The session ends. You come back the next day β and the AI remembers
|
| 984 |
+
> nothing. You start over from scratch.
|
| 985 |
+
>
|
| 986 |
+
> We asked a different question: what if we trained the AI to leave a perfect
|
| 987 |
+
> briefing for its future self?
|
| 988 |
+
>
|
| 989 |
+
> Cross-Session Continuity Env is an RL environment where an agent must complete
|
| 990 |
+
> a coding task split across two sessions with zero shared memory. Session 1
|
| 991 |
+
> works on the problem, then writes a structured handoff note. Session 2 starts
|
| 992 |
+
> completely cold β only that note exists.
|
| 993 |
+
>
|
| 994 |
+
> The agent is rewarded not for session 1 performance, but for how well its
|
| 995 |
+
> future self performs using only the note it left behind.
|
| 996 |
+
>
|
| 997 |
+
> After training, the agent learned something we didn't expect. It stopped writing
|
| 998 |
+
> long rambling summaries. It started writing surgical briefings β 180 words,
|
| 999 |
+
> six sections, exactly what session 2 needs and nothing it doesn't.
|
| 1000 |
+
>
|
| 1001 |
+
> Test pass rates went from 8% (no handoff at all) to 65%.
|
| 1002 |
+
>
|
| 1003 |
+
> No one has trained this behavior explicitly before. We think it matters."
|
| 1004 |
+
|
| 1005 |
+
---
|
| 1006 |
+
|
| 1007 |
+
## 21. Timeline [UPDATED]
|
| 1008 |
+
|
| 1009 |
+
| Day | Task | Risk & Contingency |
|
| 1010 |
+
|---|---|---|
|
| 1011 |
+
| Day 1 (pre-onsite) | Task bank: 20 tasks + holdout set. Sandbox + ulimits tested. HandoffValidator working. | Sandbox is highest-risk β do first. Fallback: relax ulimits if resource module unavailable |
|
| 1012 |
+
| Day 2 (pre-onsite) | Env class, session manager, rubric, auxiliary rewarder. Full unit tests on each. | Rubric edge cases β budget 2h for test coverage |
|
| 1013 |
+
| Day 3 (pre-onsite) | End-to-end episode: agent completes 2-session run. Client/server separation verified. | Integration bugs β if stuck, simplify tool set |
|
| 1014 |
+
| Day 4 (onsite 25th) | Colab notebook. All 3 baseline runs. First GRPO curves. WandB connected. | Compute time β run baselines overnight if needed |
|
| 1015 |
+
| Day 5 (onsite 26th am) | Full training run on HF credits. Ablations. Plots committed. | GRPO divergence β fall back to PPO results |
|
| 1016 |
+
| Day 5 (onsite 26th pm) | HF Space live. README + blog done. Demo recorded. Final checklist. | Deployment issues β test HF Space access 24h early |
|
| 1017 |
+
|
| 1018 |
+
---
|
| 1019 |
+
|
| 1020 |
+
## 22. What Good Looks Like at Submission
|
| 1021 |
+
|
| 1022 |
+
1. Judge visits HF Space β watches a live 2-session run with trained agent
|
| 1023 |
+
2. Reward curve shows clear upward trend with all 4 baselines on the same plot
|
| 1024 |
+
3. Ablation plot shows each component contributes something measurable
|
| 1025 |
+
4. Epoch 1 vs epoch 20 handoff note is visibly, strikingly different
|
| 1026 |
+
5. Per-difficulty breakdown shows where the agent is strong vs weak
|
| 1027 |
+
6. Colab notebook re-runs in under 30 minutes on a T4
|
| 1028 |
+
7. Holdout eval confirms generalization, not just memorization
|
| 1029 |
+
|
| 1030 |
+
All seven = strong submission that covers every judging criterion.
|