25.4. Refactoring Research Code into Packages#
Research code often starts as a few scripts or a notebook that grew. A package, with modules, tests, and packaging metadata, is easier to reuse and to trust, and an agent can do much of the mechanical work. The risk is a refactor that quietly changes a result, so this method keeps behavior fixed: pin what the code does now with tests, restructure in small steps you review, then package and automate. For the underlying practices, see Software Design Principles, Testing and Continuous Integration, and Package Development.
flowchart LR
A["Pin current behavior<br/>with regression tests"] --> B["Restructure<br/>one small step"]
B --> C{"Tests pass?"}
C -->|yes| D["Review the diff<br/>and commit"]
D -->|next step| B
C -->|no| E["Revert or fix<br/>the code, not the tests"]
E --> B
D -->|done| F["Package and<br/>add CI"]
classDef s fill:#14154C,color:#ffffff,stroke:#3D3E82;
classDef v fill:#A51C30,color:#ffffff,stroke:#A51C30;
class A,B,D,F s;
class C,E v;
25.4.1. Before you start#
Work on a branch with a clean git tree. Every agent change then shows as a diff, and
git restoreorgit revertundoes it. Commit before each step; see Git as the undo layer.Record the environment. Pin the package versions the current code runs with, so you compare like with like. See Environment Reproducibility.
Pick a small, representative input. Choose data that exercises the main code paths and runs in seconds or minutes. If it needs a GPU, use a short interactive job, not a login node; see Using Agentic AI on the Cluster.
Tell the agent the rules. Put the test command and the constraints in your project instructions, for example “run
pytestafter every change”.
25.4.2. Step 1: Pin current behavior with tests#
Before anything moves, have the agent write regression tests that record what the code produces now. They are the safety net for every later step, so write them first and leave the code untouched.
Before changing any code, write pytest regression tests that run
scripts/preprocess.pyandscripts/train_small.pyontests/data/sample.csvwith seed 0, save their outputs as reference files undertests/references/, and compare future outputs against those files with a relative tolerance of 1e-6. Do not modify the scripts.
Review the tests before trusting them:
Do they exercise the real code paths? A test that only checks a file exists proves nothing.
Do they fail when they should? Change a constant in the code and confirm a test fails, then undo the change. A test that never fails is not protecting you.
Is randomness controlled? Seed every random number generator the code uses, including PyTorch’s; see Randomness and Seeds.
Is the tolerance deliberate? Exact equality is fragile for floating-point results; a tolerance that is too loose hides real changes. See the regression tests in Types of Tests.
Commit the tests and reference files on their own, before any refactoring, so later diffs show if anything touches them. Then add “never edit files under tests/references/” to your project instructions, and back it with a deny rule, Edit(./tests/references/**), since an instruction alone does not stop the agent. The deny rule blocks the agent’s own edits, though not a test or script that rewrites the files; see Permission rules.
25.4.3. Step 2: Restructure in small steps#
Aim for the layout in Package Structure and Layout (Python Example): a src/ package with a pyproject.toml. Ask for one change at a time, and after each one have the agent run the tests and show you the diff.
Move the data-cleaning code from
notebooks/explore.ipynbintosrc/mypkg/cleaning.pyas functions with docstrings. Change only that, and keep the notebook working by importing the new functions. Runpytest, then show me the diff and the test results before committing.
Typical steps, each its own commit:
Move code out of notebooks and scripts into modules, without changing what it does.
Separate computation from file input and output, so the logic becomes plain functions that are easy to test. Software Design Principles shows this pattern.
Replace hard-coded paths and constants with arguments or a configuration file.
Remove duplicated code once the tests cover both copies.
Turn the old scripts into thin command-line entry points that call the package.
Warning
Watch the diff for changes you did not ask for. Agents tend to “improve” code beyond the request: they change a data type, reorder floating-point operations, drop a code path that looked unused, or switch a random seed, and any of these can shift results. Never accept a refactor that edits the reference files or loosens a tolerance to make tests pass. If a test fails, the code goes back, not the test.
25.4.4. Step 3: Package, test, and automate#
Once the code lives in the package and the regression tests still pass:
Packaging. Have the agent complete
pyproject.toml(dependencies, Python version, entry points) and install the package in editable mode. See Tools for Building Packages.Unit tests. Add focused tests for the new functions alongside the regression tests, and check coverage to find untested code. See Testing and Continuous Integration.
Continuous integration. Add a workflow that runs the tests on every push, so later changes, by you or an agent, are checked automatically.
A final full run. Run the complete pipeline once more on the cluster with the packaged code and compare against the old results.
25.4.5. Before you trust the refactor#
The tests pass, and the reference files have not changed since you committed them in step 1.
git log -- tests/referencesshould show only that commit.Each commit contains only the change you asked for.
A representative run with the packaged code matches the old results within your tolerance.
Someone else can install the package in a fresh environment and reproduce the results; see Reproducible Research.
See also
For general checks on agent output, see Evaluating and Monitoring Agents. To understand a codebase before you change it, see Working with Unfamiliar Research Codebases.