Skip to content

Repository files navigation

Assignment 9 — Make your agent prove it

Tests

In Assignment 8, you used existing code rather than asking the agent to rewrite it. Here the engineering work is small — finite-difference derivatives and their convergence rates — and the agent exercise is teaching your instructions to demand numerical evidence instead of a remembered answer.

Learning objectives and scope

Use one coding agent throughout. You review its code and results. Implement finite-difference first derivatives, measure their convergence, and replace an instruction that accepts fixed answers with one that requires calculations showing the code works. This is not yet a pressure solver or reservoir model.

1. The problem with the inherited rule

The AGENTS.md that ships with this repository contains a validation rule of the form

a scheme is correct when compute_forward_difference_convergence_rate returns 1 and compute_central_difference_convergence_rate returns 2.

Read it closely. It names two numbers, not a test. Nothing about it requires a study, and it says nothing about how you would know the numbers are right. Ask your agent to explain why that rule cannot distinguish a validated implementation from one that returns the constants.

Then ask the agent to propose a short replacement rule that says what calculations are needed to check the code: compare numerical and exact derivatives, measure how the errors change as the grid gets finer, and compare the measured convergence rate with the expected rate. The rule should still work when the function or grid changes. Review the proposed changes and explicitly approve the wording before the agent edits AGENTS.md.

Write the replacement rule in your own words, including these terms:

  • Convergence: checking whether error decreases as the grid gets finer.
  • Resolutions: the numbers of grid points used; use at least four.
  • Slope: the slope of a plot of logarithmic error against logarithmic grid-point count.
  • Order: the expected rate at which error decreases for the chosen method.
  • Tolerance: how close the measured rate must be to the expected rate.
  • Evidence: the computed errors and rates that support your conclusion.

Do not use a fixed return value as proof that the code works. The provided test checks for these terms and rejects the original rule; it does not assess the quality of your explanation.

2. Engineering task

Implement in assignment9.py:

  • FiniteDifference(function, x_range, N) — a uniform grid over x_range = (x_min, x_max) with N points, storing x, delta_x and the sampled y = function(x).
    • forward() returns the O$(\Delta x)$ forward difference at x[:-1] (N - 1 values).
    • central() returns the O$(\Delta x^2)$ central difference at x[1:-1] (N - 2 values).
  • compute_forward_difference_convergence_rate() and compute_central_difference_convergence_rate() — each returns the observed convergence rate as a positive float.
  • convergence_study(path='convergence.json') — runs both studies, writes the results in JSON format to path and returns it. convergence.json is Git-ignored and is not a submitted deliverable.

Convergence rate

The forward difference approximates $\partial u/\partial x$ with error $\mathcal{O}(|\Delta x|)$; the central difference with $\mathcal{O}(|\Delta x|^2)$. The convergence rate is the slope of the $l_2$ error against $\Delta x$ on a plot with logarithmic scales on both axes (equivalently against $N$, since $\Delta x \propto 1/N$):

$$ \left| u^{\text{numerical}} - u^{\text{exact}} \right|{l_2} = \sqrt{ \frac{1}{M} \sum{i=1}^{M} \left( u^{\text{numerical}}(x_i)

  • u^{\text{exact}}(x_i) \right)^2 }. $$

Here $M$ is the number of evaluated derivative points: $N-1$ for forward and $N-2$ for central. Compare against the exact derivative at those same points.

A rate near $1$ is first order and a rate near $2$ is second order; the plot looks like images/convergence.png.

Use $f(x) = x^4$ on $(0, 10)$, the resolutions $(10, 100, 1000, 10000)$, the exact derivative $f'(x) = 4x^3$, and a np.polyfit of $\log_{10}(\text{error})$ against $\log_{10}(N)$ for the slope. Inspect its sign: decreasing error with increasing $N$ gives a negative slope. Report its absolute value for the required interface, but do not interpret increasing errors as convergence. Fitting against grid spacing instead gives positive order; since spacing is proportional to $1/(N-1)$, the two fits give increasingly similar rates as the grids get finer. Keep the prescribed $N$ fit for grading.

3. Single-agent scientific review

Use the same agent for these steps, with you reviewing its conclusions:

  1. Review and approve the validation plan and instruction repair before edits.
  2. Inspect computed errors at all four resolutions, their reduction, the signed slopes, and agreement with theoretical order. Choose and justify a tolerance; exact equality with the theoretical orders is not required.
  3. Repeat the study locally for $f(x)=\sin(x)$ on $(0,2\pi)$ with exact derivative $\cos(x)$ and at least four resolutions. Check each scheme at its own evaluation points. Use your existing class for this check; no additional function is required.
  4. In a temporary calculation, multiply a correct derivative approximation by two. Measure its errors and explain why your validation rejects it. Restore correct code before submission; never edit protected tests to introduce the fault.
  5. Ask your agent: "How could incorrect code pass these checks?" Review its answer. Require actual errors, error reduction, and an order comparison, not a claimed rate alone.

Include a brief explanation in a docstring or comments in assignment9.py: the computed errors and rates, why you chose your tolerance, the results for the second function, and why the faulty calculation was rejected. No copy of your agent conversation or additional file is required. The numerical evidence, not the agent's confidence, supports completion.

Test, review, submit

env -u SUBMISSION_VALIDATION python -m unittest -v
git diff --check

Initially the class and rate checks fail because the required functions are not yet implemented, and the instruction check fails until the approved repair is in place. Do not weaken tests to change this starting behavior. Passing tests support the review but do not replace your responsibility to inspect the convergence evidence and approve instruction changes.

The two student-editable deliverables are:

  • AGENTS.md (approved, evidence-based validation rule)
  • assignment9.py (class, both rate functions and the study writer)

Everything else is protected, including the tests, images/, environment.yml, submission-policy.json, the workflows and the devcontainer files. No copy of your agent conversation or separate report is required; convergence.json stores your results locally. When ready, ask the agent to submit assignment 9. Review its dry-run result, then explicitly authorize execution. The goal is to write instructions that ask for evidence instead of encoding an answer.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages