Forge Fate
AI

How Do I Evaluate Code I Didn’t Write?

7 min read

In the previous installment, we set a goal for improving a list endpoint while preserving its sort order, pagination behavior, and response contract. Now suppose Codex reports that the change is complete. The useful question is less “Does the explanation sound plausible?” and more “What evidence shows that the behavior we agreed to preserve is still there?” An OpenAI account of Codex code review describes how a change that looks reasonable can still miss an existing API contract or a name that callers depend on.[S1]

The list endpoint in this article is an illustrative example. No repository diff or test run is provided. The page results below are expected values worked out from the stated contract. They are not records of what Codex returned or of passing tests.

Read the completion report, then find the diff that touches the contract

Suppose a completion report says, “I cleaned up the query logic and preserved the existing behavior.” That tells us where to look, but it does not establish that the behavior was preserved. If the explanation adds that sorting moved into the data access layer, it raises a specific question: Does the new sort expression still break ties the same way?

Each piece of evidence answers a different question. The completion report states what the author says was done; the explanation gives reasons for the choices; and the diff shows what code changed. If tests were run, their results show whether the asserted expectations held for those inputs within that run’s scope. Whether the expectations themselves are correct is a separate question.[S3]

Start with the places where the diff meets the contract. Did the sort expression change? Did the condition that determines where the next page starts change? Did the code that assembles the response move? For details such as response fields or error behavior whose contract has not yet been established, check callers and the existing specification before concluding that they were preserved. Without an actual diff, we cannot diagnose a defect in this example. We can still say precisely what evidence a particular change would require.

Work out the expected pages from five items

Here are the assumptions for this illustrative example. The data stays fixed during the queries. Results are ordered by created_at descending, then by unique id ascending when timestamps match. The page size is 2, and the first page is numbered 1. We compare only the list of item IDs in the response. These assumptions do not define other response fields, pagination metadata, errors, or permission behavior.

idcreated_at
1110:00
1209:00
1309:00
1408:00
1507:00

Under that contract, the order is 11 → 12 → 13 → 14 → 15. The following are expected values by design, worked out by hand. They are not API responses or test results.

RequestExpected item IDs
Page 1[11, 12]
Page 2[13, 14]
Last page, page 3[15]

Items 12 and 13 have the same timestamp and fall on opposite sides of a page boundary. If id disappears from the sort expression or its direction is reversed, we need to check whether that boundary still produces the expected result. Adding a tie-breaking column to make the order unambiguous is useful for this reason.[S2] If the diff changes both sorting and the condition for starting a page, read both to see whether they use the same ordering rule.

This small input lets us reason about the first and last pages, a tie across a boundary, and duplicates or omissions between pages. If we also check an empty list, the expected item ID list is []; the shape of the full empty response must come from a separately confirmed contract. Nor does a result calculated from fixed data guarantee behavior when data changes between requests. Pages fetched at different times can produce different results, which needs separate consideration.[S2]

If tests passed, read what they actually checked

Even if a test resembles the table above, a name such as pagination_works and a passing status are not enough. Check whether its input really contains all five items, which contract supplied its expected values, and whether it asserts the IDs and order on each page. Then check which tests ran, against which version of the change, and with what results. This article has no test code or run log, so we cannot say whether the example passes.

A useful question is: “Would this test fail if the sort direction were reversed or the tie-breaking key were removed?” If it appears that it would not, the test might have no tied values, examine only the first page, or assert only the item count. This is a way to read whether a test targets the risk in the change; it is not a claim that we changed the code and ran an experiment. Coverage figures alone say little about whether tests catch defects, and research into whether tests detect small, deliberately introduced faults examines that gap.[S4]

The expected values need review too. If the implementation sorts by id descending and the test is written to expect that same order, both could agree and pass while violating the contract established above. Research on the test oracle problem treats the distinction between observed output and desired output as a problem in its own right.[S3] Compare AI-suggested assertions with the contract and the results worked out by hand. A preliminary study of AI-generated test assertions from requirements also compared requirements-derived criteria with implementation behavior. Its limited findings should not be generalized into a verdict on all AI-generated tests.[S5]

Understanding the code becomes clear when the next change arrives

Alongside evidence that the behavior works, consider whether you can continue changing the code when a new requirement arrives. The standard I suggest is not memorizing every line. It is being able to explain why the key choices were made, where those choices are implemented, and which tests a change would affect. A developer’s account of generated code observes that a team may pass basic tests without understanding the concepts behind the code’s structure; that concern motivates the question here.[S6]

For a list endpoint, follow a caller’s page request through the main query path, data access, and response assembly. Where is the ordering by created_at and id defined? Does the page-start condition use the same rule? Did moving query code change where errors are handled or permissions are checked? If response assembly changed, which callers depend on those fields? Follow the boundaries touched by this diff, without assuming you understand paths you have not read. A case study describes code review as an activity that also builds understanding of a change’s context, though this reading order need not be a rule for every team.[S7]

Suppose the next request is, “When timestamps tie, show the item with the larger id first.” Which sort expression and page-start condition would need to change together? Which expected values in the boundary example would you recalculate, and which tests would you update? If you can answer, you understand how this code connects to the next change. If you cannot, ask AI to explain the path, analyze the impact, suggest counterexamples, or propose tests. Compare its answers with the diff, the agreed contract, and any actual run results. Agreement from another model or a fresh session does not, by itself, constitute an independent execution check.

Accept, request more evidence, or hold based on what is missing

Match the depth of review to the impact of the change. For a small edit to one sort expression, it may be enough to focus on that expression and the page-boundary tests. If the change also reaches a response contract shared by several callers or a permission boundary, trace more of the call path and affected behavior. The official code review example gives a reason to pay attention to changes where contracts can be easy to miss.[S1]

In this example, acceptance is reasonable when you have read the diff where it touches the contract and have enough evidence to compare contract-derived expectations with the relevant tests’ inputs, assertions, and run scope. Request more evidence when you can identify a specific gap. If the test contains tied values but does not assert the second page’s IDs, ask for that assertion. If test code exists but there is no run record, ask what ran and what happened. Narrow or hold the change when it extends into a public response or permission boundary and you cannot identify which code owns that behavior. Neither a test count nor a coverage figure should decide among these outcomes automatically.[S3][S4]

Take one recent change and identify three things: the contract it was meant to preserve, the actual evidence that checks that contract, and the code and tests you would change for the next request. Where an answer is missing, choose one gap and fill it. In the next installment, we’ll consider how much to invest in learning the tools and preparing the working environment so you can keep making these judgments.

Sources

Report an error or share feedback

Open a draft with this article’s title and URL. Review the message and recipient before sending.

To: [email protected]

Open email draft

If no email app opens, copy these details into your usual email service.

Contact information