Generated code and generated tests can agree on the wrong answer

A passing test suite can repeat the assumptions that produced an implementation. Evaluate AI-assisted work against explicit requirements, independent examples, and the consequences of being wrong.

An assistant writes a function from an ambiguous description. It also writes tests. The tests pass, the implementation looks tidy, and the review begins with a reassuring green check.

There is still an unanswered question: did both artifacts interpret the requirement incorrectly in the same way?

Tests are useful evidence of specified behavior. They do not become an independent account of what the system should do merely because they occupy a different file. If the implementation and its expected results come from the same unresolved assumption, agreement between them leaves that assumption untouched.

This matters in any development workflow. AI assistance can make it especially easy to produce a large amount of internally consistent work before someone examines the underlying decision.

Write the assignment around the outcome

Before delegating implementation, describe the objective, relevant context, permitted scope, authority, and acceptance evidence. Explain what existing behavior matters and which uncertainty should be raised rather than silently resolved.

A request to “fix the validation” gives considerable room for interpretation. A stronger assignment identifies the affected workflow, the valid and invalid conditions already established, the interfaces that must remain compatible, and the consequences of rejecting or accepting the wrong input.

The assignment does not need to anticipate every detail. It should make the important boundaries visible and name the human responsible for decisions outside them. When an unresolved domain rule changes behavior, clarifying that rule is part of the work.

Obtain expected results from somewhere meaningful

Start acceptance evidence with requirements, documented domain rules, representative examples, or behavior that the team has deliberately chosen to preserve. Examine their quality too. Existing behavior can contain a defect; documentation can omit an exception.

For consequential rules, work through examples independently of the proposed implementation. Include boundaries, missing information, conflicting inputs, and cases where the correct result is to refuse or request a decision. Ask someone with the relevant responsibility to confirm the expected behavior when software expertise alone cannot establish it.

Independence comes from the basis of the evidence. Asking the same assistant to explain its own answer may help reveal reasoning, but it does not supply a new source of truth. Asking another model may produce useful criticism without resolving an ambiguous business rule either.

Review beyond the happy path

Correct output for a few examples is only part of a complete workflow. Examine permissions, input validation, error handling, retries, data changes, and operational visibility where relevant. Consider what can appear to succeed while leaving someone with an incomplete or harmful result.

Generated tests may cover many branches while avoiding the consequential question. A test can verify that an error message renders without checking whether an unauthorized action was prevented. It can confirm a database write without verifying that a failed retry cannot create a duplicate.

Use the intended behavior and the risk of failure to guide the review. Test quantity, naming quality, and coverage percentages can help describe the suite. They cannot decide which obligation matters most.

Check the checks

For important behavior, ask whether a plausible defect would make the relevant evidence fail. Change a boundary deliberately in a safe local check, or inspect an existing failing example. If the suite remains green when the requirement is violated, find out what it actually verifies.

Keep this exercise proportionate. The purpose is to challenge a consequential assumption, not manufacture endless verification work. Strong existing evidence may already answer the question.

Read the production change as well. Look for unnecessary scope, hidden defaults, new dependencies, and decisions that were never authorized. A valid test does not make every implementation choice desirable or maintainable.

Leave a reviewable account

When the work is ready, the crew should be able to explain what changed, which evidence supports it, and what remains uncertain. Preserve the relevant reasoning beside the code or in the decision record, wherever the next maintainer will look.

A named human remains accountable for accepting the change. That responsibility can be supported by automated checks and delegated investigation. It still includes judging whether the evidence answers the actual question.

The practical habit is simple: before accepting generated code and tests, identify one meaningful source of expected behavior that does not depend on the implementation’s own interpretation. Then verify that the work meets it.

Use the expedition brief Back to The Logbook

Copy the text manually

Your browser could not copy this automatically. Select the text below and use your device’s copy command.