When Green Checks Pass but the Requirement Fails
Passing tests can validate an incomplete outcome. A telemetry example, ways to test the result, and a practical release record.
Summary: Passing tests show that the assertions in a particular run held. They do not, by themselves, show that the requirement was met. A check that a message was sent verifies an activity; a requirement that an operator sees the right temperature calls for evidence at the display. Start with the required outcome, test it at the relevant boundary, and record what remains unverified.
What a green check establishes
CI reports results for a particular build, environment and set of inputs. Unit and schema tests are valuable because they quickly locate faults and protect contracts. Their names and release records should describe the claims they actually support. Google's guidance on testing behavior makes the related point that tests of public behavior are generally more resilient than tests tied to implementation details. An interaction check can still be useful, but it cannot establish a different, user-visible result.
A fictional telemetry failure
Illustrative scenario, not a real incident. A sensor node publishes a temperature message, a back end stores it, and a monitoring screen shows the latest reading. The requirement is: “For the specified reference condition, the operator display shows the current node's temperature in °C within ±0.5 °C.” For this example, an independently measured reference is 25.0 °C and the specification's acceptance interval is 24.5 to 25.5 °C. These are assumed example values, not a universal conversion from ADC counts. Raw count 2048 has no temperature meaning without the sensor, circuit, calibration and conversion specification.
The following pseudocode checks the publishing activity. The classes and methods are illustrative, not an API provided by this site.
# Pseudocode: verifies message shape and an interaction only.
node = Node(adc=FakeAdc(raw_count=2048), broker=FakeBroker())
node.sample_and_publish()
message = node.broker.last_message
assert message["type"] == "temperature"
assert isinstance(message["value"], (int, float))
assert node.broker.publish_count == 1All three assertions can pass if a conversion defect puts the raw count into value. The message is present and numeric, while the displayed result may be wrong. Keep this test as a schema or publishing check; do not label it temperature accuracy evidence.
Test the outcome at the boundary that matters
Choose expected values from a specification or independent reference measurement, never from the same conversion code under test. Correlate the reading with the node and the fresh sample, for example with a sample ID and timestamp. Check the unit, numeric tolerance, and behavior when the value is missing, stale or in the wrong unit. An old 25.0 °C result must not satisfy a requirement for the current sample.
This second example is also pseudocode. It demonstrates an API integration assertion against a replayed, independently characterized fixture:
# Pseudocode: fixture maps sample 42 to reference 25.0 °C.
rig.replay(characterized_sample_id=42)
reading = backend_api.latest_reading(node_id=rig.node_id)
assert reading.sample_id == 42
assert reading.unit == "°C"
assert reading.age_seconds <= SPEC_MAX_AGE_SECONDS
assert abs(reading.value - 25.0) <= 0.5 # example spec toleranceThat proves the replay and API path under the tested conditions. If the requirement says the operator screen shows the value, acceptance testing must also inspect the rendered value and unit for sample 42, including how stale or missing readings appear. Until that screen assertion runs, record the display boundary as unverified. A replay does not establish sensor accuracy, analog behavior, calibration or transport on physical hardware. A bench run with a traceable reference is separate hardware verification.
Check that a relevant fault is detected
Temporarily make the conversion return the raw count, or swap °C for another unit, in a controlled test branch. The outcome assertion should fail; restore the code afterwards. This demonstrates detection of that particular injected fault under those test conditions. It does not prove complete fault coverage. Google's mutation testing discussion describes this general technique of introducing faults and observing whether tests catch them, along with the need to avoid unhelpful mutations.
Mocks keep unit tests fast and focused, but a fake sensor or broker only models the behavior its author supplied. Retain those tests for logic and schema checks. Add boundary checks where risk justifies them; every requirement does not need a heavy end-to-end test.
A release record the team can reuse
Copy this structure into the release evidence, replacing the illustrative statuses with actual results and links. An unrun check stays open, even if neighboring tests pass.
| Requirement or boundary | Evidence for this example | Status |
|---|---|---|
| Message has required fields | Schema test, CI run link | Passed in CI |
| Fresh API reading is 25.0 ±0.5 °C for sample 42 | Characterized replay, API assertion, CI run link | Passed in replay |
| Operator screen shows sample 42 value and °C, handles stale or missing data | Rendered screen assertion, result link | Unverified until run |
| Physical sensor meets specified accuracy | Bench procedure, reference instrument and result link | Unverified until run |
Record header: requirement ID and specification revision; build/firmware hash; CI run and replay fixture version; hardware revision and environment; reference instrument and calibration record when applicable; date; result; open limits; owner and next action. For the two open rows above, assign owners and a planned verification environment before a release decision. This is a template, not a claim that these tests were run on a real product.
Release checklist
- Write the required outcome with observation point, unit, tolerance and freshness rule.
- Map each claim to a named test or measurement and its build and environment.
- Use an independent expected value and a correlated sample, then check wrong units and stale or missing values.
- Separate API integration, rendered screen and physical hardware evidence.
- Inject a relevant fault once to show the intended check can detect it, then restore the code.
- List gaps with an owner, next action and release decision. Keep useful unit and schema results in their proper scope.
FAQ
Are unit tests still useful?
Yes. They isolate logic faults quickly and protect stable contracts. Their passing result supports those contracts, within the assumptions of their fakes and inputs.
Does each requirement need an end-to-end test?
No. Choose the cheapest credible evidence for each risk and boundary. A calculation may be well covered by reference cases at the component interface; a display requirement needs evidence from the rendered display.
What if hardware cannot run in CI?
Use characterized replay for the software path in CI and schedule a bench run for the hardware path. Keep the two claims and their environments separate in the release record.
See how we approach these evidence boundaries in quality engineering.