The usual testing debate asks whether a team has the right balance of unit tests and integration tests. That remains useful—but it is not the deepest problem for AI- and data-driven systems.
In conventional software, testing principally compares an implementation with a specification. In systems that extract relationships, form knowledge graphs, rank evidence or make recommendations, there is a second question: does the specification describe the relevant reality well enough in the first place?
That is an epistemic problem, not merely an engineering one.
The Test Oracle Meets the World
Every automated test needs an oracle: a criterion that says whether an outcome is right or wrong. For a deterministic function, that criterion can be simple.
assert add(2, 3) == 5
Now consider a system that turns documents into a knowledge graph. A complete oracle would need to determine which entities truly exist, which relationships are stated or implied, where ambiguity remains, which links must not be inferred, and which contextual information is missing.
If another system could answer all of these questions reliably and automatically, it would reproduce a substantial part of the original understanding problem. This is the test-oracle problem in its most demanding form: the test becomes entangled with the intelligence the product is intended to supply.
That does not make evaluation impossible, and it does not mean every oracle must be more capable than the system under test. Partial oracles can check constrained properties. Human review can judge a sample. Reference cases can preserve known expectations. Metamorphic tests can verify that a controlled change in input leads to a required change—or non-change—in output. The limit concerns complete semantic verification: no test can automatically identify every missing insight unless it already knows what that insight should be.
Integration Is Also a Meaning Question
For conventional software, an integration test asks whether components work together: does the service call the database, does the schema survive the boundary, does an error propagate correctly?
AI systems need that question too. But their integration introduces another one: does the combined system produce the intended meaning?
| Test layer | Primary question |
|---|---|
| Unit test | Do local, deterministic properties hold? |
| Contract test | Are interfaces, schemas and invariants preserved? |
| Integration test | Do components and data flows work together? |
| Semantic system test | Is the assembled result faithful, coherent and useful? |
| Exploratory evaluation | What surprising omissions, interpretations or new failure classes appear? |
| Production evaluation | Does quality hold under real data, drift and changed conditions? |
The final three layers are rarely cleanly separable. An evaluator sees an unexpected graph edge, identifies an omission, changes a prompt or transformation, then discovers that a different interpretation has shifted elsewhere. This is not an ordinary regression in a loosely coupled program. It is a behavioral change in a coupled socio-technical system.
The Specification Is an Incomplete Model of Reality
Take a familiar requirement: “As a user, I want relevant entities and relationships extracted from documents so that I can analyze connections.” It sounds actionable until the hidden terms appear.
What counts as relevant? What is a relationship? Which connection is implicit rather than explicit? How should contradictory sources be handled? What should happen when the ontology has no category for a newly encountered fact?
Teams can write more acceptance criteria. Often they should. But at a certain point the work has not eliminated uncertainty; it has moved it from implementation into specification. The specification itself is an incomplete model of the world.
This is why a green CI pipeline can coexist with an unhelpful result. The graph may be valid. The database may be consistent. Every service contract may pass. Yet a consequential relationship may be absent, and no error condition appears. A traditional regression test detects that absence only if the team already knew the connection should exist.
Three Different Kinds of Correctness
Separating the problem clarifies which evaluation methods are appropriate.
- Syntactic correctness is highly automatable. Is output well-formed, complete enough to parse and compliant with its contract?
- Semantic correctness is partly testable. Is an extracted claim faithful to the source, internally consistent and supported by evidence?
- Epistemic adequacy requires open evaluation. Did the system identify relevant relationships, preserve uncertainty and avoid unjustified conclusions?
The third level is the uncomfortable one. A system may be syntactically perfect and locally accurate, yet still fail to surface the structure that would matter to the user. It may state no falsehood, but it may fail to make a discovery.
Evaluation should therefore measure not only false positives but also blind spots: missing relations, collapsed ambiguities, unsupported bridges between sources and situations where uncertainty was silently converted into a confident fact. The exact mix of checks depends on the domain and on the cost of error.
A Development Loop for Learning Systems
The familiar delivery sequence is linear:
Specification → implementation → verification → done
For exploratory AI engineering, a more truthful model is a loop:
Hypothesis → implementation → evaluation → counterexample → revised criteria → next hypothesis
Tests remain essential within this loop. They protect known behavior, schemas, latency budgets, safety constraints and previously resolved failures. But they are not the entire quality process. The loop must also produce evidence beyond code: curated evaluation sets, documented counterexamples, taxonomies of failure modes, review decisions and more precise quality criteria.
This changes the unit of progress. A closed ticket is useful for coordination, but it cannot by itself express whether a system has learned the edge cases that matter. The durable asset is not only the implementation; it is the growing map of what the team knows about its system and its domain.
Why User Stories Still Persist
User stories, pull requests and acceptance criteria do more than describe software. They coordinate people, bound responsibility and make delivery legible. Organizations need to know what will be delivered, who owns it, when it can be accepted and what evidence supports that decision.
Open-ended epistemic evaluation fits this structure poorly. “Done” may instead mean: the observed quality is adequate for the intended context, the remaining uncertainties are visible and accepted, and there is a path to detect and correct drift.
That is harder to plan than five criteria and a successful build. It is also more honest. Translating an uncertainty problem into a conventional implementation task can make the work administratively manageable, but it does not remove the underlying uncertainty.
Build an Evaluation Architecture, Not Just a Test Suite
A practical approach makes the limits explicit and gives each method a job:
- Use unit and contract tests for deterministic transformations, schemas, permissions and safety invariants.
- Maintain a versioned evaluation set with representative cases, hard negatives and counterexamples discovered in use.
- Score provenance and uncertainty, not only answer similarity: can a claim be traced, qualified and contested?
- Run targeted human review on high-impact or ambiguous cases, recording both the decision and its rationale.
- Use metamorphic tests where a complete answer is unknown but a relationship between outputs is known—for example, removing a source must not increase confidence in a claim derived only from that source.
- Monitor production data for drift, changed input distributions and novel failure patterns.
- Treat the absence of expected relationships as a first-class evaluation category, not as invisible success.
The aim is not perfect knowledge before deployment. It is a system capable of revealing what it does not know, gathering evidence against its own conclusions and incorporating what evaluation teaches.
The Actual Divide in AI Engineering
The important divide is not unit testing versus integration testing. It is development under largely known requirements versus development under irreducible epistemic uncertainty.
In the first case, the specification can usually stabilize before implementation. In the second, parts of the specification emerge only through implementation and evaluation. Requirements engineering, estimation, acceptance and accountability must all adapt to that fact.
AI engineering becomes reliable not when it pretends that every meaningful outcome can be specified in advance, but when it turns uncertainty into an explicit, inspectable and continuously evaluated part of the system.
