Ask a generative system for a visually ambitious slide and it may give you something polished, colourful and technically competent—then quietly turn the central idea into abstract art.
That failure is easy to misdiagnose as “too much complexity.” But a complex SVG can work remarkably well when its problem remains homogeneous: many geometric elements, governed by one visual language, assembled into one kind of object. The difficulty changes when a single slide must make several different kinds of requirement true at once.
It has to communicate visually, state an argument precisely, embody an abstract concept, create a coherent design and make relationships legible. None of these dimensions is optional decoration. Each constrains the others.
This suggests a more precise working hypothesis: cross-domain constraint collapse. A generative system can successfully handle multiple requirement dimensions in isolation, yet lose consistency between them when it must realize them simultaneously in a longer sequence of generation, revision and assembly.
This is not an established scientific term, nor a claim about a known internal mechanism. It is a proposal for a failure mode that can be tested.
A Slide Is a Coupled System, Not a Checklist
Consider five requirement dimensions in a presentation task:
| Dimension | What must remain true |
|---|---|
| Visual communication | Attention is directed toward the intended perception or contrast. |
| Text communication | The claim is precise and the argument says what it means. |
| Abstract concept | The underlying idea—not just its vocabulary—is represented. |
| Design | Composition, hierarchy, rhythm and aesthetic choices support the message. |
| Connections | The relationships between claims, objects and narrative beats are visible. |
A brief can list all five. A model can even repeat all five back accurately. That does not mean it can compose them into one artifact.
A visual metaphor has to fit the abstract concept. The copy must make the same assertion as the image. Spatial placement must reveal the relationship between elements rather than merely make them fit. The slide must also serve its place in the narrative of the deck.
If a model optimizes any one of these locally, it can damage the others. A cleaner layout can erase the unusual visual tension that carried the idea. More explanatory text can weaken the visual hierarchy. A stylish metaphor can become detached from the argument it was meant to clarify. The result may contain every requested object while expressing the wrong proposition.
Why the Difficulty Rises Faster Than the Brief
The relevant complexity is not only the number of requirements. It is the number of dependencies between them.
With five dimensions there are ten possible pairwise relationships. With ten dimensions there are 45. Not every possible relationship matters in every slide, of course. But the space of potential conflicts and consistency obligations grows much faster than the list of individual instructions.
That observation explains an otherwise puzzling practical result: micromanagement often works.
When a human breaks a coupled composition into small, comparatively isolated tasks, the model can solve many of them well. The human is not merely providing more detail. They are taking over the integration work: deciding which relationship matters, accepting a component before the next one changes it, and checking whether the whole still means what it was meant to mean.
The model solves bounded subproblems. The human maintains semantic and visual coherence across the system.
Four Different Things We Call “Context”
It helps to distinguish four phenomena that are often collapsed into one.
| Capability | The real question | Typical failure |
|---|---|---|
| Context availability | Is the information inside the context window? | The requirement was never supplied, or was truncated. |
| Context retrieval | Does the model attend to the relevant information for the next decision? | A stated requirement is present but not used when a local choice is made. |
| Constraint preservation | Does the requirement remain binding throughout execution? | An early design decision satisfies it; a later revision silently removes it. |
| Composition | Can several requirements be realized together in one coherent result? | Each requirement is individually understood, but the assembled output violates some of them. |
More context primarily improves availability. Retrieval systems can improve the chance that a useful item is surfaced. Neither property, by itself, proves that an instruction will remain a constraint throughout a chain of generation and revision—or that two constraints will continue to agree with each other.
This distinction echoes the difference between retrieval relevance and authoritative state. An item can be available for retrieval without being treated as a fact that must constrain the next transition. As argued in Retrieval Relevance vs. Ontological Persistence , a system needs to distinguish what is merely useful to recall from what is still true and binding.
Why SVG Often Works Better Than a Slide
The contrast between a single SVG element and a complete slide is revealing.
One SVG element usually has a compact solution space. The task is local; the result can be checked immediately; and the relevant requirements map fairly directly onto visible properties: position, colour, shape, label or stroke.
A slide has to make several kinds of decision at once. It must express a claim, establish hierarchy, allocate visual attention, make aesthetic trade-offs and leave room for a narrative. A presentation introduces global dependencies too: repeated visual grammar, consistent terminology, progression across slides, and the discipline not to destroy an earlier decision while improving a later one.
The difficulty therefore does not grow linearly with the number of elements. It grows with the number of interacting requirements, the heterogeneity of those requirements and the number of decisions that can invalidate an earlier choice.
This is why a model can produce an excellent icon, chart fragment or SVG component and still deliver an inconsistent deck. Local capability is real; reliable composition is a separate capability.
A Testable Hypothesis: Cross-Domain Constraint Collapse
Constraint interference names the broad pattern in which requirements are missed, displaced or regressed during composition. Cross-domain constraint collapse is a narrower hypothesis about why the pattern becomes especially acute: requirements come from different representational domains and must preserve their relationships across those domains.
Imagine a slide with twenty requirements. A model may understand every one and may be able to satisfy each in isolation. But while generating the complete slide, satisfying one requirement can displace another from effective attention—even when the requirements are logically compatible. More importantly, it can keep all twenty while losing the relationship that made them meaningful.
The model might preserve the requested visual hierarchy but omit the unusual metaphor that made the slide distinctive. It might add the requested data detail but flatten the narrative. It might correct the layout and accidentally remove an approved visual element. Or it might render a beautiful image, correct wording and a neat diagram whose combined message contradicts the intended causal relationship.
Constraint interference is deliberately a hypothesis about observable behaviour, not a claim about a model’s hidden internal mechanism. It does not say that requirements literally compete inside a known neural subsystem. It gives teams a practical label for a pattern they can measure: a rising rate of unmet or regressed requirements as the task becomes more compositional.
The “lost in the middle” literature is relevant here because it shows that information placed in a long input may be used unevenly. But position is only one possible cause. A requirement can be recently stated, highly salient and still lose force after several tool calls, intermediate artifacts and visual revisions. Production systems should therefore evaluate persistence across a process, not only recall from a static prompt.
Measure Relation Fidelity, Not Just Item Completion
An evaluation that asks only, “Is each requested element present?” will miss the central failure. A slide can include a title, a metaphor, a diagram and the specified colours yet still say the wrong thing.
The decisive measure is relation fidelity: the degree to which required relationships are correctly preserved in the completed artifact.
relation fidelity
= correctly preserved required relationships
÷ required relationships evaluated
The relationships should be defined before generation. For example: “the visual imbalance must make the tension in the argument perceptible”; “the callout must identify the cause rather than the consequence”; or “the final element must resolve the question posed by the opening image.” A scorer can judge each relation independently of whether the individual objects are present.
This allows a second, complementary score:
cross-domain fidelity
= correctly preserved relationships spanning two or more domains
÷ cross-domain relationships evaluated
The distinction matters. A model may retain relations within text while losing text-to-image relations, or create a coherent composition whose visual logic no longer corresponds to the stated concept.
An Experiment That Can Falsify the Hypothesis
Hold the model, tools, token budget, scope and evaluation rubric constant. Then progressively increase only the number of interdependent requirement dimensions. Keep the amount of content comparable: a simple conceptual claim, a small amount of copy and a comparable visual footprint.
One useful design compares four conditions:
| Variant | Procedure | What it probes |
|---|---|---|
| A | Meet requirements from one homogeneous domain, such as geometry. | Component capability baseline. |
| B | Add independent requirements from multiple domains. | Whether heterogeneity alone reduces performance. |
| C | Add explicit dependencies between those domains. | Relation fidelity under coupled composition. |
| D | Decompose the same coupled brief with an external constraint register and acceptance checks. | Whether external integration restores fidelity. |
Score item completion, relation fidelity, cross-domain fidelity, regressions after an item was first correct and unintended changes to already accepted elements. Score generic visual polish separately from intent preservation; otherwise an attractive default may conceal a semantic failure.
The hypothesis predicts a sharper decline in relation fidelity from B to C than can be explained by merely adding more items. It also predicts that D improves results by moving some integration work into explicit state, bounded subtasks and validation. If that pattern does not appear across prompts, layouts, models and evaluators, the hypothesis should be rejected or narrowed.
One experiment cannot establish an internal causal explanation. Repeated trials can establish something operationally more useful: which task shapes and production architectures preserve authorship and meaning most reliably.
The More Revealing Metric: What Gets Optimized Away?
Not every omission is equally informative. In creative work, the most interesting details are often unusual, costly to render or difficult to justify through generic design conventions. If those elements disappear more often than ordinary layout rules, the system is not merely forgetting at random.
It may be converging toward a generic, easily realizable solution: clean hierarchy, familiar palette, standard illustration, reduced specificity. That is often attractive at a glance—and a failure of intent at a second glance.
This suggests a further evaluation dimension:
intent preservation
= retained distinctive constraints
÷ distinctive constraints requested
Track distinctive constraints separately from functional ones. “Use the approved title” and “preserve the strange orange object that creates the visual tension” should both count, but they tell different stories when omitted. The second measure tests whether the system preserves authorship rather than merely producing acceptable defaults.
The Architectural Consequence
An LLM can generate components and formulate a plan. It cannot be assumed to keep the relationship between plan, components and accepted whole stable on its own.
That is the role of an external state and control architecture. A constraint register can make requirements and relationships inspectable. An execution graph can assign ownership of a constraint or relation to a component. Validators can reject regressions. Versioned assets can prevent a late edit from silently overwriting an approved decision. Human review can remain focused on the choices where taste and authorship matter.
brief → constraints and relationship register → bounded component tasks
→ per-component validation → composition
→ relation-fidelity and regression checks → accepted artifact
This does not make the model less creative. It gives creativity a durable frame. The model remains valuable where ambiguity, synthesis and variation are useful; the surrounding system retains the responsibility to preserve intent.
That is also the practical argument for process sovereignty : let models propose and generate, but keep state, acceptance and irreversible transitions in a system that can inspect what the model forgot, changed or optimized away.
The next frontier is therefore not only longer context windows or stronger component benchmarks. It is production systems that can tell the difference between having seen a requirement, using it now, keeping it binding, and preserving its relationship to everything else.
