A Reliable Harness Is Not Enough
August 25, 2026
A Reliable Harness Is Not Enough
Why Legal AI Requires Adequacy Standards, Genuine Human Judgment, and Evidence of Justified Reliance
By Alok Priyadarshi and Krishna Puranik
The Earthquake Nobody Is Governing For
The legal profession's AI governance conversation has been organized around one failure mode: hallucination. AI systems that fabricate citations. The assumption embedded in every bar guidance, court sanction protocol, and AI detection tool is the same: the error signal is present, and the governance task is catching it.
That is the tremor. The earthquake is different.
The real risk is AI that produces fluent, well-structured work product that looks complete and is not, and that actively suppresses the review posture that might catch what is missing. A contract analysis that reads as thorough but omits the obligation type the matter turns on. A document review that leaves a material gap in responsive production. The output is excellent. The problem is what it does not contain.
This is not theoretical. A recent empirical study found generative AI achieves roughly 88% recall on complex document review protocols, routinely cited as a success. What it also means is that one in eight responsive documents was not found, and the review appeared complete.
The governance frameworks being built for legal AI do not yet answer the harder question: how do we determine whether apparently successful work is adequate for the professional purpose it must serve? That is the problem this article addresses.
The Harness Has Already Won the Argument
When legal teams evaluate AI tools, the instinct is still to focus on the model: which performs best on their matter types, which scores highest on legal reasoning benchmarks, which vendor produces the most impressive demonstration. That instinct is increasingly out of date.
Harness engineering, the discipline of designing everything that surrounds the model, the instructions it operates under, the data it can access, the workflow it sits inside, the human checkpoints it hands off to, and the controls that govern all of it, has rapidly become an established part of the AI engineering conversation. Across software development and agentic systems generally, the field is now organized around orchestration, context, permissions, tools, verification, and control. Ryan McDonough, Head of Engineering at KPMG Law, captured the legal-specific version of this shift when he observed that by 2026 legal AI would be judged less on interface quality and more on operational discipline, with procurement hardening to demand task-level evidence and traceability of outputs.
That advance is real, and it matters. The model determines what an AI system can do. The harness determines what it actually does. But professional work introduces a further question. Even if a harness executes exactly as designed, routes the right task to the right model, enforces the right permissions, and produces a fully traceable record, how do we know that what it produced was sufficient for the purpose someone intends to rely on it for? A harness can be operationally reliable and professionally inadequate at the same time.
Harness engineering cannot, by itself, answer that question. Not because engineers have not gotten to it yet, but because professional adequacy is not fundamentally an engineering property. It requires domain standards, context, purpose, and supervisory judgment, the particular requirements that govern whether reliance is justified in a specific matter.
Capability Is Not Adequacy
Sirisha Gummaregula, CEO of QuisLex, made a distinction in her analysis of Harvey's Legal Agent Benchmark that deserves to be carried forward. A benchmark score tells you the model is capable. It does not tell you whether this output, on this matter, is sufficient for reliance, that is, whether it is adequate. Capability is a property of the system. Adequacy is a determination about a specific output in a specific context, and it requires a standard against which the output is measured.
ABA Formal Opinion 512 states that what constitutes an appropriate degree of independent verification of AI output depends on the tool and the specific task. The profession has named the adequacy question and explicitly declined to answer it, placing the standard-setting obligation on individual lawyers without defining what that standard must meet. That is the gap this article addresses.
Much of current harness engineering is designed to ensure that a system executes reliably against its specified requirements, through verification steps, permission scopes, failure attribution, and deterministic assertions. Those mechanisms can establish whether the system operated as specified. Adequacy asks something different: whether what was specified, and ultimately produced, was sufficient for the professional purpose the work was intended to serve.
Increasingly sophisticated benchmarks do not eliminate this distinction. A model optimized against a detailed rubric may become exceptionally good at satisfying everything that rubric anticipates. But improved performance against specified criteria says nothing about what the specification itself failed to capture. No matter how capable a system becomes against known requirements, residual risk remains in what those requirements did not know to ask for.
An audit trail that records every agent action is not a governance standard. A certification that covers how an AI system is managed is not a determination that any specific output from that system was adequate for reliance. These are different claims, and the legal profession needs to be precise about which one it is making.
A harness can verify execution against a standard. It cannot establish, merely by doing so, that the standard itself was adequate. Defining the adequacy standard is not a technical problem. It is a legal domain problem, and it must be solved before the harness is built.
Harness Integrity: What Professional Reliance Requires
The distinction between a capable harness and a trustworthy one has a name, and QuisLex is proposing what that name should mean in professional legal work.
Harness integrity is the property of an AI deployment that ensures human judgment at every checkpoint is genuine, structured, and auditable, and that the adequacy standard being enforced reflects what the work actually requires.
This is a different problem from capability. A capable harness routes work correctly, applies the right model to the right task, and surfaces outputs efficiently. A harness with integrity does something harder: it ensures that the humans embedded in that workflow are genuinely exercising judgment rather than nominally approving outputs, and that the standard against which outputs are evaluated is the right one for the purpose the work will serve.
Harness engineering can address the first problem. The second requires something more: a professionally grounded adequacy standard and a way to establish that genuine judgment was exercised against it. That is where the earthquake-level failures accumulate quietly, behind clean audit trails and smooth workflows.
Five Ways a Harness Fails
Before describing what harness integrity requires, it is worth naming what its absence produces. QuisLex's taxonomy of the Five Failure Modes of Legal AI identifies the patterns through which AI-assisted legal work breaks down even when the harness appears to be functioning.
Silent omission is the failure mode the earthquake framing describes: the output is fluent and well-structured, and the problem is what it does not contain. No error is flagged because no error, in the system's terms, occurred. Even where a review is correctly scoped, material information can fall outside the model's effective engagement without ever producing a visible signal that anything was missed.
Boundary failure occurs when an AI system correctly answers the question it was asked but fails to surface information that lies just outside the defined scope of that question. The answer is accurate. The analysis is incomplete in a way that is consequential.
Confident inconsistency occurs when an AI system produces materially different outputs in response to the same query, applied to equivalent documents or provisions, at different times or across different parts of a matter portfolio. Each individual output appears confident and complete. The inconsistency is invisible without systematic cross-matter comparison.
Context drift occurs in multi-step or agentic AI workflows where the system's understanding of its task, its scope constraints, and its risk parameters shifts as it processes more information. The output produced at the end of a long workflow may not reflect the instructions, constraints, or analytical framework defined at the beginning.
Hallucination is the failure mode the field has already organized around: fabricated facts, invented citations, confident assertion of things that are not true.
Each of these is a harness design problem as much as a model problem. A well-designed harness can detect, mitigate, or govern all five. None of them can be reliably prevented by capability alone.
For example, an NDA abstraction that misses a non-standard carve-out to the confidentiality obligation does not look wrong. A contract summary that characterizes a termination right differently from the underlying clause does not trigger an error flag. Both pass through a capable harness without friction. Neither passes through a harness with integrity.
The Failure Mode That Enters Before Output Evaluation
There is a sixth failure vector that an output-focused harness may not address: AI sycophancy, because it operates upstream of output evaluation, in the human-model interaction that shapes the task before generation even begins.
A February 2026 study from the UK AI Security Institute (Dubois, Ududec, Summerfield, and Luettgau, “Ask Don't Tell,” arXiv:2602.23971) established the following experimentally: model agreement with the user's stated position increases monotonically with the epistemic certainty the user expresses. A reviewer who frames a prompt as “confirm that this clause does X” receives a systematically more agreeable response than one who asks “what does this clause do.” This effect is stronger than simply instructing the model not to be sycophantic.
The implication is precise: framing-induced deference, the tendency of the model to yield to the reviewer's implicit position before the task is even executed, is a failure mode that operates before output evaluation. A harness that governs only outputs has already missed it, the framing shaped the result before there was anything to measure. A well-designed harness can govern framing itself, the prompts, the task construction, the interaction upstream of generation, but an output-only verification regime cannot detect that the output was biased by how the question was posed.
What Harness Integrity Requires: Engineered Judgment
The assumption embedded in most harness design is that human checkpoints deliver human judgment. They deliver human presence. Whether genuine judgment occurs at those checkpoints depends entirely on how they are designed.
This distinction matters because the human-in-the-loop is the harness's last line of defense against every failure mode described above. Silent omissions surface when reviewers interrogate outputs, not when they ratify them. Framing-induced deference is interrupted when reviewers are required to pose questions rather than confirm positions. Boundary failures are caught when reviewers are asked to assess scope, not just sign off on conclusions.
The design question that harness integrity requires is not “did a human approve this output?” It is “did a human engage with this output in a way that constitutes genuine scrutiny, and is there a record that distinguishes that engagement from nominal sign-off?”
Answering that question requires structured engagement gates: checkpoints designed not merely to capture a decision but to ensure that decision emerged from deliberate reasoning. A structured engagement gate governs what the reviewer is asked to do, not just when they are asked to appear. It surfaces the AI's reasoning, not just its conclusion. It requires the reviewer to interrogate specific dimensions of the output rather than assess it in aggregate. And critically, it governs how the task is framed before it reaches the model, addressing framing-induced deference at the point where it originates. In practice, this means a contract reviewer working inside a structured engagement gate is asked to pose a question about the clause rather than confirm an interpretation of it, a design choice, not a training exercise.
It is worth being precise about what structured engagement gates address. Framing-induced deference and cognitive surrender are distinct failure modes that compound each other. Framing-induced deference is the model deferring to the reviewer's implicit position before the task executes. Cognitive surrender is the reviewer deferring to the model's output during review, adopting it with minimal scrutiny. Both operate simultaneously in an unengineered checkpoint. The structured engagement gate addresses both: it governs how the task is framed upstream of execution, and it structures how the reviewer interrogates the output downstream of it.
When genuine judgment has occurred, it should be distinguishable from nominal approval, and that distinction must survive challenge. Judgment provenance is the structured record of how an adequacy determination was reached: the standard applied, the dimensions interrogated, and the basis on which reliance was justified. An approval timestamp records that a human was present. Judgment provenance records what that human actually did. Where AI-assisted work may be tested by a client, regulator, or court, the difference between those two records is the difference between a defensible governance position and an audit trail that creates the appearance of one.
This design philosophy has a name: Artisanal Intelligence™. Artisanal Intelligence™ is the discipline of designing AI systems that protect and deepen human judgment at high-consequence decision points, rather than accelerating past it. It is not a rejection of AI scale. It is the recognition that scale without judgment quality is a systematic risk that compounds silently until it surfaces in a matter where it cannot be contained. In legal operations, where reviewer expertise creates strong epistemic confidence and volume creates relentless time pressure, Artisanal Intelligence™ is not a design aspiration. It is the design requirement that separates a harness that looks trustworthy from one that is.
What a Harness with Integrity Actually Looks Like
A harness with integrity is not a philosophy bolted onto a model. It is a system with controls at every stage at which the adequacy standard can be compromised: before generation, so the model operates against a defined standard rather than an open brief; during execution, so the framing governs what the model is required to surface, not just what it chooses to; at the human determination checkpoint, so that genuine judgment is the mechanism by which the gate clears; and after determination, so that the audit trail captures the quality of engagement, not just the presence of a human reviewer.
Who Defines the Standard
The harness engineering conversation has converged on a true observation: the harness will define institutional performance. What it has not answered is the question that observation raises.
The harness measures output completeness against a standard of adequacy. That standard did not come from the model, the platform, or a benchmark. It cannot be inherited from a vendor's default configuration, built for the average use case and not for this matter, this workflow, or this client's reliance requirements.
Adequacy standards for legal AI work, what constitutes a complete contract review, which obligation categories must be captured, what document universe must be covered for a review to be defensible, are domain determinations. They require the kind of empirical grounding that only comes from deep operational history: structured records of where legal work succeeds and fails across real matters, at scale, as evaluated by professional supervisory judgment.
They also require something beyond data. They require the judgment to know which adequacy signals are stable across matter types and which are context-dependent, which gaps are tolerable and which are not, and how the standard should evolve as AI capability improves and failure profiles shift. Adequacy standards are not static. As models improve, the gap between what they produce and what the work requires narrows, but it never closes, and the shape of the gap changes. The durable advantage lies not in data volume alone, but in the institutional discipline required to interpret adequacy signals and translate them into defensible standards.
Two decades of ISO 9001-certified, Six Sigma-disciplined legal service delivery at QuisLex produced something that only becomes legible now. Two decades of structured legal quality data provide an empirical foundation that cannot easily be replicated. The reason that foundation exists is deliberate methodology, not accumulated volume.
This is where legal domain expertise becomes irreplaceable. Not as a check on AI outputs after the fact, but as the foundation on which adequacy standards are defined before the harness is designed.
The organizations that define institutional performance in legal AI will not be those with the most sophisticated models or the most elaborately engineered harnesses. They will be those who brought the operational depth to define what adequate looks like, and then built harnesses designed to enforce and defend it.
Three questions can start that assessment. What happens before the model runs: is the adequacy standard defined, the document universe scoped, and reviewer framing governed, or is each left to individual habit? What do your human checkpoints actually demand: are they capturing presence or judgment, and does your audit trail carry judgment provenance or only a timestamp? Who defined the adequacy standard your harness enforces, and was that definition grounded in operational expertise and empirical history, or inherited from a vendor's default configuration?
Scott Milner, partner and eData practice leader at Morgan Lewis, observed in the same National Law Review piece that as legal AI outputs continue to improve, hallucinations become harder, not easier, to detect, with risk shifting from obviously wrong answers to confidently delivered, plausibly incorrect ones that evade surface-level review. That is a precise external description of silent omission and boundary failure. The earthquake is already happening. The governance question is whether legal organizations are building harnesses designed to feel it.
The harness tells you whether the system operated as designed. Harness integrity tells you whether its output was subjected to the standard and judgment required for professional reliance. The harness is necessary. It is not sufficient. The adequacy standard inside it, defined before the model runs, enforced through genuine human judgment, and documented with the provenance to defend it, is what makes reliance on AI-assisted legal work defensible when it matters most.
Download our Governance Taxonomy for AI in Legal Workflows to dive deeper into the Five Failure Modes of Legal AI.
Alok Priyadarshi is Vice President of Strategic AI Advisory and Legal Transformation at QuisLex. Krishna Puranik is Head of Information Technology at QuisLex. QuisLex is an Alternative Legal Services Provider specializing in AI-enabled legal operations, governance, and managed services.