Survey · 2026

Harness-AwareEvaluation ofLLM Agents

Systems, Benchmarks, and Protocols

A benchmark score characterizes a complete evaluation configuration, not the model alone.

Vector Institute University of Agder

arXiv: XXX Repository
MModel
HHarness
EEnvironment
VEvaluator
Reported score SB fB(M, H, E, V)
Evaluation protocol
Scroll to explore

Abstract

The system around the model changes what the score means.

LLM-based agents combine model capabilities with memory, tools, and execution control. Changing this surrounding harness can alter performance and even reverse model rankings, yet benchmark reports often omit the configuration, resource budget, or scoring procedure needed to interpret the result. This survey introduces a four-factor framework for harness-aware evaluation, maps representative systems and benchmarks to it, and provides reporting guidance for more consistent and reproducible comparisons.

The framework

Evaluate the configuration, not an isolated model.

Every result emerges from four interacting factors, under a protocol that controls tasks, budgets, repeated runs, and aggregation.

M

01

Model

Backbone weights, decoding choices, and tool-use capability.

H

02

Harness

Context, tools, memory, control flow, verification, and permissions.

E

03

Environment

Tasks, data, applications, services, users, and execution sandbox.

V

04

Evaluator

Tests, rubrics, judges, outcome checks, and success criteria.

Protocol
  • Task splits
  • Resource budgets
  • Repeated runs
  • Aggregation
Overview of the model, harness, environment, and evaluator in harness-aware agent evaluation.
Overview of harness-aware evaluation. Reported performance belongs to the complete configuration.

Inside the harness

Six mechanisms shape agent execution.

01

Context & memory

Select, retain, retrieve, and compress information across steps.

02

Tool interfaces

Define operations, schemas, dispatch, and returned observations.

03

Planning & execution

Decompose tasks and govern how plans become actions.

04

Coordination

Assign roles, share state, and route work across multiple agents.

05

Verification & recovery

Check intermediate work, retry failures, and revise trajectories.

06

Guardrails & permissions

Constrain actions, enforce boundaries, and escalate to people.

Same model, different harness

One mechanism is enough to move the score.

Each circle is one change, with the model and benchmark held fixed. A model swap on the same interface, GPT-4 Turbo to GPT-4o, moved τ-bench by 3.5 points. These contrasts come from separate studies, so they are not a ranking.

+3.0 15% to 18% Context
+6.0 12% to 18% Tools
+0.2 73.0 to 73.2 Verifier
−1.6 73.0 to 71.4 Planning
Fixed model SB

Evidence base

A rapidly shifting research landscape.

The selected corpus is recent and unevenly distributed, revealing where evidence is strongest and where attribution remains limited.

110works reviewed
75%published in 2025–26
56%preprints or system docs
16titles using “harness”
Reviewed works by research topic and year from 2022 to 2026.
Research topics in the selected corpus.
Reviewed works divided by publication type.
Publication types in the selected corpus.
Terms used in paper titles, sized by frequency and colored by whether they skew toward recent work.
Title terms show the recent shift toward explicit harness research.

What this survey contributes

01

Unified evaluation taxonomy

Connects model, harness, environment, and evaluator in one configuration, interpreted under an explicit protocol.

02

Mapping of empirical evidence

Distinguishes factors that studies explicitly evaluate from those merely present in the execution setting.

03

Reporting guidance and gaps

Identifies what evaluations should report and where evidence remains thin across adaptation, resources, and interactions.

Timeline from LLM foundations through agent execution and frameworks to harness control and evaluation.
Selected milestones from language-model foundations to harness control and evaluation.

Open challenges

What harness-aware evaluation needs next.

01

Adaptive harnesses

Report the initial configuration, modification space, feedback, update frequency, and resulting harness.

02

Fair resource budgets

Compare systems under matched or normalized tokens, calls, time, tool use, compute, and human intervention.

03

Generalizable effects

Test interventions across multiple models, environments, and task families with controlled comparisons.

04

Factor interactions

Measure how model, harness, environment, and evaluator affect one another without exhaustive search.

05

Complete reporting

Develop standardized, machine-readable descriptions of configurations, budgets, and repeated-run procedures.

Citation

Cite this survey.

The publication identifier will be updated when the paper is posted.

@article{radwan2026harness,
  title   = {Harness-Aware Evaluation of LLM Agents:
             Systems, Benchmarks, and Protocols},
  author  = {Radwan, Ahmed Y. and Vasilakos, Athanasios V.
             and Raza, Shaina},
  journal = {arXiv preprint arXiv:XXXX.XXXXX},
  year    = {2026}
}

Acknowledgements

Resources used in preparing this research were provided, in part, by the Province of Ontario, the Government of Canada through CIFAR, and companies sponsoring the Vector Institute. This research was funded by the European Union’s Horizon Europe AIXPERT project (Grant Agreement No. 101214389).