1. Introduction and problem
A ranking circulates. The top row shows an agent name and a number. The laboratory that publishes it does not deliver the step trace, does not declare the token or time budget, does not name the oracle that decides success, and does not decompose why the second place failed. The screenshot travels through networks and slides. A video, cut at the instant the interface seems to “solve X,” accompanies the announcement. That is not an evaluation of a language-model agent: it is a poster. The problem of this article is not that benchmarks exist —they do, and several of them build evidence environments— but that the market for “AI agents” has naturalized substituting the craft of eval with the display of a number, a clip or a lucky demo (Ma et al., 2024; Guo et al., 2024; Aleithan, 2025; Wang et al., 2024; Xi, Chen et al., 2025).
The thesis is restrictive. A cherry-picked demo, a leaderboard screenshot, a single-run score or a video of “the agent solved X” without a fixed evaluation harness do not constitute evaluation of an agent. Reliable eval distinguishes demo capability from evidence of work and requires a verifiable criterion outside the model itself. It relates, without duplicating it, the 16 September 2026 argument on the agent harness: the harness converts generative capacity into work; eval measures whether that work is reliable against an oracle, a budget and a failure taxonomy that do not coincide with the completion itself. They are not the same. A system can have an execution harness and lack eval; it can exhibit a ranking and lack both. Yang et al. (2024) show, as an empirical finding of interface, that the agent–computer interface changes what can be done and therefore what can be measured. Ma et al. (2024) decompose multi-turn failure instead of compacting it into a score. Guo et al. (2024) treat the stability of tool-learning benchmarking as a design problem. Aleithan (2025) recalls that a SWE-Bench score without a data audit is not a verdict. That authorizes asking what was measured: a clip, a table row —or a protocol.
The problem is aggravated by six confusions that product discourse crushes into the word eval. First: eval is not a harness —the harness produces work; eval adjudicates reliability with external evidence (Yang et al., 2024; Xie et al., 2024; Liu et al., 2026)—. Second: eval is not perplexity; environments shift the metric toward the task (Deng et al., 2023; Trivedi et al., 2024; Lee et al., 2025; Zhuang et al., 2023). Third: eval is not a leaderboard without a protocol (Ma et al., 2024; Guo et al., 2024; Xi, Ding et al., 2025). Fourth: eval is not RAG, a fine-tune or the agent label (Patil et al., 2024; Schick et al., 2023; Wang et al., 2024). Fifth: Toolformer, HuggingGPT and Reflexion are pieces, not an evaluation harness (Schick et al., 2023; Shen et al., 2023; Shinn et al., 2023). Sixth: Mind2Web, VisualWebArena, AppWorld, OSWorld, AgentBoard, AgentGym, StableToolBench or SEC-bench are gardens of evidence as protocol and posters as prize (Deng et al., 2023; Koh et al., 2024; Trivedi et al., 2024; Xie et al., 2024; Ma et al., 2024; Xi, Ding et al., 2025; Guo et al., 2024; Lee et al., 2025). Inference: eval is fulfilled by the craft of measurement, not by model scale or clip virality.
There is, besides, an economy of coordination. Artifacts fit in an advertisement; a protocol with oracle, budget, taxonomy and trace does not. Wang et al. (2024), Xi, Chen et al. (2025) and Liu et al. (2026) describe the rise of agents —profile, memory, planning, action— without authorizing a reading as if a verifiable criterion were already in place. Park et al. (2023) and Li et al. (2023) articulate memory, plan and communication: system frameworks, not a unique ranking. Qian et al. (2024) organize roles for software development: organizational craft, not a chat demo. The contributions are three: separating artifacts from craft; tracing boundaries against the harness, RAG, fine-tune, agent marketing, the protocol-less leaderboard and perplexity; and offering four design tests, not an industrial standard. No N, d, r, AUC, leaderboard percentages or DOI is invented.
2. State of the art: situated practice vs artifact
It is useful to separate four strata that the market of “agent evals” usually mixes. The first is the construct of an LLM agent as a system that perceives an environment, decides and acts with tools, not as a longer completion or a slogan (Wang et al., 2024; Xi, Chen et al., 2025; Park et al., 2023). The second is the craft of eval in a real or controllable environment: typed tasks, an oracle external to the model, budgets, a failure taxonomy, trace logging, reproducibility and interfaces that leave a measurable trail (Yang et al., 2024; Ma et al., 2024; Trivedi et al., 2024; Xie et al., 2024; Guo et al., 2024; Xi, Ding et al., 2025). The third is the evidence of capability artifacts —self-supervised tool-use, model orchestration, verbal reflection, massive API invocation, multi-agent communication— without equating them to an evaluation harness (Schick et al., 2023; Shen et al., 2023; Shinn et al., 2023; Patil et al., 2024; Li et al., 2023; Qian et al., 2024). The fourth is the audit of the measuring instrument itself: data quality of the task bank, stability of the tool ranking, analytical decomposition of the multi-turn episode (Aleithan, 2025; Guo et al., 2024; Ma et al., 2024). Inference: a leaderboard does not observe the repository, the web page or the state of the apps; a protocol with an oracle and a trace does.
In construct, Wang et al. (2024) organize profile, memory, planning and action (framework survey). Xi, Chen et al. (2025) review rise and limits (2025 survey). Liu et al. (2026) close the map in software engineering: the agent is played as a system (TOSEM survey). Park et al. (2023) compose memory, reflection and planning as simulacrum architecture, not as a leaderboard. Li et al. (2023) explore communicative agents as a society of models, not as a unique ranking. In capability components, Schick et al. (2023) show typed tool-use —tool capacity ≠ eval with an oracle—. Shen et al. (2023) orchestrate models and tools: planning is not a verifiable criterion. Shinn et al. (2023) contribute verbal reflection: the same channel is not an oracle. Patil et al. (2024) document API hallucination when retrieval evidence is missing (empirical finding). Qian et al. (2024) describe ChatDev: organizational craft, not a chat demo. Inference: the construct authorizes asking for architecture and eval; “piece = eval” is illegitimate.
In evidence environments —here, the operational garden where the agent is measured, not an early-childhood classroom— Deng et al. (2023) and Koh et al. (2024) type web tasks, the latter with visual evidence. Trivedi et al. (2024) build AppWorld, with app state external to the LLM. Xie et al. (2024) situate OSWorld on a real computer. Lee et al. (2025) take the principle to security tasks (SEC-bench). Zhuang et al. (2023) contribute ToolQA: the oracle does not live in the prompt. Status: empirical environment benchmarks or datasets. In the instrument, Ma et al. (2024) —AgentBoard— decompose the multi-turn episode; Xi, Ding et al. (2025) —AgentGym— taxonomize environments; Guo et al. (2024) —StableToolBench— treat the stability of the tool ranking; Yang et al. (2024) show that the ACI enables measurement; Aleithan (2025) audits SWE-Bench data quality. Restrictive inference: the state of the art distinguishes component, environment, board, stability and audit; the market crushes them into a number. These works are not slide trophies; they are protocols that leave a trail.
3. Review method
A critical narrative review centered on agent benchmarks, tool-use and framework surveys (2021–2026) was conducted, not a primary meta-analysis. The purpose was not to estimate a homogeneous effect size nor to combine leaderboard rates —rates this article does not invent— but to articulate a category argument: what counts as evaluation of an LLM-based agent and what remains at the ceiling of a display artifact. Inclusion criteria: (a) 2021–2026; (b) LLM agents, tool-use, agent–computer interfaces, evidence environments (web, desktop, apps, security, QA with tools), analytical boards, benchmarking stability, task-bank data quality, or agent surveys; (c) proceedings or journal with Crossref DOI verified on 18 September 2026 (slot 09:02 America/Mexico_City); (d) relevance to the eval-craft / demo-artifact contrast. Early-childhood axes of this series, product catalogs without a paper, and any figure not present in the verified sources were excluded. The twenty-one sources of fuentes.md were used. SWE-bench, WebArena or AgentBench, when named, are historical context, without invented rates or DOI.
Search and verification were executed on 18 September 2026 against Crossref records of the kit. Each source was read by the object it actually measures or proposes. Empirical finding, framework and design inference were distinguished. When an artifact —cherry-picked demo, leaderboard screenshot, single-run score, video without a trace, chatbot labeled agent, ranking without budgets or taxonomy— lacks a verified trial that equates it to an eval with an external oracle, it is discussed as a category ceiling, not as a quantitative finding. The method does not apply its own PRISMA. No N, d, r, AUC, resolution percentages or DOI was invented. The four tests of section 7 are category hypotheses of this article, anchored in the corpus, not an ISO standard or a substitute leaderboard. The relation to the 16 September harness article is treated as a distinction of object: producing work is not measuring reliability. The unit of analysis is the measurement episode, not a classroom. The leaderboard-without-trace vignette is a limit case of display, not ethnography of a named laboratory; that is why heading 4 is titled axis, not case.
4. Axis 1. The artifacts do not constitute support
The support this axis denies is not that of an adult in a kindergarten: it is the epistemic support a laboratory claims when it declares that it “already evaluated the agent.” The first artifact is the cherry-picked demo: a lucky episode is cut as if it were the distribution of performance. Schick et al. (2023) show typed tool-use; Shen et al. (2023) show orchestration; neither authorizes treating a clip as eval. The demo may be an example of capability; it is not a protocol. The second is the video of “the agent solved X”: it compresses time, hides retries and hallucinated calls, and hides the criterion with which someone off-camera decided it “was solved.” Patil et al. (2024) document API hallucination when retrieval evidence is missing: the failure mode is the one the video does not show. Restrictive inference: “it was seen doing it” is not an oracle; it is edited testimony.
The third artifact is the leaderboard screenshot. Ma et al. (2024) build AgentBoard because an aggregate number does not say where the multi-turn episode broke. Guo et al. (2024) treat instability of the tool ranking as a design object. Xi, Ding et al. (2025) evaluate across diverse environments: a ranking of a single environment is not the craft. The screenshot inherits the visual authority of a table and none of the obligations of the protocol. The fourth is the single-run score —luck of sampling, of temperature or of a poorly filtered issue—. Aleithan (2025) insists on SWE-Bench data quality: without an audit, the score does not distinguish capacity from instrument noise. No percentage is invented here; it is denied that the percentage suffices.
The fifth artifact is the chatbot labeled agent without an oracle. Wang et al. (2024), Xi, Chen et al. (2025) and Liu et al. (2026) describe architecture and the SE field; none equates the label to an eval. Shinn et al. (2023) contribute verbal reflection: an agent that declares itself successful in the same channel has not left the context. Park et al. (2023) and Li et al. (2023) show that memory and communication can be architecture; without an oracle they remain a simulacrum of success. The sixth is the ranking without budgets or taxonomy. Xie et al. (2024) and Trivedi et al. (2024) build environments with observable state; Deng et al. (2023) and Koh et al. (2024) build web tasks with action evidence. A ranking that does not declare steps, tokens or time, nor classify the failure, is a list, not eval. Inference: the demo speaks for the distribution; the video for the oracle; the screenshot for the protocol; the score for the taxonomy; the label for the architecture; the ranking for the craft.
The ceiling does not deny that Toolformer, HuggingGPT, Reflexion, Gorilla or CAMEL are real contributions (Schick et al., 2023; Shen et al., 2023; Shinn et al., 2023; Patil et al., 2024; Li et al., 2023). It denies that their existence authorizes the gesture “there is already eval.” A larger model may produce more coherent traces; Wang et al. (2024) and Xi, Chen et al. (2025) do not authorize reading that coherence as measurement. Qian et al. (2024) organize roles: multiplying characters in a prompt does not install an oracle. Liu et al. (2026) treat the software agent as a system; Yang et al. (2024) recall that the interface is part of what is measured. Undetected error compounds: a hallucinated tool produces a false state; the next step reasons over that state; the output looks finished and is wrong. Without an external oracle there is no one to cut the chain. Design inference: the laboratory that only shows artifacts has not evaluated; it has announced.
5. Axis 2. Relational craft in the garden
The “garden” of this axis is not an early-childhood classroom or an assembly circle: it is the operational measurement environment —repository, browser, desktop, apps, tools, security— in which the craft of eval is recognized as a relational practice among model, interface, oracle, budget and trace. The floor is an episode with typed tasks, an external criterion, a budget, a taxonomy, a log and, in principle, reproducibility (Yang et al., 2024; Ma et al., 2024; Trivedi et al., 2024; Xie et al., 2024; Guo et al., 2024; Zhuang et al., 2023). There is eval when that craft is protected, not when a demo is exhibited. The measured object is the coupling among model, execution harness and instrument. The harness produces the work; eval judges it. Relation, not identity.
The first gesture is typed tasks and an oracle external to the LLM. Zhuang et al. (2023) —ToolQA— formulate questions that require a tool and a criterion outside the model. Trivedi et al. (2024) —AppWorld— build a world of apps with observable state; the oracle is not the agent’s text. Deng et al. (2023) —Mind2Web— and Koh et al. (2024) —VisualWebArena— anchor success in the page and in the action, including visual evidence. Lee et al. (2025) —SEC-bench— take the principle to software security: verification is not “it looks like a patch.” Inference: the oracle is the piece that the video and the judge-chatbot do not replace.
The second gesture is step, token and time budgets, and trace logging. Xie et al. (2024) —OSWorld— situate the agent on a real computer: without a budget the episode is not comparable; without a trace it is not auditable. Yang et al. (2024) —SWE-agent— show that the ACI is part of the system and of what an honest eval must declare (empirical finding of interface). Guo et al. (2024) require stability of the tool bank. The log is the external memory that Reflexion does not replace (Shinn et al., 2023). A run without a trace is an anecdote; a run with a trace and without a budget is an expensive anecdote.
The third gesture is the failure taxonomy. Ma et al. (2024) —AgentBoard— ask in which sub-skill the multi-turn episode broke. Xi, Ding et al. (2025) —AgentGym— taxonomize environments. Aleithan (2025) adds that the failure may be in the data, not in the agent. Restrictive inference: compacting a multi-turn or security episode into a single number is the anti-craft. The fourth gesture is reproducibility and the interface that leaves a measurable trail. Park et al. (2023), Qian et al. (2024) and Li et al. (2023) recall that memory, roles and communication are architecture: they must be declared in the protocol, not hidden behind the LLM’s name. Reproducing is not repeating the screenshot: it is being able to reconstruct tasks, oracle, budget, taxonomy, trace and interface.
The craft is verified when the measurement loop leaves a trail —logs, app states, web actions, oracles, stops— not when the model narrates what it “would have done” or when a human edits a video. The garden of eval is relational because the criterion does not live in the model: it lives between the model and a world that can contradict it. The ranking that hides that contradiction is not a garden: it is a display case.
6. Contrast. Category boundaries
The first boundary is eval versus harness. The 16 September article argued that scale, prompt or a tools demo without an observation–action–verification loop do not constitute a harness; this text uses that as a shore, not as its thesis. The harness produces controlled execution; eval produces a judgment with external evidence (Yang et al., 2024; Liu et al., 2026; Xie et al., 2024). One can have an ACI without a task oracle, or a task bank without isolation. Confusing them allows saying “we have eval” when there is only a runtime, or “we have a harness” when there is only a ranking. The second boundary is eval versus RAG: retrieving passages does not observe a desktop or verify a patch. Patil et al. (2024) show that tool calling requires retrieval evidence; that does not turn RAG into eval. The third is eval versus fine-tune: Schick et al. (2023) teach tools in the model; the checkpoint is not a measurement protocol.
The fourth boundary is eval versus perplexity: Deng et al. (2023), Koh et al. (2024), Trivedi et al. (2024), Xie et al. (2024), Lee et al. (2025) and Zhuang et al. (2023) shift the object toward the task. Predicting the next token is not solving a visual web task or a security issue. The fifth is eval versus agent marketing: Wang et al. (2024), Xi, Chen et al. (2025), Liu et al. (2026), Park et al. (2023), Li et al. (2023) and Qian et al. (2024) describe potential, architecture and roles, not a chatbot label. The sixth is eval versus a leaderboard without a protocol: Ma et al. (2024) decompose; Guo et al. (2024) ask for stability; Aleithan (2025) asks for data quality; Xi, Ding et al. (2025) ask for environment diversity. A table row without oracle, budget, taxonomy and trace does not distinguish model, prompt, harness and luck. Memory, plan, communication, orchestration and verbal reflection are components: without an oracle they can memorize the error or theatricalize consensus (Park et al., 2023; Li et al., 2023; Qian et al., 2024; Shen et al., 2023; Shinn et al., 2023). The corpus environments are gardens of evidence as protocol and trophies as announcement. The category “eval” is conjunctive: one piece missing and the system falls back into artifact, even if the slide says SOTA.
An additional boundary protects this article from colonizing the harness piece and the early-childhood series. Measuring work is not producing it. The garden of this axis is the measurement environment, not a classroom. SWE-bench, WebArena and AgentBench are treated as historical context, without fabricated rates; the objects with DOI in the corpus are read as environments or as critiques of the instrument, not as medals. Naming the bank is not having evaluated. Using the bank with a protocol is.
7. Four tests of support (not an artifact)
The framework that follows is a design inference of this article, anchored in the axes and in the verified sources. It is not an ISO standard or a leaderboard. It distinguishes four tests. If a platform, a paper or a product does not pass them, it cannot declare that the demo, the screenshot, the single score, the video, the chatbot labeled agent or the ranking without a protocol constitute evaluation of an agent.
7.1. Test of typed tasks and an oracle external to the LLM, not of the demo nor of the video nor of the model that evaluates itself. Zhuang et al. (2023) require tools and a criterion outside the context; Trivedi et al. (2024) expose app state; Deng et al. (2023) and Koh et al. (2024) anchor web action, including visual; Lee et al. (2025) anchor security; Xie et al. (2024) anchor the real computer. If the “evidence” is that the model wrote “done” or that a human cut a video, there is theater of eval, not eval.
7.2. Test of fixed budgets and logging of reproducible traces, not of the unique run nor of the opaque ranking. Yang et al. (2024) situate the ACI as a trail; Xie et al. (2024) require a bounded episode; Guo et al. (2024) require stability; Shinn et al. (2023) are reread in the negative: verbal reflection does not replace the log. There is eval when another laboratory can reconstruct steps, tokens, time, interface and trace.
7.3. Test of a failure taxonomy and analytical decomposition, not of the unique score. Ma et al. (2024) decompose the multi-turn episode; Xi, Ding et al. (2025) taxonomize environments; Aleithan (2025) decomposes the instrument; Patil et al. (2024) name API hallucination, which an aggregate number hides. A laboratory that only reports a percentage has not evaluated the how.
7.4. Test of category distinction and of eval design judgment, not of the product catalog nor of the environment trophy. Eval ≠ harness ≠ RAG ≠ fine-tune ≠ agent marketing ≠ leaderboard without a protocol ≠ perplexity ≠ demo ≠ video. Wang et al. (2024), Xi, Chen et al. (2025) and Liu et al. (2026) prevent reducing the field to a slogan. Park et al. (2023), Li et al. (2023), Qian et al. (2024), Schick et al. (2023) and Shen et al. (2023) are architecture or capability, not a license to omit the protocol. The work of evaluating is fulfilled by designing the coupling of tasks–oracle–budget–taxonomy–trace and declaring which piece is missing.
The framework admits tool-use, reflection, orchestration, APIs, roles and memory as pieces (Schick et al., 2023; Shinn et al., 2023; Shen et al., 2023; Patil et al., 2024; Qian et al., 2024; Park et al., 2023). It refuses to declare eval by any of them alone. The four tests are read together: passing one and failing the others is, again, an artifact.
8. Discussion
Three tensions organize the discussion. The first is between exhibiting the artifacts and exercising the craft of eval. Schick et al. (2023), Shen et al. (2023), Shinn et al. (2023), Patil et al. (2024) and Li et al. (2023) sustain real components —tools, orchestration, reflection, APIs, communication— that the market inflates until they pass as measurement. Ma et al. (2024), Guo et al. (2024), Yang et al. (2024) and Aleithan (2025) sustain the contrast: analytical board, stability, interface, data audit. Inference: the piece is true in its domain; “piece = agent eval” is an illegitimate category inference. The ranking screenshot is the gesture that best summarizes that illegitimacy: it inherits the form of science and none of its protocol obligations.
The second is between producing work and measuring it. The 16 September harness and the eval of this text need each other and do not substitute for each other. Yang et al. (2024) show that the ACI changes possible work; Xie et al. (2024) and Trivedi et al. (2024) show environments with observable state; Liu et al. (2026) recall that in software engineering the agent is played in the system. Omitting the harness and exhibiting a ranking is a bias; omitting eval and exhibiting a runtime is the symmetric bias. Xi, Chen et al. (2025) and Wang et al. (2024) describe the rise of more capable agents: that rise does not, by itself, sign either harness or eval. Reported empirical gain is not read honestly if it is attributed to a clip or a unique score and the protocol is hidden.
The third is between the failure mode and the instrument. Patil et al. (2024) name API hallucination; Koh et al. (2024) and Deng et al. (2023) distinguish perception error from plan error; Lee et al. (2025) situate security as a domain in which “it looks solved” does not suffice; Aleithan (2025) situates the error in the data. Ma et al. (2024) and Xi, Ding et al. (2025) prevent compacting that heterogeneity; Guo et al. (2024) recall that the ranking itself can be unstable. Without oracle, taxonomy and audit, the error is not seen or is attributed to the wrong place. Park et al. (2023) and Qian et al. (2024) illustrate that memory and roles can rehearse the same error with more eloquence.
The four tests read these tensions. A caller, a role or an invocation can live inside an evaluated system; it is not to invert the sequence: first the announcement of an “evaluated agent,” then the hope that the model will verify itself (Xi, Chen et al., 2025; Wang et al., 2024; Schick et al., 2023; Patil et al., 2024; Qian et al., 2024). If there are typed tasks, oracle, budget, trace, taxonomy and design judgment, there is eval; if there is only a demo, a screenshot, a single-run score or a video, there are artifacts. Zhuang et al. (2023) and Xie et al. (2024) recall that the criterion lives outside: in the tool and on the desktop, not in the model’s own applause.
9. Limits
This review is narrative. It does not apply its own PRISMA nor estimate combined effects. No leaderboard percentages, N, d, r or AUC is invented. Sources on tool-use, reflection, orchestration and APIs are read by the object they propose, not as trials of “production eval” (Schick et al., 2023; Shinn et al., 2023; Shen et al., 2023; Patil et al., 2024; Li et al., 2023). Yang et al. (2024) anchor interface, without generalizing to every domain. The environments and datasets (Deng et al., 2023; Koh et al., 2024; Trivedi et al., 2024; Xie et al., 2024; Lee et al., 2025; Zhuang et al., 2023) have marked scenario transfer. AgentBoard, AgentGym and StableToolBench are board, multi-environment and stability, not a trial of the artifact ceiling (Ma et al., 2024; Xi, Ding et al., 2025; Guo et al., 2024). Park et al. (2023) and Qian et al. (2024) are architecture and roles; Wang et al. (2024), Xi, Chen et al. (2025) and Liu et al. (2026) are surveys; Aleithan (2025) is data quality of a specific bank.
No trials were located that equate a demo, a screenshot, a single-run score or a video without a trace to an eval with oracle, budget, taxonomy and reproducibility; they are discussed as a category ceiling. Commercial products are not evaluated. The inferences of section 7 are design hypotheses, not evidence of a particular runtime. The “garden” of heading 5 names the measurement environment; it does not transfer early-childhood findings. SWE-bench, WebArena and AgentBench appear as historical context, without invented rates or DOI. This text does not re-demonstrate that the harness is not the model: it demonstrates that eval is neither the harness nor the poster.
10. Conclusions
A cherry-picked demo, a leaderboard screenshot, a single-run score or a video of “the agent solved X” without a fixed evaluation harness do not constitute evaluation of an agent. Neither do a chatbot labeled agent, RAG, a fine-tune, a ranking without a protocol or perplexity. Schick et al. (2023), Shinn et al. (2023), Shen et al. (2023), Patil et al. (2024) and Li et al. (2023) confirm components without equivalence to eval. Wang et al. (2024), Xi, Chen et al. (2025) and Liu et al. (2026) fix the map. When there is eval there is craft: tasks and an external oracle (Zhuang et al., 2023; Trivedi et al., 2024; Deng et al., 2023; Koh et al., 2024; Lee et al., 2025); budgets, traces and interface (Xie et al., 2024; Yang et al., 2024; Guo et al., 2024); taxonomy, not a unique score (Ma et al., 2024; Xi, Ding et al., 2025; Aleithan, 2025); architecture and roles declared (Park et al., 2023; Qian et al., 2024). Eval is distinguished from the harness, from the LLM alone, from the demo and from the slogan.
Where the sources do not measure a production runtime, this article does not claim one. Where they measure tools, environments, boards, stability, data quality or surveys, it does not translate them into “the agent is already evaluated” via a clip or a screenshot. Distinguishing demo capability from evidence of work is exercising typed tasks, external oracles, budgets, a failure taxonomy, trace logging and reproducibility, with a criterion that can contradict the model. The rest is a cherry-picked demo, a screenshot, a single-run score and a video without a trace. It is not evaluation, and it should not be presented as what it is not. The harness produces work; eval says whether that work is reliable. Confusing them is the gesture this text refuses to sign.
Laboratorio Editorial de NEXTECH.IA / Ingeniero Mitre.
References
- Aleithan, R. (2025). Revisiting SWE-Bench: On the importance of data quality for LLM-based code models. En 2025 IEEE/ACM 47th International Conference on Software Engineering: Companion Proceedings (ICSE-Companion). https://doi.org/10.1109/icse-companion66252.2025.00075
- Deng, X., Gu, Y., Zheng, B., Chen, S., Stevens, S., Wang, B., Sun, H. y Su, Y. (2023). Mind2Web: Towards a generalist agent for the web. En Advances in Neural Information Processing Systems 36. https://doi.org/10.52202/075280-1220
- Guo, Z., Cheng, S., Wang, H., Liang, S., Qin, Y., Li, P., Liu, Z., Sun, M. y Liu, Y. (2024). StableToolBench: Towards stable large-scale benchmarking on tool learning of large language models. En Findings of the Association for Computational Linguistics ACL 2024 (pp. 11143-11156). https://doi.org/10.18653/v1/2024.findings-acl.664
- Koh, J. Y., Lo, R., Jang, L., Duvvur, V., Lim, M., Huang, P. Y., Neubig, G., Zhou, S., Salakhutdinov, R. y Fried, D. (2024). VisualWebArena: Evaluating multimodal agents on realistic visual web tasks. En Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (pp. 881-905). https://doi.org/10.18653/v1/2024.acl-long.50
- Lee, H., Zhang, Z., Lu, H. y Zhang, L. (2025). SEC-bench: Automated benchmarking of LLM agents on real-world software security tasks. En Advances in Neural Information Processing Systems 38. https://doi.org/10.52202/085713-3878
- Li, G., Hammoud, H., Itani, H., Khizbullin, D. y Ghanem, B. (2023). CAMEL: Communicative agents for “mind” exploration of large language model society. En Advances in Neural Information Processing Systems 36. https://doi.org/10.52202/075280-2264
- Liu, J., Wang, K., Chen, Y., Peng, X., Chen, Z., Zhang, L. y Lou, Y. (2026). Large language model-based agents for software engineering: A survey. ACM Transactions on Software Engineering and Methodology. https://doi.org/10.1145/3796507
- Ma, C., Zhang, J., Zhu, Z., Yang, C., Yang, Y., Jin, Y., Lan, Z., Kong, L. y He, J. (2024). AgentBoard: An analytical evaluation board of multi-turn LLM agents. En Advances in Neural Information Processing Systems 37. https://doi.org/10.52202/079017-2365
- Park, J. S., O’Brien, J., Cai, C. J., Morris, M. R., Liang, P. y Bernstein, M. S. (2023). Generative agents: Interactive simulacra of human behavior. En Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology (pp. 1-22). https://doi.org/10.1145/3586183.3606763
- Patil, S., Zhang, T., Wang, X. y Gonzalez, J. (2024). Gorilla: Large language model connected with massive APIs. En Advances in Neural Information Processing Systems 37. https://doi.org/10.52202/079017-4020
- Qian, C., Liu, W., Liu, H., Chen, N., Dang, Y., Li, J., Yang, C., Chen, W., Su, Y., Cong, X., Xu, J., Li, D., Liu, Z. y Sun, M. (2024). ChatDev: Communicative agents for software development. En Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (pp. 15174-15186). https://doi.org/10.18653/v1/2024.acl-long.810
- Schick, T., Dwivedi-Yu, J., Dessi, R., Raileanu, R., Lomeli, M., Hambro, E., Zettlemoyer, L., Cancedda, N. y Scialom, T. (2023). Toolformer: Language models can teach themselves to use tools. En Advances in Neural Information Processing Systems 36. https://doi.org/10.52202/075280-2997
- Shen, Y., Song, K., Tan, X., Li, D., Lu, W. y Zhuang, Y. (2023). HuggingGPT: Solving AI tasks with ChatGPT and its friends in Hugging Face. En Advances in Neural Information Processing Systems 36. https://doi.org/10.52202/075280-1657
- Shinn, N., Cassano, F., Gopinath, A., Narasimhan, K. y Yao, S. (2023). Reflexion: Language agents with verbal reinforcement learning. En Advances in Neural Information Processing Systems 36. https://doi.org/10.52202/075280-0377
- Trivedi, H., Khot, T., Hartmann, M., Manku, R., Dong, V., Li, E., Gupta, S., Sabharwal, A. y Balasubramanian, N. (2024). AppWorld: A controllable world of apps and people for benchmarking interactive coding agents. En Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (pp. 16022-16076). https://doi.org/10.18653/v1/2024.acl-long.850
- Wang, L., Ma, C., Feng, X., Zhang, Z., Yang, H., Zhang, J., Chen, Z., Tang, J., Chen, X., Lin, Y., Zhao, W. X., Wei, Z. y Wen, J. (2024). A survey on large language model based autonomous agents. Frontiers of Computer Science, 18(6). https://doi.org/10.1007/s11704-024-40231-1
- Xi, Z., Chen, W., Guo, X., He, W., Ding, Y., Hong, B., Zhang, M., Wang, J., Jin, S., Zhou, E., Zheng, R., Fan, X., Wang, X., Xiong, L., Zhou, Y., Wang, W., Jiang, C., Zou, Y., Liu, X., Yin, Z., Dou, S., Weng, R., Qin, W., Zheng, Y., Qiu, X., Huang, X., Zhang, Q. y Gui, T. (2025). The rise and potential of large language model based agents: A survey. Science China Information Sciences, 68(2). https://doi.org/10.1007/s11432-024-4222-0
- Xi, Z., Ding, Y., Chen, W., Hong, B., Guo, H., Wang, J., Guo, X., Yang, D., Liao, C., He, W., Gao, S., Chen, L., Zheng, R., Zou, Y., Gui, T., Zhang, Q., Qiu, X., Huang, X., Wu, Z. y Jiang, Y. G. (2025). AgentGym: Evaluating and training large language model-based agents across diverse environments. En Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (pp. 27914-27961). https://doi.org/10.18653/v1/2025.acl-long.1355
- Xie, T., Zhang, D., Chen, J., Li, X., Zhao, S., Cao, R., Hua, T., Cheng, Z., Shin, D., Lei, F., Liu, Y., Xu, Y., Zhou, S., Savarese, S., Xiong, C., Zhong, V. y Yu, T. (2024). OSWorld: Benchmarking multimodal agents for open-ended tasks in real computer environments. En Advances in Neural Information Processing Systems 37. https://doi.org/10.52202/079017-1650
- Yang, J., Jimenez, C., Wettig, A., Lieret, K., Yao, S., Narasimhan, K. y Press, O. (2024). SWE-agent: Agent-computer interfaces enable automated software engineering. En Advances in Neural Information Processing Systems 37. https://doi.org/10.52202/079017-1601
- Zhuang, Y., Yu, Y., Wang, K., Sun, H. y Zhang, C. (2023). ToolQA: A dataset for LLM question answering with external tools. En Advances in Neural Information Processing Systems 36. https://doi.org/10.52202/075280-2180