Situated Action and Affordance Theory Applied to AI Evaluation
ChatGPT Deep Research report (2026-03-23)
Query Has anyone applied Sellen & Harper's affordance analysis (from "The Myth of the Paperless Office," 2002) or Suchman's situated action framework (from "Plans and Situated Actions," 1987) to the question of AI capability evaluation or AI benchmarking? Are there papers arguing that AI benchmarks fail because they strip tasks of their situated context — the practices, artifacts, and feedback loops that surround real-world use? Search in HCI, CSCW, STS, information science, and AI evaluation literature. Also check: has anyone connected "The Myth of the Paperless Office" to AI adoption patterns? The structural parallel is direct — paper persisted where digital tools couldn't replicate paper's affordances, just as human work persists where AI can't replicate the extensions surrounding human practice. Has anyone made this connection explicitly? --- Full Report Situated Action and Affordance Perspectives on AI Evaluation and Benchmarking Research question and scope This report investigates whether (a) Suchman’s situated action framework and (b) Sellen & Harper’s affordance analysis of paper have been applied explicitly to AI capability evaluation, AI benchmarking, or closely related evaluation practices (e.g., sociotechnical “readiness,” in-the-wild assessment, auditing, and adoption). It also searches for scholarship making the structural analogy explicit: paper persisted where digital tools did not replicate paper’s affordances, and (by parallel) human work persists where AI evaluation/benchmarks do not capture the surrounding practices, artifacts, and feedback loops that constitute real-world capability. Evidence below is drawn from AI evaluation/benchmarking scholarship, human-centered AI and HCI evaluation work, and STS/critical studies of benchmarks and ground truthing. Key findings: • Yes, there are AI-evaluation papers that (1) explicitly cite situated action (Suchman) and (2) argue for ecological validity / “messy” real-world context rather than stripped-down benchmark tasks—especially in work on agentic workflows and interaction-based evaluation. • Yes, there is robust STS work showing that benchmarks and “ground truths” are produced through standards, alignment work, and infrastructures—casting benchmarks as sociotechnical constructions rather than neutral measures. • For “The Myth of the Paperless Office” → AI adoption patterns, there are explicit, high-signal connections in recent discourse and research (including a major U.S. government AI report and a detailed public interview where paperless-office lessons are applied directly to AI workflow injection/adoption). • Direct, explicit application of Sellen & Harper’s specific affordance analysis as a method for AI benchmark design appears uncommon in peer‑reviewed AI benchmarking literature; instead, the connection is more visible in (i) AI adoption and workflow critiques, and (ii) paper/digital-work scholarship that is now being extended with LLMs. Explicit uses of situated action ideas in AI evaluation and benchmarking A strong “direct hit” is recent work on evaluating LLMs (and multi-agent “debate” workflows) that explicitly invokes situated action as an argument for ecological validity. In “Deliberative Dynamics and Value Alignment in LLM Debates” (2026 preprint), the authors argue that sociotechnical alignment is often studied with single-turn, static evaluations, but agentic deployment turns behavior into multi-turn workflows; they then state that ecological validity is paramount, citing that reasoning and decision-making are context-dependent and emerge in response to circumstances (Suchman, 1987), and that evaluation should reflect “messy, everyday” conditions. This is an unusually explicit bridge: “situated action” is named as a rationale for changing evaluation design. In HCI-centered AI evaluation, a recurring pattern is shifting evaluation away from decontextualized model outputs and toward interaction- and lifecycle-based metrics. “From Accuracy to Readiness: Metrics and Benchmarks for Human–AI Decision-Making” (CHI EA 2026) proposes evaluating human–AI decision-making through interaction traces and an onboarding lifecycle (calibration, error recovery, governance), explicitly positioning this as a move from evaluating models in isolation to evaluating socio-technical teams. While not framed as “situated action theory,” its operational target (interaction over time, governance enacted in practice) overlaps strongly with situated-action critiques of plan-like evaluation abstractions. A closely aligned evaluation critique in AI safety/auditing is the argument that prevailing evaluations focus on foundation model capabilities in isolation while real deployments are composite, interactive, and feedback-driven. In “AI Safety Evaluations Need To Consider Cascading Effects” (2026), and argue that AI systems are “algorithmic supply chains” whose behavior is mediated by many components (system prompts, guardrails, tool use, personalization, etc.), and that audits focusing on models alone (or components independently) miss how components interact and compound downstream effects, calling for systems-oriented audits that incorporate such cascades. This is not a citation of Suchman, but it directly targets the same failure mode the user describes: benchmarks strip away the surrounding socio-technical assemblage (components + stakeholder practices + feedback loops). Finally, beyond “situated action” proper, there are new evaluation papers explicitly diagnosing benchmark/context mismatch as a cultural and situational problem. “Who Benchmarks the Benchmarks? A Case Study of LLM Evaluation in Icelandic” (2026) argues that translating culturally specific English benchmarks can reward English‑centric assumptions rather than culturally situated understanding in the target language. This is another concrete instantiation of the “stripping context” argument—here, the missing context is linguistic/cultural practice rather than workflow artifacts. Benchmarks as de-situating abstractions in AI and HCI evaluation Several recent benchmark-design papers argue directly that mainstream benchmark construction methods lack ecological validity because they generate tasks and questions outside the real situation of use. A particularly on-point example is “‘Is This It?’: Towards Ecologically Valid Benchmarks for Situated Collaboration” (ICMI Companion 2024). The paper contrasts common benchmark construction (post hoc question–answer generation over preexisting/synthetic datasets) with an “interactive system-driven” approach where questions are generated by users in context during interaction with an end-to-end situated AI system. It argues that current benchmarks do not capture the kinds of questions users ask in real-time tasks, leading to mismatch between measured performance and user experience. Crucially, it proposes iterative cycles: organize real interaction questions into a dataset, train/evaluate, then re-integrate into the original application and test again in live interactions, surfacing new questions in cycles of refinement. This is extremely close to the user’s stated concern about missing “practices, artifacts, and feedback loops.” A second explicit benchmark-design response is “Towards Ecologically Valid LLM Benchmarks: Understanding and Designing Domain-Centered Evaluations for Journalism Practitioners” (2025). It argues that benchmark scores are widely used to communicate capability, but they face criticisms about construct validity (whether the benchmark measures what it claims) and ecological validity (whether evaluation represents how models are used “in the wild”). The authors propose a human-centered, practitioner-engaged process (workshop with journalism professionals) to design domain-oriented benchmarks tied to real professional values, use scenarios, and contextual variability (“Would a model act differently or fail in a slightly different context?”). This is a direct, peer‑review adjacent attempt to rebuild benchmarking around situated professional practice. Beyond benchmark construction, evaluation scholarship in domains like healthcare has also emphasized that “bench performance” (model-centric evaluation) is not well suited to systems embedded in clinical settings and calls for “situated evaluation” and co-creation with end users. This line reinforces the broader claim—across HCI and applied AI—that evaluation must attend to the real “combination space” of humans + tools + institutions, not just a model’s isolated output. STS accounts of benchmarks and ground truths as sociotechnical constructions The user’s hypothesis that benchmarks fail by stripping context is strongly supported by STS/critical studies showing that benchmarks and “ground truths” are not mere representations; they are instituted through practices, infrastructures, and negotiated standards. In “Agreements ‘in the wild’: Standards and alignment in machine learning benchmark dataset construction” (2024), presents an ethnographic case study of a corporate–academic group constructing a benchmark dataset of daily activities. The paper conceptualizes the dataset as a socio-technical “knowledge object” stabilized via standards and “alignment work,” and—most importantly for your question—calls into question the claim that the dataset captures unscripted activities occurring naturally “in the wild,” because alignment work bleeds into data capture (guidance on how to record, negotiated parameters, etc.). This is a precise empirical demonstration of the claim that “in-the-wild” benchmark rhetoric can mask the situated labor that makes a benchmark possible. In “Groundwork for AI: Enforcing a benchmark…” (2023), analyzes the Tumor nEoantigen SeLection Alliance (TESLA) ground-truthing project as a benchmark-making endeavor and argues that ground truth datasets establish what can be retrieved computationally and evaluated statistically; the “truth” depends on a broad technoscientific network and infrastructures, and the project produced a contestable reference rather than an indisputable truth, while still instituting a benchmark that enabled AI development in that field. This aligns tightly with situated-action sensibilities: evaluation artifacts stabilize and coordinate action, but do not “contain” the world they claim to represent. Finally, “AI and the Everything in the Whole Wide World Benchmark” (2021) argues that AI subfields valorize influential benchmarks that become stand-ins for broad “general” progress; it critiques the construct validity of such benchmarks and emphasizes that these measures are inherently specific, finite, and contextual—despite being framed as general. The paper has become a cross-disciplinary anchor for later benchmark critique and review work. Together, these STS and critical benchmark studies offer a strong empirical and conceptual base for your claim: benchmarks are not merely “missing context”; they participate in redefining the task world by selecting what counts as the task, the relevant artifacts, and the acceptable feedback channels. “The Myth of the Paperless Office” and explicit links to AI adoption patterns The most explicit, direct connection found between paper-affordance arguments and AI adoption has appeared in (a) reflective/governmental AI discussions and (b) public HCI research discourse—more than in mainstream benchmark papers. A particularly strong explicit link is in a 2024 interview transcript hosted by. In that conversation, the host explicitly asks (paraphrasing) why “in the age of AI” paper still matters, referencing The Myth of the Paperless Office. Sellen answers by emphasizing that the book argued for understanding the work paper does in context and warns that if organizations “inject AI into workflows” or replace parts of workflows without deep understanding of how work is currently done and what value workers bring, it can reproduce chaos similar to poorly executed “paperless” transitions. This is almost exactly the structural parallel you describe—made explicit by one of the coauthors, in an AI-era reflection. A second explicit connection is in a major U.S. government report on AI and research & development. In a passage discussing the productivity effects of AI writing tools, the report explicitly references a “version of the Jevons paradox (or the ‘myth of the paperless office’)”: increased writing efficiency may be converted into more writing produced, rather than time savings. This is not about benchmarking per se, but it is a clear policy-level articulation of the same kind of rebound/translation effect that the paperless-office literature made legible. Peer-reviewed and preprint HCI work that cites The Myth of the Paperless Office is also now appearing in LLM-augmented document and reading systems. For example, RealitySummary (mixed-reality reading assistants combining OCR and LLMs) situates itself in a long-standing consensus that people prefer physical paper over screens and explicitly cites Sellen & Harper in that context. Even when the citation is used primarily to justify continued attention to paper/digital hybridity, it functions as a bridge: LLM systems are being positioned as an attempt to add capabilities without forcing “paperless” replacement. Relatedly, a 2025 seminar abstract by Richard Harper (hosted by ) critiques the “abstraction” between users and LLMs and argues that chatbot-based interaction is built around abstractions that are not thought through, leading to misunderstandings of what LLM outputs mean and how they should be used. While not an adoption study, it explicitly frames LLM use problems as interaction-design/meaning-making problems—consistent with paper-affordance and situated-practice sensibilities. Finally, there are additional signals that the paperless-office analogy is entering workplace/AI discourse in HCI-for-work venues (e.g., papers explicitly referencing “ChatGPT and Co-pilot” alongside The Myth of the Paperless Office in their bibliographies), though access to full texts is uneven across venues. Implications for AI evaluation design through a situated-action and affordance lens The literature above converges on a specific diagnosis: many benchmarks behave like plans—clean specifications with stable inputs/outputs—while real AI capability emerges (or fails) through situated action: interaction histories, local contingencies, artifacts that externalize cognition, organizational policies, selective trust, contestation, and repair. The most explicit version of this argument in an LLM-evaluation paper is the claim that evaluation should reflect “messy, everyday” contexts because reasoning and decision-making are circumstance-dependent. Across sources, a situated/affordance-informed AI evaluation agenda has several recurring building blocks: First, evaluation objects shift from “model” to “socio-technical assemblage.” Cascading-effects work argues that deployed systems are supply chains of components and stakeholders; evaluating an FM alone misses interactions among prompts, guardrails, tools, personalization, and organizational practices. Human–AI decision-making metrics similarly argue that accuracy is insufficient, and evaluation should incorporate reliance behaviors, calibration, harms, learning over time, and governance as enacted in use. Second, benchmarks should incorporate the artifacts and feedback loops that constitute real work. Interactive system-driven benchmarking for situated collaboration explicitly proposes cyclic evaluation: collect in-context questions during actual task performance, build a benchmark from those questions, then reintegrate models back into the system and test again—turning benchmarking into an iterative sociotechnical loop rather than a one-shot test set. The journalism benchmark work similarly insists that evaluation must represent actual domain practice, including contextual variability and the procedural know-how that professionals deploy to handle that variability. Third, “in the wild” claims should be treated as empirically contestable. STS studies show that “in-the-wild” benchmark datasets can depend on alignment work that shapes what is recorded, how it is recorded, and what counts as compliant data, undermining naïve assumptions that “wildness” automatically implies ecological validity. Fourth, the paperless-office analogy suggests a practical evaluative heuristic: when a new system claims replacement, look for the “affordance debt” it incurs—the often invisible functions done by incumbent artifacts (paper, human collaboration, tacit routines). In the AI-era interview, Sellen explicitly warns that inserting AI without understanding what the existing workflow “does for people” can reproduce the failure dynamics of earlier “paperless office” hype cycles. The U.S. government report’s invocation of the “myth of the paperless office” likewise highlights that seemingly straightforward productivity gains can translate into changed work volumes and practices rather than simple substitution. Gaps and what has and has not been done explicitly The search suggests an asymmetric picture: • Suchman → AI evaluation: There are now clear, explicit instances where situated-action ideas are used to motivate ecologically valid, multi-turn, workflow-like evaluations (e.g., LLM deliberation/agentic alignment work). There is also broader HCI evaluation work that reframes measurement around interaction and lifecycle rather than model scorecards. • Sellen & Harper → AI benchmarking: Direct application of the specific affordance analysis method (as “benchmark design method”) is harder to find in the benchmarking canon; instead, the connection appears more often in AI adoption/workflow discussions and in document-centric HCI systems that cite paper persistence as a constraint and design premise. • Paperless-office analogy made explicit: Yes—explicitly—in (i) an AI-era reflection by Sellen and (ii) a U.S. government report that uses “myth of the paperless office” as a named analogy for AI productivity impacts. • Benchmarks fail because tasks are stripped of situated context: Yes—explicitly—in multiple benchmark-design and evaluation-critique strands: interactive system-driven benchmarking argues post hoc QA generation lacks ecological validity; domain-centered benchmarks critique construct/ecological validity; STS work shows benchmark construction “bleeds” into capture and stabilizes contestable references; AI safety evaluation work argues for holistic assessment of component interactions rather than isolated capability probes.