The Multidimensional Frontier
ChatGPT Deep Research report (2026-03-23)
Query The jagged frontier (Mollick et al., 2023) is measured along one axis: task output quality. But when AI is introduced into a practice, the output is not the only thing that changes. The work itself changes — the skills involved, the pace, the relationship between practitioners, the nature of expertise, what newcomers learn, what gets valued. Search for literature on how AI adoption changes work practices beyond output metrics. Not "will AI replace jobs" — that's a different question. Rather: what happens to the practice when AI becomes part of it? Specific threads to investigate: • Deskilling and reskilling: When AI handles tasks that newcomers used to learn from (writing boilerplate code, drafting initial briefs, taking clinical notes), what happens to skill formation? Look for literature connecting AI adoption to legitimate peripheral participation (Lave & Wenger, 1991) — if the peripheral tasks are automated, how do newcomers learn? • Changes to craft and professional identity: When a programmer's work shifts from writing code to reviewing and prompting, when a lawyer's work shifts from drafting to checking AI drafts — how do practitioners describe the change? Not just efficiency gains but changes in their relationship to the work. • What success metrics miss: The handoff analytic (Mulligan & Nissenbaum, 2020) argues that functional equivalence does not equal ethical equivalence — things change in the reconfiguration even when the output looks equivalent. Has anyone applied this framework to AI-in-the-workplace evaluation? • Constructed experience (Barrett): Lisa Feldman Barrett's work on constructed emotion suggests that the act of externalizing and labeling an experience participates in constructing it. When an LLM helps you articulate what you think or feel, it's not just reflecting — it's participating in the construction. Has anyone connected Barrett's framework to AI interaction? The hypothesis: the jagged frontier is multidimensional. The frontier that Mollick measures (output quality) is one axis. There are other axes — skill formation, craft identity, relationship to the work, what newcomers learn, how the practice reproduces itself — and the frontier may be jagged on those axes in completely different places than on the output axis. --- Full Report The Multidimensional Frontier: How AI Changes Work Practices Beyond Output Metrics Why output-only frontiers are incomplete Empirical work on AI-as-assistant in knowledge work has productively operationalized a "jagged frontier" by measuring task performance (for example, output quality and throughput) under AI assistance. Yet a large, older stream of organizational research argues that introducing a technology rarely produces "same work, better output"; instead, it reconfigures the practice: who does what, in what sequence, with what forms of coordination, authority, and accountability. citeturn6search1turn6search2 Two canonical lenses make this point in ways that map directly onto AI adoption. First, studies of technological change in organizations show that "identical" technologies can occasion divergent role relations and work structures depending on local arrangements and interpretations—meaning outcome metrics can look similar while the social order of work becomes different. citeturn6search1turn6search17 Second, a "practice lens" emphasizes technology-in-use: structures are enacted through recurrent human–tool interaction, not embedded in the artifact alone. citeturn6search2turn6search6 Recent qualitative theorizing on AI specifically pushes the same direction: treating AI not as a stable "thing" but as phenomena "in-the-making" foregrounds practices, enactments, and how agency and accountability become distributed across models, data pipelines, infrastructures, and workplace routines. This framing implies that "frontiers" are plural: a practice may cross the output frontier (high-quality deliverables) while remaining jagged on other axes (learning, responsibility allocation, professional identity, relational trust, or organizational pace). Skill formation when AI automates the periphery A key practice-level concern is not whether AI "replaces jobs," but what happens to skill formation pathways when AI absorbs tasks that historically served as entry points for novices (drafting, boilerplate code, first-pass analyses, routine documentation). In situated learning terms, many occupations rely on "legitimate peripheral participation," where novices learn through low-stakes peripheral tasks, observation, and gradual increases in responsibility. citeturn1search6 Automating the periphery risks altering the apprenticeship gradient: novices may produce acceptable artifacts sooner, but with weaker internalization of tacit judgment, tool sense, and error-detection skills. turn9search0 Evidence of this pattern is emerging across domains, often described as flattening (rapid performance gains) coupled with fragility (weaker underlying understanding). A study of undergraduate programming with AI code generation found substantial improvements in completion speed and correctness during assisted work, while also capturing learner concerns about missing explanations ("I never got an explanation on why…") and the social substitution effect ("solve my own problems without having to turn to someone"). In a separate controlled study of computing students doing "brownfield" work in an unfamiliar codebase, AI assistance shifted activity away from manual code writing and even away from web search—an empirical sign that the learning ecology of programming (searching, reading, comparing sources, synthesizing) is being reconfigured, not simply accelerated. In professional settings, AI can also change who coaches whom. Work on "novice risk work" reports a case where junior professionals—often expected to be early adopters and translators of new technologies—may fail when the technology's capabilities are uncertain and rapidly changing, leading to organizational learning failures rather than smooth diffusion. Complementary research on expertise and new technologies documents "inverted apprenticeships," where senior experts structure interactions with juniors to develop new technological expertise while preserving status—often at a cost to junior learning and to the traditional expert–novice developmental pipeline. citeturn23search1turn23search6 In medicine, the deskilling/reskilling question is now explicit. A mixed-method review of "AI-induced deskilling" in medical literature links concerns to over-reliance, automation bias, reduced critical thinking, and declining judgment, and frames these as issues for training and system resilience—not merely near-term performance. A more recent clinical-education argument distinguishes deskilling from "never-skilling," warning that if trainees rely on automation early, they may never acquire foundational reasoning and procedural competence in the first place. The "GenAI wall effect" is also relevant to skill formation because it suggests that AI does not erase expertise gradients uniformly: in a field experiment at a firm, GenAI helped "adjacent outsiders" more than "distant outsiders," and helped conceptualization more than execution—implying that where novices can "skip steps" depends on task type and knowledge distance. If organizations redesign roles assuming AI universally bridges gaps, they may unintentionally create developmental dead-ends for novices in some job families while accelerating others. Craft, expertise, and professional identity under AI-mediated workflows Practice changes are often experienced as changes in craft: the felt sense of "doing the work," what counts as expertise, and what professionals take pride in. Multiple empirical studies suggest that AI introduces new categories of labor that are not captured by output quality scores: prompt formulation, context packaging, iterative steering, verification, error triage, and accountability work. In software development, grounded observational work on AI code assistants characterizes distinct usage modes (for example, accelerating familiar work vs. exploring unfamiliar APIs), indicating that "writing code" becomes interleaved with "sampling, selecting, and integrating" model outputs—closer to editorial control than direct authorship in many moments. In data analysis, participatory prompting studies similarly report benefits to information foraging and sensemaking while introducing barriers around query formulation, specifying context, and verification—again emphasizing that the locus of difficulty moves, rather than disappearing. Workplace studies focused on role redesign reinforce that "success" can mean role expansion, not time savings. Recent ethnographic findings summarized in Harvard Business Review report that generative AI adoption can intensify work by increasing pace, expanding task scope, and extending work into more hours of the day—even without mandates—because AI reduces the friction of initiating tasks and makes "doing more" feel feasible and rewarding. This is a practice-level transformation: downtime, sequencing, and boundary norms shift, even if output quality improves. Professional identity shifts are especially visible when the practitioner becomes a reviewer of AI drafts. In a study of product managers delegating work to GenAI, participants describe a new "0–80%" drafting rhythm ("go from 0–80%… much faster") and tool-mediated gatekeeping ("run content through this tool before I review"), which changes both the manager's craft and the social organization of review. In education, phenomenological work on teachers' identity highlights how AI integration can prompt teachers to renegotiate autonomy, roles, and the meaning of expertise—again pointing to identity and values, not just efficiency. turn8search22 In high-stakes expert work, the practice change can be about judgment formation. Field research in diagnostic radiology examines how AI tools get incorporated (or not) into professionals' processes of making critical judgments, with ambiguity shaping whether AI becomes authoritative, advisory, or contested in practice. In law, empirical/doctrinal work on generative AI in legal practice emphasizes that adoption pressures collide with professional duties of competence and confidentiality—suggesting a shift toward "verification as craft" and risk management as routine professional work, not an occasional compliance step. citeturn8search13turn8search21 What success metrics miss: reconfiguration, accountability, and ethics A core reason output quality alone is insufficient is that AI often changes handoffs: who is responsible for what, how decisions travel across sociotechnical boundaries, and what values are embedded in the reconfiguration. citeturn10search6turn3search4 The "handoff model" articulated by Deirdre K. Mulligan and Helen Nissenbaum was developed to analyze why functional equivalence does not imply ethical equivalence when a function moves from humans to automated systems (or between institutions). citeturn3search12turn10search20 While the handoff model did not originate in workplace GenAI evaluation, it has been operationalized in adjacent domains precisely to show how "technical upgrades" shift participation, transparency, and accountability. For example, a FAccT paper on the U.S. Census Bureau's adoption of differential privacy explicitly uses the handoff lens to reveal value shifts and participation impacts that would not be visible through output-focused evaluation alone. citeturn10search6turn10search10 A recent policy-oriented governance model similarly invokes the handoff model to move risk thinking away from model-centric metrics toward a sociotechnical view of responsibility and control. citeturn3search4 Related accountability frameworks underscore a similar point: responsibility often "slides" toward the nearest human operator even when that operator has limited control over system behavior. Madeleine Clare Elish names this dynamic the "moral crumple zone," describing how complex automated systems can preserve the appearance of technological faultlessness while positioning humans as the absorbent layer for blame. turn10search0 In organizational contexts, this implies that "human-in-the-loop" designs can create new invisible burdens (monitoring, second-guessing, liability exposure) that are not captured by output metrics and may ultimately degrade safety and trust. Measurement research in human–AI decision-making is beginning to formalize what "output-only" misses. A recent framework proposes moving "from accuracy to readiness," emphasizing metrics spanning not only outcomes but also reliance behavior, safety signals, and learning-over-time—explicitly arguing that many failures arise from miscalibrated reliance rather than low average accuracy. citeturn10search18turn10academia41 This aligns with sociotechnical accountability work that treats accountability as relational and systemic, not reducible to model properties. citeturn10search14turn12view6 Finally, algorithmic impact assessment traditions—developed largely for public-sector algorithmic systems—offer templates for multidimensional evaluation: documenting objectives, stakeholders, risks, governance, and feedback loops, not just performance benchmarks. citeturn10search28turn10search32 These approaches are conceptually well-suited to AI-in-practice settings because they treat "system success" as a bundle of technical, social, and institutional properties. citeturn10search32turn3search4 The email example: when "better writing" changes the social meaning of communication Email is a useful microcosm because it looks superficially like a pure output task ("write a clearer email"), but it is deeply relational: it carries signals about effort, status, intent, and authenticity. turn7search36 The AI-mediated communication (AI-MC) framework explicitly treats AI as acting "on behalf of" a communicator by modifying or generating messages to accomplish interpersonal goals—shifting analysis from message quality to how mediation changes social outcomes and norms. citeturn7search36turn17view3 Empirically, several "practice-level" effects show up even when messages are fluent and professional. Linguistic comparison work finds that LLM-generated emails tend to be more formal, verbose, and complex, while human emails are more concise and personalized; it also reports user ambivalence that explicitly targets the relationship between writer and words (e.g., "most complete… but… not… pleasant," and concerns about impersonality). This is not merely "quality critique"—it is evidence that AI shifts genre norms and perceived voice, which changes how recipients interpret intent and social distance. turn7search36 Social perception studies add a crucial asymmetry: when AI authorship is disclosed or strongly suspected, people often impose an "AI penalty" (harsher judgments), but under realistic uncertainty they may not suspect AI at all. "Blissful (A)Ignorance" findings show that impressions can remain positive (and similar to fully human-written conditions) when recipients are uninformed about possible AI use—implying that AI can change the social economy of signaling by lowering the cost of producing high-polish messages without reliably updating recipient inference. turn7search0 In market-facing communication, experiments on the "AI-authorship effect" show that messages believed to be AI-generated are perceived as less authentic and can provoke negative emotional responses (e.g., moral disgust), reducing loyalty intentions. turn7search1 In interpersonal contexts, experimental work using predictive text assistance in trust-game communication finds that AI assistance can increase efficiency while having minimal effects on behavioral and self-reported trust—suggesting that some relational outcomes may be surprisingly robust, but also that "trust" is not a single dimension and may depend on disclosure norms and context. citeturn7search29turn7search3 Work in AI-mediated social support illustrates a related "practice transformation" mechanism: different patterns of human–AI collaboration (AI-only, modified-AI, AI-guided) systematically change message features such as emotional support, self-disclosure, and perceived authenticity. The important point for a multidimensional frontier is that "good output" (helpful, clear messages) can coexist with changing norms about authorship, disclosure, reciprocity, and what counts as a sincere interpersonal act. citeturn7search36turn17view3 AI-mediated externalization and constructed experience The last thread in the prompt pushes beyond workplace production into experience construction: when AI helps articulate thoughts or feelings, it may participate in shaping what the experience becomes. Lisa Feldman Barrett's theory of constructed emotion (in an active inference framing) emphasizes categorization, concepts, and context in how emotion episodes are made, rather than emotions being read out as fixed internal modules. turn3search6 On this view, externalizing and labeling experience is not a neutral transcription step; it can be part of the constructive process by which ambiguous affect becomes a specific, nameable emotion in context. turn4search17 Direct literature explicitly connecting Barrett's framework to everyday LLM interaction is still relatively thin and scattered, but several adjacent research areas strongly imply the connection (and make it empirically researchable). First, LLM-based journaling and reflection systems are explicitly designed to scaffold articulation of daily experiences, emotions, and thoughts—i.e., to change the practice of externalization. MindfulDiary, developed with mental health professionals and evaluated in a four-week field study with psychiatric patients and psychiatrists, reports that AI-supported journaling helped patients enrich daily records and helped clinicians gain deeper, more empathetic understanding of patients' contexts and thoughts. turn18search4 MindScape similarly frames contextual AI journaling as a "new frontier," combining behavioral sensing and LLM prompts to encourage self-reflection and emotional development. turn18search17 Second, research on affective use of chatbots suggests that high-intensity engagement can correlate with dependence-like indicators and that impacts on emotional well-being can be nuanced, varying by modality and initial emotional state. This matters for "constructed experience" because it points to feedback loops: if an AI becomes a frequent interlocutor for interpreting experience, it may shape not only expression but also attention, categorization habits, and relational expectations. Third, early design-science and HCI work explicitly explores operationalizing "constructed emotion" in computational systems (for example, context-dependent emotion analysis pipelines) and argues for focusing on situations rather than discrete emotion labels—an orientation compatible with analyzing LLM-mediated self-description as a constructive act. Systems like ExploreSelf, which use adaptive LLM guidance for reflecting on personal challenges, similarly build on the premise that guided articulation changes reflection outcomes, even when the goal is not "emotion detection." A careful synthesis, consistent with Barrett's theory and these systems, is that AI-mediated externalization is plausibly interventionist rather than representational: the prompts offered, the vocabulary proposed, the framing of causes and actions, and the conversational rhythm can all shift what gets noticed and how it is categorized. This claim is an inference (not yet a settled empirical conclusion), but it is strongly supported by the convergence of (a) constructed emotion theory's emphasis on context and categorization and (b) the demonstrated ability of LLM systems to structure reflective practice through conversational scaffolding. Toward a multidimensional jagged frontier A multidimensional view treats "the frontier" as a vector of practice properties, not a scalar of output quality. This is consistent with practice-theoretic accounts of technology-in-use, and with recent AI scholarship arguing for sociomaterial, performative approaches to AI in organizations. citeturn6search2turn12view6 It is also compatible with the empirical pattern in productivity/quality studies: benefits are real but uneven, creating discontinuities across task types and worker groups. One way to formalize the hypothesis in the prompt is to treat AI adoption as moving a practice across multiple "frontiers," each with its own jaggedness: Skill formation frontier: whether novice-to-expert development remains viable when peripheral tasks are automated; includes risks of deskilling and "never-skilling," plus altered mentorship and inverted apprenticeships. turn23search1 Craft frontier: whether practitioners experience the work as authorship, editorial control, or governance labor (prompting, steering, verifying), and how pride, autonomy, and professional jurisdiction shift. Pace-and-boundary frontier: whether AI reduces work or intensifies it via workload creep, scope expansion, and blurred time boundaries. Accountability frontier: whether responsibility is appropriately distributed or collapses into "moral crumple zones," and whether handoffs preserve legitimacy, participation, and ethical equivalence. turn10search6turn10search10 Relational communication frontier: whether AI-mediated messages change social meaning, disclosure norms, authenticity expectations, and signaling equilibria (including "AI penalty" vs "blissful ignorance"). citeturn7search36turn11view5turn17view1 Constructed-experience frontier: whether AI-mediated articulation participates in shaping emotion/thought categories and attention habits over time, especially in reflective or mental-health-adjacent use. Methodologically, this multidimensional frontier view implies that evaluating AI-in-workplace interventions requires mixed methods and longitudinal designs: performance experiments should be paired with ethnography, diary studies, interaction-trace analysis of reliance and error recovery, and assessments of learning-over-time and governance practices. citeturn10search18turn10academia41turn14view0 Policy and organizational-tooling lineages such as algorithmic impact assessments offer practical templates for documenting stakeholder impacts and accountability structures alongside performance results. citeturn10search28turn10search32 The main takeaway is that "where the frontier is jagged" depends on which axis you measure. The same AI system can yield strong gains in output quality while degrading apprenticeship pathways, intensifying work pace, shifting responsibility onto thin human oversight layers, and altering communicative authenticity norms.