Extended Capability Library
A working library of sources on how AI capability is extended, validated, and repaired in practice.
- Applied AI Handoff Atlas
Daniel S. Griffin
A field-evidence companion to the Extended Capability Library: Curiosity Builds read as small handoffs of judgment, memory, practice, access, and representation into AI-shaped systems.
note · 2026
- Hermes Agent README
Nous Research
The Hermes Agent README presents an open agent harness with model-provider switching, terminal and messaging interfaces, scheduled automations, isolated subagents, toolsets, persistent memory, session search, and a closed learning loop around skills.
doc · 2026
- An open-source spec for Codex orchestration: Symphony
Alex Kotliarskyi, Victor Zhu, and Zach Brock
OpenAI describes Symphony, a spec and reference implementation that turns issue trackers such as Linear into always-on control planes for coding agents, shifting humans from supervising sessions to managing work.
blog · 2026
- What Is an Agent Harness
Aparna Dhinakaran
Defines the modern agent harness as an out-of-the-box architecture that emerged from coding agents: an iteration loop over tools, context management, skill/tool discovery, permissions, hooks, session persistence, sub-agents, and project-context injection.
tweet · 2026
- Thin Harness, Fat Skills
Garry Tan
Short, practitioner-facing ethos doc arguing that the durable leverage in agent systems comes from model-resident skills (markdown) and deterministic code at the edges, with the harness kept as thin as possible so each model upgrade flows through.
doc · 2026
- LLM Knowledge Bases
Andrej Karpathy
Describes a personal research workflow where raw source documents are compiled by an LLM into a markdown wiki, maintained through index files, health checks, generated outputs, and lightweight tools rather than a heavyweight RAG stack.
tweet · 2026
- Skill Issue: Harness Engineering for Coding Agents
HumanLayer
Case-study framing of harness engineering for coding agents, with specific claims about what does and does not work (notably: role-based sub-agents don't work; sub-agents for context control do).
blog · 2026
- Standard Signal: AI-native hedge fund announcement
Michael Royzen
Launch announcement for a YC-backed hedge fund where AI models both generate hypotheses and execute trades. Included here as a domain-claim entry: markets-with-P&L are a paradigmatically favorable domain — clean outcome signal, fast feedback, offline backtestable, institutionally-ratified wrapper (a fund).
tweet · 2026
- OpenEstimate: Evaluating LLMs on Reasoning Under Uncertainty with Real-World Data
Alana Renda, Jillian Ross, Michael Cafarella, Jacob Andreas
OpenEstimate is a multi-domain benchmark for testing whether language models can express calibrated Bayesian priors for numerical estimation tasks under uncertainty, using real-world datasets in healthcare, employment, and finance.
paper · 2025
- Resurrecting deceased darlings: The Missing Foreword to AI and the Art of Being Human
Andrew Maynard
Maynard publishes the cut foreword to AI and the Art of Being Human, describing months of close collaboration with Claude while emphasizing human agency, manual refinement, AI tells, fictional allegories, and practical tools for staying human with AI.
essay · 2025
- Equipping agents for the real world with Agent Skills
Anthropic
Anthropic's engineering announcement of Agent Skills: a markdown-based pattern for extending Claude's capabilities by progressive disclosure. Important as an *institutional* ratification of the thin-harness / fat-skills framing.
blog · 2025
- Claude Skills are awesome, maybe a bigger deal than MCP
Simon Willison
Practitioner synthesis of Anthropic's Agent Skills feature, arguing the markdown-file pattern is conceptually simpler and more token-efficient than MCP, and that the ease of sharing a single file is the feature.
blog · 2025
- Good and Bad Harness Engineering
Daniel Miessler
Argues that good harness engineering focuses on who the user is and what they're trying to accomplish — the 'what' — and lets the model handle the 'how'. Pairs with Miessler's 'Bitter Lesson Engineering' as a design discipline for scaffolding that extends capability rather than compensating for model weakness.
essay · 2025
- Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains
Unknown (OpenReview: 21UFlJrmS2)
Proposes rubrics as a reward source for reinforcement learning in domains where a crisp verifiable outcome does not exist. A deliberate extension of RLVR-style methods past the easy cases.
paper · 2025
- Building an AI-ready public workforce
OECD
OECD full report on how public-sector workforces are (and are not) prepared to deploy AI. Brought into the library as a governance-piece anchor: the argument is that whether an AI system is capable *in practice* depends on the institutional scaffolding around its use, not only on the model or the harness.
doc · 2025
- Bitter Lesson Engineering
Daniel Miessler
Leans on Richard Sutton's 'The Bitter Lesson' to argue that prescriptive scaffolding around AI systems is a losing strategy in the limit: you should specify intent precisely and let the best available model figure out the path.
essay · 2025
- Measurement to Meaning: A Validity-Centered Framework for AI Evaluation
Olawale Salaudeen, Anka Reuel, Ahmed Ahmed, Suhana Bedi, Zachary Robertson, Sudharsan Sundar, Ben Domingue, Angelina Wang, Sanmi Koyejo
Proposes a validity-centered framework for AI evaluation that reasons explicitly about which evaluative claims the evidence actually supports, with detailed vision and language case studies. The operational companion to the Wallach/Jacobs position paper.
paper · 2025
- Expanding RL with Verifiable Rewards Across Diverse Domains
Ma et al.
Arxiv paper investigating how reinforcement learning with verifiable rewards (RLVR) generalises beyond the easy cases (math, code) to more diverse domains. The technical paper whose conceptual shadow Royzen's domain-claim entry sits in.
paper · 2025
- Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge
Hanna Wallach, Meera Desai, A. Feder Cooper, Angelina Wang, Chad Atalla, Solon Barocas, Su Lin Blodgett, Alexandra Chouldechova, Emily Corvi, P. Alex Dow, Jean Garcia-Gathright, Alexandra Olteanu, Nicholas Pangakis, Stefanie Reed, Emily Sheng, Dan Vann, Jennifer Wortman Vaughan, Matthew Vogel, Hannah Washington, Abigail Z. Jacobs
ICML 2025 position paper arguing that generative AI evaluation is fundamentally a social-science measurement problem, and presenting a four-level framework grounded in measurement theory for constructs related to GenAI capabilities, behaviors, and impacts.
paper · 2025