Curious Containers
An alien ethnographer comes to document intelligence on Earth and finds baskets, bodies, labels, playrooms, refrigerators, and one compartment it cannot classify.
An alien ethnographer comes to document intelligence on Earth and finds baskets, bodies, labels, playrooms, refrigerators, and one compartment it cannot classify.
Automation bias predicts that smooth operation trains humans out of checking. manrun is a built response—a gate that forces the practitioner to re-engage before an agent's blocked command executes. The tool exists because the literature said it should.
If the extensions framework is right, it should be falsifiable. Here's the specific within-domain prediction: the same task, in the same domain, will show different AI performance depending on whether the practice's extensions are engaged. Not across domains—within one.
Cross-trained practitioners—doctors who code, lawyers who build tools—see the frontier differently because they carry extensions from one domain into another. Their accounts reveal the mechanism: it's not the model that differs, it's what the practice provides.
The AI discourse obsesses over output quality. But for most of life, the harder problem is upstream—how do you turn 'something feels wrong' into a question worth asking?
Verification catches the error. Repairability determines whether you can do anything about it. They're different properties of the work, and repairability is itself an extension—one that can be designed in or absent.
Task-exposure models count what AI can do. Bundle theory asks whether the task can be separated from the job. The extensions framework asks a different question: does 'can do' mean the same thing with and without the practice's feedback loops?
Productivity and quality are one outcome among many. When AI enters a practice, skill formation, craft, pace, accountability, and repairability all shift. The handoff analytic is what makes these dimensions visible.
Each extension has specific capacities. The compiler grounds 'does it run?' A security review grounds 'is it safe?' Map the capacities of the extensions in a practice, and you've mapped where the frontier is smooth and where it isn't. The handoff analytic reveals what happens when those capacities shift.
The jagged frontier of AI capability isn't random. It's predictable from the practices, artifacts, and feedback loops that constitute actual work. The frontier is smooth where work is extended. It's jagged where work has been stripped to an isolated task.
I built an index of all the questions in my dissertation — research questions, interview excerpts, cited provocations.
Announcing a lecture and workshop at UCLA's Information Studies department on March 1st about maintaining and improving search engines.
Work log on search evaluation tooling including Searchevals, Searchjunct, and SearchRights.org, plus community engagement on AI search benchmarking.
Task list for upcoming Searchevals updates, Searchwhence improvements, and preparation for a UCLA lecture on search engine evaluation.
Research notes on Trieve's search-before-generate approach, ArcSearch, Perplexity AI, and testing how distractor hints affect search accuracy across multiple tools.
We have a big chance to really change search for the better. Will we?
Social media request asking if anyone is developing a comparative benchmark between major search and AI platforms.