Planning CSUIRx

I have a wide range of criteria by which to provide some marks and remarks on these systems. I'll need to narrow them down and gradually work through them. I'm not thinking about them as goals, requirements, or desired-specifications, and some even may contradict. For some of the criteria I will provide citations as reference or support. Some criteria are drawn from examples in previous search systems (including shuttered, speculative, and experimental systems). My goal here is not to simply do an accounting of searching today, but to get some sense of where we might want search to go.

Sample

Here is a sample of a few questions as applied to several web search systems.

Initial Search Systems

In this initial set of reviews, I'm focusing on these search engines, listed alphabetically:

These are my initial examples of new approaches to searching in generative web search systems. I may provide come contextualizing comments about other systems, like the explicit search-focused tools from Google and Microsoft, and chat-based systems like ChatGPT, Anthropic's Claude, etc., that support search and search-like interactions.

To they extend that they support public-facing search, I will also be examining newer search libraries and services (including RAG frameworks), like the offerings from LangChain, from LlamaIndex, and Weaviate's Verba, with comparison to (the also adapting) existing tools like those from Algolia and Elasticsearch.

Broad Criteria

My criteria are broad. I'm focused on concerns my research and training best prepares me to engage with. These are broadly questions related to the explicit and implicit articulation of the search system, the interactions around queries and results, the ability to share the burden of search, and the formalized methods of complaint. I'll do some explicit evaluations of atomic performance related to "hallucination" or "groundedness", but my focus is more on how people perceive and perform-with tool outputs than the outputs themselves. How are the searchers ushered into their searches? What do they see as searchable? How can they engage with search results (or responses)? Are they expected to vet the responses for hallucinations? How is automation bias addressed? What post-search activities are supported by the search system itself?

I'll ask about features or uses that might perhaps be refused or reimagined, while situating this period of search amidst a longer history of search. There are importance concerns about misleading results, sources of training and reference data, oversight, and the future of work. I'm very much developing these reviews to acknowledge that these systems will keep changing. Where there are very important concerns that I am less well-versed in, like accessibility, I will leverage other resources.

{% include weblinks.html url="https://twitter.com/danielsgriffin/status/1698387713198522645" %}

Scoping

This is not intended to be an introductory guide to these systems, but focused on making sense of what new search tools are providing and what they might become. These reviews may be useful to heavy users, developers, and others looking to understand changes in system support for various searching practices.

I will largely be looking at systems for web search, including those more focused to particular subject areas. Though important, these reviews will not (yet at least) engage with new search systems for:

I will also not be focused on in-editor code generation tools (like GitHub's Copilot) and writing tools (like Lex.page) that replace or subsume some searching tasks.

I will not very focused on various metrics related to speed, unless it is very noticeable in frequent use.

I am concerned about questions of bias, but here only insofar as these systems are markedly different from the prior problems found in search.

I am not focused on explainability or transparency of these systems, though some question will definitely engage with those questions. I will be more focused on examining questions around seamfulness, tractability, and traceability. I will be thinking about how practical algorithmic knowledge [@cotter2022practical] is built up and valued.

I'm less focused on responding to or rehashing and regurgitating arguments about "model collapse", than perhaps looking at how these search tools and their users imagine supporting or working towards unsealing knowledge, whether through articulations that help users doubt & dig deeper, providing multiple drafts, or RAG adaptations.

The most important work would be work looking at how these search systems and tools are imagined and used (or not) by other people. I am not looking at that right now, but I will look at aspects of the systems identified publicly by different users or others.

Acknowledgements

I've long wondered been inspired by the work at Ranking Digital Rights, and wondered what a related approach might help us think about in relation to our curiosity, questions, ignorance, doubts, and claims-making. (Of course digital rights are heavily implicated in the design and operations of search engines and in search itself.) Much of my thoughts around this were developed in conversations with Emma Lurie while we worked on our paper on Google's Search Liaison [-@griffin2022search]. I've thought also of her comparison of platform research API requirements [@lurie2023comparing] while thinking through this. It was also a help to write-up and share The Need for ChainForge-like Tools in Evaluating Generative Web Search Platforms. I've also been inspired by recently seeing Search Smart, which looks at the academic search domain with largely different criteria of evaluation. And thanks Bill Chambers for a recent nudge.

Daniel Griffin