
@kabir2023answers ([a preprint on arXiv](https://arxiv.org/abs/2308.02312)) compares ChatGPT 3.5 Turbo API answers[^3_5] with Stack Overflow answers to 517 questions found on Stack Overflow. Two attention-grabbing tweets about the paper with slightly different quotes raised questions for me about their method for identify an answer as incorrect:

[^3_5]: A common critique on Twitter is that the paper did not evaluate answers from the latest model, GPT-4. That is a critique that should at least be addressed in revisions of the paper.


<div class="card m-2 border mx-auto justify-content-center">
<div class="card-body text-center p-0 py-1 m-0">
<p class="small p-0 m-0">
Qualifier: @kabir2023answers is more than the exciting takes on Twitter.
</p>
</div>
<div class="row card-body p-1 m-0 justify-content-center">
<div class="col-sm-10 col-lg-5 m-1 p-0">

{% include weblinks.html url="https://twitter.com/timnitGebru/status/1689768597437169664" %}

</div>
<div class="col-sm-10 col-lg-5 m-1 p-0">

{% include weblinks.html url="https://twitter.com/GaryMarcus/status/1689851180657397760" %}

</div>
</div>
</div>
</div>


While I didn't find it explicitly noted in the paper, they seem to mark an answer as a whole "incorrect" if it contains any anything incorrect among four types? Here is p. 4:

> For _**Correctness**_, we compared ChatGPT answers with the accepted SO answers and also resorted to other online resources such as blog posts, tutorials, and official documentation. Our codebook includes four types of correctness issues— _Factual_, _Conceptual_, _Code_, and _Terminological_ incorrectness. Specifically, for incorrect code examples embedded in ChatGPT answers, we identified four types of code-level incorrectness—_Syntax_ errors and errors due to _Wrong Logic_, _Wrong API/Library/Function Usage_, and _Incomplete Code_.

This answer-level "correctness" is striking to me (given my interest in [how people perceive and perform-with tool outputs](/2023/08/21/people-perceive-and-perform)). For many multi-part problems a single incorrect step in an answer produces a failure. But that does not hold for all problems. There are many problems that are resolved iteratively, where drafts of initial answers are build upon for subsequent answers.


<div class="card m-2 border border-info mx-auto">
<div class="card-body">
<p class="small p-0 m-0">
Aside: I do not believe it is clear that "Stack Overflow is being destroyed" nor what the costs or causes may be. Clearly a full analysis of LLM answers would need to incorporate many downstream interactions with the same people and systems that produce so much training data, but it seems strange to me that more granularity is not used here, especially since practical utility tests were not conducted or generalized-to.</p>
</div>
</div>