52% are incorrect or 52% contain inaccuracies?

@kabir2023answers (a preprint on arXiv) compares ChatGPT 3.5 Turbo API answers[^3_5] with Stack Overflow answers to 517 questions found on Stack Overflow. Two attention-grabbing tweets about the paper with slightly different quotes raised questions for me about their method for identify an answer as incorrect:

[^3_5]: A common critique on Twitter is that the paper did not evaluate answers from the latest model, GPT-4. That is a critique that should at least be addressed in revisions of the paper.

{% include weblinks.html url="https://twitter.com/timnitGebru/status/1689768597437169664" %}

{% include weblinks.html url="https://twitter.com/GaryMarcus/status/1689851180657397760" %}

While I didn't find it explicitly noted in the paper, they seem to mark an answer as a whole "incorrect" if it contains any anything incorrect among four types? Here is p. 4:

For Correctness, we compared ChatGPT answers with the accepted SO answers and also resorted to other online resources such as blog posts, tutorials, and official documentation. Our codebook includes four types of correctness issues— Factual, Conceptual, Code, and Terminological incorrectness. Specifically, for incorrect code examples embedded in ChatGPT answers, we identified four types of code-level incorrectness—Syntax errors and errors due to Wrong Logic, Wrong API/Library/Function Usage, and Incomplete Code.

This answer-level "correctness" is striking to me (given my interest in how people perceive and perform-with tool outputs). For many multi-part problems a single incorrect step in an answer produces a failure. But that does not hold for all problems. There are many problems that are resolved iteratively, where drafts of initial answers are build upon for subsequent answers.

Daniel Griffin