Daniel Griffin

Practitioner Accounts Across Five Domains

ChatGPT Deep Research report (2026-03-23)

Query Investigate how people in software engineering, medicine, law, education, and creative writing describe where AI tools work and where they fail. Focus on first-person accounts — blog posts, forum threads, interviews, conference talks, podcast transcripts. Are people in domains with rich feedback loops (code, data engineering) less surprised by AI failures than people in domains without them (writing, strategy, consulting)? Do practitioners in feedback-rich domains describe AI errors as routine and recoverable, while practitioners in feedback-poor domains describe them as surprising and dangerous? Collect specific quotes and organize by domain. Pay attention to whether people describe building personal habits or workflows to compensate for AI failures — these are extensions in the making. --- Full Report AI Capability Boundaries in First-Person Practitioner Accounts Across Five Domains Executive summary Across software engineering, medicine, law, education, and creative writing, practitioners consistently describe today’s AI tools as high-leverage helpers for drafting and synthesis—and high-risk when treated as authoritative in tasks where mistakes are costly or hard to detect. In first-person accounts, the “boundary line” is rarely phrased as AI can/can’t do X; instead it is framed as what the practitioner can reliably verify (and how fast). In domains with rich, immediate feedback loops (especially software engineering), failures are often described with low surprise—engineers expect hallucinations and rely on compilation/tests as the truth serum. In domains with weaker or delayed feedback loops (law filings, clinical reasoning, and creative writing quality), failures produce more intense emotional language (mortification, fear, distrust) because errors can look plausible and may be discovered only after reputational or real-world harm. Research approach and coding scheme This report prioritizes primary, first-person sources in English: practitioner blog posts, professional outlets that use first-person experimentation, and practitioner-heavy forums (notably Hacker News and Reddit) where users speak in the first person about real workflows. Each quote is kept short (≤25 words) and presented verbatim, along with author/date/URL as requested. Codes used per quote • Task type: what the practitioner was trying to do (e.g., refactor code, draft an appeal letter, grade essays, brainstorm plot). • Tool used: the AI system named in the account (or the closest description, when the author names a product family). • Feedback loop strength: • Immediate testable (compile/tests/run output right away) • Immediate verifiable (facts/citations can be checked quickly against an external authority, like Westlaw) • Delayed objective (outcomes appear later but can be measured, e.g., clinical outcomes, appeal results) • Delayed subjective (quality is judged later and is partly taste/voice-dependent, e.g., fiction quality). • Valence: Positive (reliably helpful), Negative (failure/risk dominates), Mixed (helpful but bounded/fragile). • Emotional tone markers: surprise, trust, frustration, fear/avoidance, relief. Domain findings Software engineering | Verbatim first-person quote | Source with URL, author, date | Context classification | |---|---|---| | “It gave me a brilliant answer. It wrote a clean function.” | Story Crafter, AWS in Plain English, Feb 10, 2026, https://aws.plainenglish.io/dont-trust-your-ai-co-pilot-the-rise-of-package-hallucination-0897ecf93336 | Task: code generation for Node.js parsing. Tool: ChatGPT. Feedback loop: immediate testable (run code, install deps). Valence: Positive (local code quality). Tone: impressed, then cautious contextually. | | “you can generate a lot of code, really fast, but you can’t trust what comes out.” |, Blog, Dec 31, 2024, https://stackoverflow.blog/2024/12/31/generative-ai-is-not-going-to-build-your-engineering-team-for-you/ | Task: production engineering framing (trust boundary). Tool: “tools like Copilot/chatGPT” (LLMs for code). Feedback loop: immediate testable + code review norms. Valence: Mixed (speed vs. trust). Tone: skeptical, admonitory. | | “If you're using an LLM to write code without even running it yourself, what are you doing?” |, Mar 5, 2025, https://simonw.substack.com/p/hallucinations-in-code-are-the-least | Task: defining safe usage posture. Tool: LLM coding tools (incl. agentic loops). Feedback loop: immediate testable (run code!). Valence: Negative toward “blind use.” Tone: blunt, norm-setting. | | “Trouble showed up fast.” |, Aug 19, 2025, https://gordonbeeming.com/blog/2025-08-19/i-let-copilot-refactor-my-app-so-you-dont-have-to-the-wrong-way | Task: cross-solution refactor (rename entity). Tool: GitHub Copilot. Feedback loop: immediate testable (build/tests) but broad refactor increases surface area. Valence: Negative. Tone: rueful, cautionary. | | “When AI code breaks, one probably doesn't.” |, Jul 11, 2025, https://dev.to/anchildress1/github-copilot-agent-mode-the-mistake-you-never-want-to-make-1mmh | Task: describing operational debugging cost. Tool: Copilot “agent mode.” Feedback loop: immediate testable, but diagnosis requires understanding. Valence: Negative. Tone: wary, experience-based. | | “Devin was pretty bad and honestly soaked up more time than it saved.” | Hacker News comment (user phil917), Feb 6, 2025, https://news.ycombinator.com/item?id=42964327 | Task: agentic coding evaluation. Tool: Devin / Cursor / Copilot agent. Feedback loop: immediate testable, but overhead dominates. Valence: Negative. Tone: exasperated. | | “the moment it tries to make any sort of structural decision I find it to be thoroughly unhelpful” | Reddit comment (user Kirne), ~2024 (thread shows “2y ago”), https://www.reddit.com/r/programming/comments/1ac7cb2/newgithubcopilotresearchfinds_downward/ | Task: code completion vs architecture. Tool: Copilot. Feedback loop: immediate testable. Valence: Mixed (autocomplete good, design bad). Tone: pragmatic. | | “flat out wrong somewhere around 50% of the time” | Reddit comment (user lqstuart), ~2024 (thread shows “2y ago”), https://www.reddit.com/r/programming/comments/1ac7cb2/newgithubcopilotresearchfinds_downward/ | Task: infra/deep learning code assistance. Tool: Copilot. Feedback loop: immediate testable. Valence: Negative. Tone: resigned frustration. | Two–three sentence synthesis: In software engineering accounts, the capability boundary is repeatedly framed as “speed until you hit reality”—the reality check being compilation, tests, running the program, and code review. Emotional tone is often low-surprise skepticism (“you roll the dice”), with frustration aimed less at the model’s existence and more at workflows that skip verification and then pay the debugging tax. Medicine | Verbatim first-person quote | Source with URL, author, date | Context classification | |---|---|---| | “ChatGPT works fairly well as a diagnostic assistant — but only if you feed it perfect information” |, Inflect Health (Medium), Jun 14, 2023, https://inflecthealth.medium.com/im-an-er-doctor-here-s-how-i-m-already-using-chatgpt-to-help-treat-patients-a023615c65b6 | Task: ED support; diagnostic assistance boundary. Tool: ChatGPT. Feedback loop: delayed objective (true diagnosis emerges via workup/outcomes). Valence: Mixed (conditional usefulness). Tone: caution, realism. | | “I anonymized my History of Present Illness notes… and fed them into ChatGPT.” | Same author, , Mar 13, 2023, https://www.fastcompany.com/90863983/chatgpt-medical-diagnosis-emergency-room | Task: real-case trial of differential generation. Tool: ChatGPT. Feedback loop: delayed objective (chart/labs/imaging vs output). Valence: Mixed (experiment framing). Tone: investigative curiosity. | | “transform medical shorthand writing… into a… clinical note.” | blog by, Aug 18, 2023, https://www.idsociety.org/science-speaks-blog/2023/how-i-use-artificial-intelligence-as-an-infectious-diseases-clinician/ | Task: clinical documentation. Tool: ChatGPT-4 (named in article). Feedback loop: immediate verifiable (clinician reads note) + privacy constraints. Valence: Positive. Tone: practical optimism. | | “it would not be safe or ethical to use ChatGPT’s logic… to determine a final diagnosis” | blog by, Jul 26, 2023, https://biomedicalodyssey.blogs.hopkinsmedicine.org/2023/07/chatgpt-in-medicine-part-1-learning-about-the-differential-diagnosis/ | Task: differential diagnosis learning. Tool: ChatGPT. Feedback loop: delayed objective (clinical truth) + training context. Valence: Negative (for final decisioning). Tone: safety-conscious. | | “I asked both programs about the potential hazards of using them.” | () by, Jul 2023, https://journals.lww.com/em-news/fulltext/2023/07000/aiassistedmedicinepossiblyhelpful%2Cpossibly.17.aspx | Task: hazard identification and bedside use reflection. Tool: ChatGPT + Google Bard. Feedback loop: mixed—some immediate verifiable (patient education), some delayed (clinical outcomes). Valence: Mixed. Tone: intrigued but wary. | | “30 min later I realized that it had hallucinated the references as well.” | Reddit comment (user pdxiowa), ~mid‑2025 (thread shows “9mo ago”), https://www.reddit.com/r/hospitalist/comments/1ksv02u/youusingchatgpt/ | Task: evidence lookup for clinical question. Tool: ChatGPT w/ citations request. Feedback loop: immediate verifiable (try to find paper). Valence: Negative. Tone: time-wasted frustration. | | “I say that as someone using an AI scribe on all my shifts.” | Reddit comment (user worldbound0514), ~early‑2026 (thread shows “3 weeks ago”), https://www.reddit.com/r/medicine/comments/1rfth1l/chatgpthealthfailstosend52of_simulated/ | Task: clinical documentation/throughput. Tool: AI scribe + ChatGPT Health discussion context. Feedback loop: immediate verifiable (note review), but downstream medico-legal risk delayed. Valence: Mixed (scribe useful; advice risk rejected). Tone: hard-line caution. | Two–three sentence synthesis: Medical practitioners describe a sharper boundary between documentation/patient communication (where correctness is often quickly reviewable) and diagnosis/clinical decision support (where missing a rare-but-lethal condition is catastrophic and may be discovered late). Emotional tone mixes curiosity and surprise at surface competence with strong safety language (“not safe or ethical”), and repeated concern about fabricated evidence and privacy exposure. Law | Verbatim first-person quote | Source with URL, author, date | Context classification | |---|---|---| | “Since last year, I have been utilizing ChatGPT… [for] summarizing cases… proofreading.” |, May 27, 2024, https://conlinpa.com/2024/05/27/trust-but-verify-chatgpt-is-a-tool-but-not-a-replacement-for-human-legal-research/ | Task: research summarization and drafting polish. Tool: ChatGPT + Westlaw + Lexis AI (named). Feedback loop: immediate verifiable (read case; compare tools). Valence: Positive. Tone: efficiency-seeking, measured. | | “Do not believe it’s true. Verify. Every. Detail.” |, Jul 28, 2025 (page notes update Aug 8, 2025), https://www.dorianinsurancelaw.com/blog/does-your-ai-lawyer-have-a-fool-for-a-client | Task: warning about AI-drafted legal content. Tool: ChatGPT / Gemini (named). Feedback loop: immediate verifiable (cite-check), but consequences delayed/high-stakes. Valence: Negative. Tone: alarmed, emphatic. | | “Now it's DEFINITELY up to me to check the statutes” |, ~Aug 2025 (LinkedIn shows “7mo”), https://www.linkedin.com/posts/sarakubik_chatgpt-activity-7356694484849491968-OhIh | Task: entry into new substantive area; statutory pointers. Tool: ChatGPT. Feedback loop: immediate verifiable (check statutes). Valence: Mixed. Tone: enthusiastic but careful. | | “I've been sitting on it feeling quietly mortified” | Reddit post (user anuj_meme), Feb 27, 2026 (from createdutc), https://www.reddit.com/r/LawFirm/comments/1rgpfko/aihallucinatedafederalcourtcitationinmy/ | Task: motion drafting; supporting case law. Tool: “AI” for research (not named in excerpt). Feedback loop: immediate verifiable (search case), but failure surfaced via opposing counsel—late in pipeline. Valence: Negative. Tone: shame, relief, fear. | | “I’d never submit something to the court that I didn’t double check on Westlaw.” | Reddit comment (user sxzzyw), ~2023 (thread shows “3y ago”), https://www.reddit.com/r/LawSchool/comments/13szgl6/anattorneyfoundoutthehardwaywhyone_should/ | Task: litigation filing guardrails. Tool: ChatGPT. Feedback loop: immediate verifiable (Westlaw). Valence: Mixed (use is bounded). Tone: norm-enforcing pragmatism. | | “Chatbots are incredibly useful for legal writing… ChatGPT helped me cut paragraphs… while keeping key details.” | Reddit comment (user Sweaty_Resist_5039), ~late‑2025 (thread shows “6mo ago”), https://www.reddit.com/r/Lawyertalk/comments/1nnvbel/californiaattorneyfined10000forfilingan/ | Task: compressing a brief to page limits; rewriting. Tool: ChatGPT. Feedback loop: immediate verifiable (read output), but legal strategy quality partly delayed. Valence: Positive. Tone: pragmatic optimism. | Two–three sentence synthesis: Legal practitioners describe AI as most useful for language work (rewriting, shortening, drafting scaffolds) and least trustworthy for authority-bearing outputs (citations, holdings, jurisdiction-specific nuance). Emotional tone is high-intensity compared to software engineering: “mortified” and emphatic verification language reflects that failures can look authentic and may be discovered only after filing, when the cost of correction rises sharply. Education | Verbatim first-person quote | Source with URL, author, date | Context classification | |---|---|---| | “I use ChatGPT… But I worry… it may get overused in education and… damage our ability to learn.” |, May 11, 2023, https://alexquigley.co.uk/why-might-chatgpt-damage-learning/ | Task: learning process vs product critique. Tool: ChatGPT. Feedback loop: delayed/subjective (learning quality over time). Valence: Mixed. Tone: worried, protective. | | “I can give detailed and accurate feedback on 120 essays two days after they’ve been turned in.” | Reddit comment (user Linkpharm2), ~2025 (thread shows “1y ago”), https://www.reddit.com/r/Teachers/comments/1k6yug2/grading_confession/ | Task: grading + feedback at scale. Tool: ChatGPT (trained to rubric). Feedback loop: immediate verifiable (teacher reviews comments) but downstream learning delayed. Valence: Positive. Tone: relieved, efficiency-driven. | | “ChatGPT did my professional growth plan… it saved me hours. No one said a word.” | Reddit post (OP), ~2023 (thread shows “3y ago”), https://www.reddit.com/r/Teachers/comments/12jyusx/iusechatgptforschoolpaperworkwarningiam/ | Task: administrative paperwork. Tool: ChatGPT. Feedback loop: immediate verifiable (admin acceptance), low scrutiny. Valence: Positive (time saved) with implicit ethical tension. Tone: mischievous, candid. | | “It's not as good as the news… but it's also more useful than many new users think.” | Reddit comment (user ErusTenebre), ~2023 (thread shows “3y ago”), https://www.reddit.com/r/Teachers/comments/14vnut7/hasanyoneusedchatgptat_work/ | Task: district training; classroom use cases. Tool: ChatGPT. Feedback loop: mixed; some immediate (drafts), some delayed (student understanding). Valence: Mixed. Tone: calibrated realism. | | “the results were fantastically bad… wouldn't consider it for grading.” | Reddit comment (user Black_RL), ~2024 (thread shows “2y ago”), https://www.reddit.com/r/artificial/comments/1b8u06g/someteachersarenowusingchatgptto_grade/ | Task: grading experiment. Tool: ChatGPT. Feedback loop: immediate verifiable (compare to rubric/expectations). Valence: Negative. Tone: blunt disappointment. | | “everyone agreed they’d need to heavily edit… because… the AI voice was ‘off’… and… invented fake details.” |, ~Jan 2026 (LinkedIn shows “2mo”), https://www.linkedin.com/posts/sarahcsutton_when-i-instructed-the-students-in-my-bc-science-activity-7414293003604795392-s7C8 | Task: teaching AI literacy via classroom exercise. Tool: ChatGPT. Feedback loop: immediate verifiable (spot fakes/voice issues), but writing development delayed. Valence: Mixed. Tone: reflective, concerned. | Two–three sentence synthesis: Educators emphasize two boundaries: AI is highly effective for teacher throughput work (paperwork, drafts, rubric-aligned feedback) but risky when it substitutes for student cognition and authenticity, where the feedback loop is delayed and confounded. Emotional tone is split—teachers express relief at time savings but also worry, even moral discomfort, because errors (or “off voice” and invented details) can quietly propagate into student work and understanding. Creative writing | Verbatim first-person quote | Source with URL, author, date | Context classification | |---|---|---| | “Some of ChatGPT’s ‘stories’ were useless… It is also good as a soundboard for brainstorming” |, May 5, 2023, https://tzbarry.com/2023/05/05/chatgpt-has-no-voice-and-it-must-mimic/ | Task: fiction drafting (short stories). Tool: ChatGPT. Feedback loop: delayed subjective (quality/voice). Valence: Mixed. Tone: experimental, discriminating. | | “ChatGPT said of my memoir… ‘you could add more dialogue’… how authentic… to remember… conversations verbatim?” |, Mar 24, 2025, https://sanjidakay.substack.com/p/10-ways-ai-can-help-writers | Task: feedback on memoir craft. Tool: ChatGPT. Feedback loop: delayed subjective (genre/authenticity judgment). Valence: Mixed. Tone: thoughtful skepticism. | | “Not too shabby… capturing Hemingway’s rhythm.” |, Jun 29, 2023, https://www.writersdigest.com/be-inspired/chatgpt-a-writers-best-friend-for-now | Task: style imitation micro-fiction. Tool: ChatGPT. Feedback loop: immediate subjective (reader ear/voice). Valence: Positive. Tone: pleasantly surprised. | | “Using Chat GPT AI to write stories… is… cheating, in my mind!” |, May 24, 2023, https://medium.com/dose-of-wonder/when-i-asked-a-chat-gpt-ai-to-write-a-short-mystery-610fee29d432 | Task: AI short-story generation experiment. Tool: Nova App (ChatGPT-like). Feedback loop: subjective/ethical evaluation. Valence: Negative. Tone: principled disapproval. | | “I’m here to tell you: THIS IS A TERRIBLE IDEA!” | (Paulina Cossette LLC), Dec 11, 2023, https://acadiaediting.com/2023/12/11/3-reasons-why-ai-is-terrible-for-editing-your-writing-the-last-one-will-surprise-you/ | Task: editing academic writing. Tool: ChatGPT 3.5 (explicit). Feedback loop: partly verifiable (grammar) but meaning/style subjective and field-specific. Valence: Negative. Tone: emphatic warning. | | “I use Chatgpt to outline… It doesn't write for me” | Reddit comment (user MapleWafer), ~2025 (thread shows “1y ago”), https://www.reddit.com/r/authors/comments/1gvq1p0/foundoutmyauthorfrienduseschatgptinher/ | Task: outlining/structuring plot. Tool: ChatGPT. Feedback loop: delayed subjective (does outline produce good story). Valence: Positive. Tone: pragmatic, legitimizing. | | “wow! did it murder my dialogue” | Reddit comment (thread includes Gemini + GPT comparison), ~2025 (thread shows “1y ago”), https://www.reddit.com/r/WritingWithAI/comments/1hhs5z4/fellowwritersneedyourinsightsaboutai/ | Task: “streamline” rewrite of a chapter. Tool: Gemini (named) and GPT comparisons. Feedback loop: immediate subjective (readability/voice). Valence: Negative. Tone: visceral frustration. | Two–three sentence synthesis: Creative writing practitioners separate AI’s usefulness for structure and ideation (outlines, brainstorming, developmental prompts) from its weakness in voice, dialogue, and authentic genre intent, where “plausible” text can still feel dead or wrong. Emotional tone ranges from curiosity and occasional admiration (micro-style mimicry) to strong hostility when AI outputs overwrite voice or violate creative/ethical self-concepts (“cheating,” “murder my dialogue”). Cross-domain comparison: feedback loops and surprise Counts of positive, negative, and mixed first-person accounts The table below counts the quotes included in this report by valence (a coarse proxy for “AI worked vs failed vs both”). This is not a population estimate; it is a structured summary of the collected, representative first-person accounts surfaced here. | Domain | Positive | Negative | Mixed | What “feedback” most often looks like in the quotes | |---|---:|---:|---:|---| | Software engineering | 1 | 5 | 2 | Compilation/tests + debugging time as ground truth | | Medicine | 2 | 1 | 4 | Clinician review + evidence lookup; outcomes and safety constraints | | Law | 2 | 2 | 2 | Citation verification + filing consequences (late-stage detection) | | Education | 2 | 1 | 3 | Surface correctness vs student learning/voice (delayed) | | Creative writing | 2 | 3 | 2 | Reader/author taste and “voice” judgments; iterative rewriting | Do rich feedback-loop domains report less surprise at failures? Pattern in the evidence collected here: domains with strong, immediate feedback loops—especially software engineering—tend to describe AI failures as expected and operationally manageable. Engineers talk about “not trusting” outputs and imply a routine verification posture (“run it yourself”), suggesting low surprise and high procedural control when failures appear. In domains where feedback is weaker, delayed, or socially mediated (creative writing quality, legal filings), first-person accounts use higher-arousal emotional language. The legal “mortified” narrative illustrates a distinctive failure mode: the output looked legitimate enough to pass initial review, and the error was discovered by the adversary—late in the pipeline—producing shame and a desire for tighter “process controls.” Medicine sits between these extremes: clinicians express both surprise at surprisingly good surface performance and hard safety boundaries around diagnosis and evidence, where hallucinated references or missed conditions are unacceptable. This yields a tone of cautious experimentation rather than either engineer-like cynicism or writer-like aesthetic outrage. Typical failure discovery and correction cycles by domain mermaid flowchart TD subgraph SE[Software engineering] SE1[Prompt / spec] --> SE2[Generate code] SE2 --> SE3[Run / compile / tests] SE3 -->|Fail fast| SE4[Inspect error + diff] SE4 --> SE5[Edit code or re-prompt] SE5 --> SE3 SE3 -->|Pass| SE6[Commit + review] end subgraph MED[Medicine] M1[Prompt / summarize / draft] --> M2[Clinician review] M2 --> M3[Check guideline / evidence / labs] M3 -->|Mismatch| M4[Correct + document rationale] M4 --> M2 M3 -->|Accept| M5[Use in note / patient education] M5 --> M6[Downstream outcomes + medico-legal risk] end subgraph LAW[Law] L1[Prompt / draft argument] --> L2[AI output looks plausible] L2 --> L3[Cite-check in Westlaw/Lexis] L3 -->|Hallucination| L4[Replace with real authority] L4 --> L5[Rewrite + verify] L5 --> L6[File] L6 --> L7[Opposing counsel / court scrutiny] end subgraph EDU[Education] E1[Prompt lesson/feedback/admin text] --> E2[Teacher edits + sanity check] E2 --> E3[Deploy to students] E3 --> E4[Student work + class discussion] E4 --> E5[Assessment / learning signals over time] E5 -->|Problems| E6[Adjust prompts + pedagogy] E6 --> E1 end subgraph CW[Creative writing] C1[Prompt outline/draft/rewrite] --> C2[Read for voice + coherence] C2 -->|Off-voice / cliche| C3[Rewrite by human] C3 --> C4[Peer/beta feedback] C4 --> C5[Iterate] C5 --> C2 end Implications for tool design and deployment Design and deployment implications follow directly from the first-person boundary descriptions in the sources: Make verification the default user experience, not an optional virtue. Engineers and lawyers repeatedly frame safe use as “trust but verify,” but in law and medicine the cost of late discovery is high. Tooling that automatically supports cite-checking, provenance display, and “show me where this came from” flows aligns with how practitioners already police boundaries. Separate “drafting modes” from “authority modes.” Many accounts endorse AI for drafting (emails, briefs, discharge instructions, lesson scaffolds) and reject it for authoritative claims (citations, final diagnosis logic, invented details). Products should reflect this with explicit modes, stronger guardrails, and different UI affordances for risk—e.g., “draft language only” vs “evidence-backed answer.” Tune the tool to the domain’s feedback loop—and compensate when the loop is weak. Where feedback is immediate (code), users adapt quickly and failures are less surprising; where it’s delayed or subjective (education outcomes, writing voice), users need structured critique tools (style provenance, contradiction checks, “invented detail” detectors) and workflow scaffolds (revision checklists, comparison to source materials). Support “partial automation” patterns that practitioners actually trust. Many first-person accounts converge on a stable pattern: AI for speed on the “blank page” step, then human judgment for final decisions. Tool design should emphasize iteration, diff-based editing, and transparent limitations rather than marketing an illusion of autonomy that practitioners describe as time-wasting or dangerous.