All Research Projects

Large Language Models Outperform Human Coders in Neurotology Operative Note Coding

Exploring Group Improvisation Training to Improve Decision Making and Well-being in Otolaryngology Surgical Residents

Research area

HEALTHTECH

Project Leads

Stephanie Younan, MPH BS

Medical Student

Nicole Jiam, MD

Executive Director and Assistant Professor

Additional Authors: Pearl Doan, BS, Vanessa S. Reyes, MPA, Jolie L. Chang, MD, Yew Song Cheng, BM BCh

Share

Green Fern

Introduction

Surgical coding determines how procedures are billed and reimbursed. Every operative note is assigned CPT codes that translate into relative value units (RVUs), which form the basis for institutional revenue. When codes are missing or incorrect, the financial impact compounds over time and is rarely caught until a formal audit.

Many academic medical centers use centralized coding teams to manage this process. These teams handle high volumes across many specialties, which can make it difficult to maintain the depth of knowledge needed for technically complex subspecialties like neurotology. Large language models offer a different approach: AI that reads operative notes and generates code assignments based on the specific details of each case.

Highlight sentence

Coding errors, whether from missed codes or incorrect assignments, accumulate quietly over time and can represent a substantial source of revenue loss for surgical practices.

AI vs. Human Coders in Neurotology

We reviewed 124 consecutive neurotology operative notes from UCSF, covering procedures performed between July 2024 and June 2025. Each note was independently coded by UCSF's institutional LLM and the centralized human coding team. A ground truth standard was established through blinded surgeon review using AAPC and AAO-HNS coding guidelines.


The LLM achieved 86.3% coding accuracy compared to 49.2% for human coders, a difference of 37.1 percentage points (p<0.001). Human coders had a mean RVU variance of -5.04 relative to the ground truth, reflecting consistent under-coding. The LLM's mean variance was +0.27, close to neutral. Every one of the 61 human coding errors involved a missing or incorrect code. LLM errors were split between missing codes and extraneous codes. Across the full 5-surgeon division, the pattern of human under-coding projects to roughly 1,950 lost RVUs per year, equivalent to $145,178 in annual revenue.

The results suggest a hybrid approach may offer the best outcome: LLM-generated code drafts reviewed by specialty-trained human coders. This would combine the breadth and consistency of AI with the clinical judgment needed to catch edge cases, improving both accuracy and revenue integrity for subspecialty surgical programs.

Address

505 Parnassus Ave, 14th Floor

San Francisco, CA 94143

(415) 353-2757

OIC@ucsf.edu

Follow Us

  • Otolaryngology Innovation Center

© 2026 UCSF Innovation Center. All rights reserved.

Address

505 Parnassus Ave, 14th Floor

San Francisco, CA 94143

(415) 353-2757

OIC@ucsf.edu

Follow Us

  • Otolaryngology Innovation Center

© 2026 UCSF Innovation Center. All rights reserved.

Address

505 Parnassus Ave, 14th Floor

San Francisco, CA 94143

(415) 353-2757

OIC@ucsf.edu

Follow Us

  • Otolaryngology Innovation Center

© 2026 UCSF Innovation Center. All rights reserved.

Address

505 Parnassus Ave, 14th Floor

San Francisco, CA 94143

(415) 353-2757

OIC@ucsf.edu

Follow Us

  • Otolaryngology Innovation Center

© 2026 UCSF Innovation Center. All rights reserved.