← 返回

使用大语言模型从接受 CAR-T 治疗患者的临床记录中提取不良事件性质、严重程度、时间线和由此产生的干预措施

英文原题:Extracting adverse event nature, severity, timelines and resulting interventions from clinical notes of patients receiving CAR-T therapy using large language models.

查看英文原题

Extracting adverse event nature, severity, timelines and resulting interventions from clinical notes of patients receiving CAR-T therapy using large language models.

PubMed 2026/05/05(内容时间) medRxiv

分数与星级只用于站内排序 —— 不代表疗效、安全性或个人适用性。

中文摘要

CAR-T 细胞疗法通过基因工程改造患者 T 细胞,使其靶向肿瘤抗原,已改变血液系统恶性肿瘤的治疗,但需要仔细追踪不良事件(AE),而这些事件通常仅记录在非结构化电子健康记录(EHR)病程笔记中。

本研究在 UCSF 安全环境下评估基于大型语言模型(LLM)的方法,能否提取 6 种商业 CAR-T 产品输注后 30 天内的 AE、发生日期、分级和干预措施(2012–2023 年),并以两名评估者为基准。研究在零样本模式下使用 GPT-4 0314 和 4 种提示(预设 AE、非预设 AE、CRS、ICANS),将结果与随机抽取的 50 份记录的双重标注进行比较,指标包括准确率、精确率、召回率、F1 值和 Cohen κ。研究分析了 293 名患者的 4,762 份病程记录,患者中位年龄 65.6 岁。CRS 发生率为 80.2%(中位起病 4 天);中性粒细胞减少为 70.0%(16 天);发热性中性粒细胞减少为 64.8%(4 天);ICANS 为 34.8%。

干预措施包括托珠单抗和糖皮质激素。毒性分级常未记录(CRS 62.3%,ICANS 56.1%);有记录的病例主要为 CRS 1 级(59.4%)和 ICANS 2 级(28.0%)。模型对 CRS 和 ICANS 分级的表现较好,准确率分别为 0.97 和 0.91。对预设 AE 的提取表现中等(准确率 0.62–0.76),非预设 AE 的准确率为 0.76–0.84。CRS/ICANS 是否发生及其分级的评估者间一致性强至近乎完美(κ=0.86–0.96);日期和干预措施一致性中等,较广泛 AE 属性的一致性较弱。LLM 生成的洞见可通过提取非结构化临床细节和 CAR-T 后典型时间线,辅助 AE 监测和真实世界证据生成。

不过,模型对较广泛 AE 属性的表现不一,应谨慎使用。其在检测 CRS/ICANS 是否发生及分级方面表现最佳,评估者间一致性强至近乎完美。鉴于本研究中表现存在差异,对广泛 AE 提取应谨慎使用 LLM;但结果支持将高性能 CRS/ICANS 提取工具整合进 EHR 流程。

展开英文摘要原文

Chimeric Antigen Receptor T-cell (CAR-T) therapy, where genetically engineered patient T cells target tumor antigens, has transformed care for hematologic malignancies but requires careful tracking of adverse events (AEs) often documented only in unstructured EHR notes.

We evaluated a Large Language Model (LLM)-based approach in UCSF's secure environment to extract AEs, dates, grades, and interventions within 30 days post-infusion for six commercial CAR-T products (2012-2023), benchmarking against two evaluators. Using GPT 4 0314 in a zero shot setting with four prompts (prespecified AEs, non prespecified AEs, CRS, ICANS), we compared outputs against dual annotations on a random sample of 50 notes using accuracy, precision, recall, F1, and Cohen's kappa. From 4,762 progress notes for 293 patients (median age 65. 6), CRS occurred in 80. 2% (median onset 4 days); neutropenia 70. 0% (16 days); neutropenic fever 64. 8% (4 days); ICANS in 34. 8%. Interventions included tocilizumab and corticosteroids.

Grades were frequently undocumented (CRS 62. 3%, ICANS 56. 1%); documented cases were mainly CRS grade 1 (59. 4%) and ICANS grade 2 (28. 0%). Performance was high on CRS and ICANS grading (accuracy of 0. 97 and 0. 91, respectively). Moderate performances were assessed for prespecified AE extraction (accuracies 0. 62-0. 76), and non prespecified AEs (accuracies 0.

76-0. 84). Inter rater reliability was strong to near perfect for CRS/ICANS presence and grade (kappa 0. 86-0. 96), moderate for dates and interventions, and weaker for broader AE attributes. LLM-derived insights can augment AE monitoring and real-world evidence generation by unlocking unstructured clinical detail and characteristic timelines after CAR T.

However, performance varied for broader AE attributes, warranting cautious use. Performance was highest for detecting the presence and grade of CRS and ICANS, with strong to near-perfect inter-rater reliability. While cautious use of LLMs for broad AE extraction is warranted due to the variable performance observed in this study, these results support integrating high performing CRS/ICANS extraction into EHR workflows.

论文信息

作者
Guillot J、Miao B、Suresh A、Sushil M、Williams CY、Vashisht R、Oskotsky TT、Sirota M
单位
Bakar Computational Health Sciences Institute, University of California, San Francisco, San Francisco, CA, USA.United States
文献类型
预印本
期刊
medRxiv : the preprint server for health sciences2026 May 5
原文标识
PubMed 42145597 · DOI 10.64898/2026.04.28.26351782