← 返回

使用大语言模型从接受 CAR-T 细胞治疗患者的临床记录中提取不良事件性质、严重程度、时间线及相应干预

英文原题:Extracting adverse event nature, severity, timelines and resulting interventions from clinical notes of patients receiving CAR-T cell therapy using large language models.

查看英文原题

Extracting adverse event nature, severity, timelines and resulting interventions from clinical notes of patients receiving CAR-T cell therapy using large language models.

PubMed 2026/08/27(内容时间) PLOS Digit Health Q1 · IF 7.7(JCR 2025)

分数与星级只用于站内排序 —— 不代表疗效、安全性或个人适用性。

中文摘要

CAR-T 细胞疗法,即经基因工程改造、靶向肿瘤抗原的患者 T 细胞,已改变血液系统恶性肿瘤的诊疗,但需要仔细追踪不良事件(AEs),而这些事件往往仅记录在非结构化的电子健康记录(EHR)笔记中。

我们在 UCSF 的安全环境中评估了一种基于大语言模型(LLM)的方法,用于提取六种商业化 CAR-T 产品(2012-2023)输注后 30 天内的 AEs、日期、分级和干预措施,并以两名评估者为基准进行对比。在零样本设置下使用 GPT 4 0314,并采用四个提示(预设 AEs、非预设 AEs、CRS、ICANS),我们将输出与 50 份笔记随机样本的双重标注进行比较,使用准确率、精确率、召回率、F1 和 Cohen's kappa。在 293 名患者(中位年龄 65.6)的 4,762 份病程记录中,CRS 发生于 80.2%(中位发生时间 4 天);中性粒细胞减少症 70.0%(16 天);中性粒细胞减少性发热 64.8%(4 天);ICANS 为 34.8%。

干预措施包括 tocilizumab 和皮质类固醇。分级经常未记录(CRS 62.3%,ICANS 56.1%);有记录的病例主要为 CRS 1 级(59.4%)和 ICANS 2 级(28.0%)。在 CRS 和 ICANS 分级方面表现较高(准确率分别为 0.97 和 0.91)。预设 AE 提取(准确率 0.62-0.76)和非预设 AEs(准确率 0.76-0.84)评估为中等表现。评分者间信度(IRR)在 CRS/ICANS 存在和分级方面较强(kappa 0.86-0.96),在日期和干预措施方面为中等,在更广泛的 AE 属性方面较弱。LLM 衍生的见解可通过解锁 CAR-T 后的非结构化临床细节和特征性时间线,增强 AE 监测和真实世界证据生成。

然而,对于更广泛的 AE 属性,性能存在差异,因此需要谨慎使用。检测和分级 CRS 和 ICANS 的性能最高,IRR 为强至近乎完美。尽管由于本研究中观察到的性能差异,LLMs 需要谨慎使用,但这些结果支持在受控的基于 EHR 的研究或试点环境中进一步评估有监督的 CRS/ICANS 提取,并进行监测和前瞻性验证。

展开英文摘要原文

Chimeric Antigen Receptor T-cell (CAR T) therapy, genetically engineered patient T cells targeting tumor antigens, has transformed care for hematologic malignancies but requires careful tracking of adverse events (AEs) often documented only in unstructured electronic health record (EHR) notes.

We evaluated a Large Language Model (LLM)-based approach in UCSF's secure environment to extract AEs, dates, grades, and interventions within 30 days post infusion for six commercial CAR T products (2012-2023), benchmarking against two evaluators. Using GPT 4 0314 in a zero shot setting with four prompts (prespecified AEs, non prespecified AEs, CRS, ICANS), we compared outputs against dual annotations on a random sample of 50 notes using accuracy, precision, recall, F1, and Cohen's kappa. From 4,762 progress notes for 293 patients (median age 65. 6), CRS occurred in 80. 2% (median onset 4 days); neutropenia 70. 0% (16 days); neutropenic fever 64. 8% (4 days); ICANS in 34. 8%. Interventions included tocilizumab and corticosteroids.

Grades were frequently undocumented (CRS 62. 3%, ICANS 56. 1%); documented cases were mainly CRS grade 1 (59. 4%) and ICANS grade 2 (28. 0%). Performance was high on CRS and ICANS grading (accuracy of 0. 97 and 0. 91, respectively). Moderate performances were assessed for prespecified AE extraction (accuracies 0. 62-0. 76), and non prespecified AEs (accuracies 0.

76-0. 84). Inter rater reliability (IRR) was strong for CRS/ICANS presence and grade (kappa 0. 86-0. 96), moderate for dates and interventions, and weaker for broader AE attributes. LLM derived insights can augment AE monitoring and real world evidence generation by unlocking unstructured clinical detail and characteristic timelines after CAR T.

However, performance varied for broader AE attributes, warranting cautious use. Performance was highest for detecting and grading CRS and ICANS, with strong to near-perfect IRR. While cautious use of LLMs is warranted due to the variable performance observed in this study, these results support further evaluation of supervised CRS/ICANS extraction in controlled EHR-based research or pilot settings, with monitoring and prospective validation.

论文信息

作者
Guillot J、Miao B、Suresh A、Sushil M、Williams CYK、Vashisht R、Oskotsky TT、Sirota M
单位
Bakar Computational Health Sciences Institute, University of California, San Francisco, San Francisco, California, United States of America.United States
期刊
PLOS digital health2026 Aug
原文标识
PubMed 42658846 · DOI 10.1371/journal.pdig.0001426