Skip to content

Jingxi Qiu

I work on the reliability of large language models: how to verify whether retrieved evidence actually supports an answer, how to detect when it is insufficient, and how to make retrieval-augmented generation selective rather than credulous. My recent work (SURE-RAG, NEI-CAP) shows that common fact-verification benchmarks carry construction artifacts that inflate "not enough information" performance, and proposes ways to audit and mitigate them.

Before that I completed an M.S. in Data Science & Analytics at Georgetown University, where I built offline reinforcement-learning pipelines over large-scale longitudinal patient records and analysed disparities in access to CAR-T therapy across 125,000+ patients. I care about models whose confidence can be trusted — in language and in medicine alike.

Email / CV / Google Scholar / GitHub

Selected publicationsall publications →

Recent writingAll posts →