About Me

Present: Incoming CS PhD at Stony Brook University.

Previously: Graduated from the Viterbi School of Engineering at USC for my MSCS, where I worked with Sai Praneeth Karimireddy and collaborated with Robin Jia.

During my undergrad, I was advised by Dr. Jinho Choi at Emory NLP, and I explored Biblical dialogue with LLMs. Here is an overview of this project. I also spent time auditing Carl Yang’s lab with additional guidance from Jiaying Lu on ML and healthcare.

Research Interests

I am broadly interested in developing trustworthy AI through a data-centric lens, where I’ve explored adversarial attacks, biases, and privacy. I’m currently thinking about improving diversity of outputs in LLMs and understanding causes of mode-collapse.

News

  • July 2026: ContextLeak has been accepted to COLM 2026. Excited for SF!
  • Apr 2026: I’ve committed to starting my CS PhD in Fall 2026 at Stony Brook! 🥳
  • Dec 2025: Also check out our ContextLeak preprint!
  • Jun 2025: Our two workshop papers, “Auditing Privacy-Preserving In-Context Learning Methods.” was accepted at L2M2 Workshop @ ACL, and “ContextLeak: Auditing Leakage in Private In-Context Learning Methods.” was published at MemFM Workshop @ ICML!
  • Sep 2024: Our paper, What is Your Favorite Gender, MLM? Gender Bias Evaluation in Multilingual Masked Language Models, was published at Information!
  • Aug 2024: Started my MSCS at USC! ✌️
  • May 2024: I graduated from Emory College of Arts and Science with a B.S. in Computer Science with Highest Honors and with a minor in German studies! 🦅
  • Aug 2023: Completed my REU at University of Colorado, Colorado Springs with Jugal Kalita.

Research Experience

  • Graduate Student Researcher, Foundations of Robust And Trustworthy (FORT) ML, USC, 01/2025 - 05/2026
  • Undergraduate Student Researcher, Emory NLP. 08/2023 - 08/2024
  • Visiting Researcher, University of Colorado at Colorado Springs, Colorado Springs. 05/2023 - 08/2023

Papers

  • ContextLeak: Auditing Leakage in Private In-Context Learning Methods.
    COLM 2026
    ICML Workshop 2025 - The Impact of Memorization on Trustworthy Foundation Models (MemFM)
    Jacob Choi*, Shuying Cao*, Xingjian Dong*, Amin Banayeeanzade, Wang Bill Zhu, Robin Jia, Sai Praneeth Karimireddy
    paper; poster

  • Auditing Privacy-Preserving In-Context Learning Methods.
    ACL Workshop 2025 - L2M2: The First Workshop on Large Language Model Memorization
    Jacob Choi*, Shuying Cao*, Xingjian Dong*, Sai Praneeth Karimireddy

  • What is Your Favorite Gender, MLM? Gender Bias Evaluation in Multilingual Masked Language Models.
    Information, 2024. Jeongrok Yu, Seong Ug Kim, Jacob Choi, Jinho Choi.
    Project; Poster; Slides

Additional Projects

  • Large Language Models With Religious Text.
    SouthNLP Symposium, 2024. Jacob Choi, Jinho Choi.
    paper; poster

  • Latent Separability of Backdoor Attacks on Language Models
    UCCS REU Symposium, 2023. Jacob Choi, Jugal Kalita.
    paper; slides

Awards

Misc. Writing

During my undergrad, I served as treasurer and wrote for Emory In Via’s Journal of Christian Thought (Augustine Collective), where you can find blog posts, poems, songs, and personal/academic pieces written by students that relate to faith. Here are some of my pieces: