
About Me

I am a Ph.D. student in Computer Science at the University of Colorado Boulder, advised by Prof. Daniel Acuña. My research focuses on AI safety and alignment, particularly how to evaluate and improve the reliability of large language models and agentic AI systems in scientific and other high-stakes settings. I develop benchmarks and evaluation methods to study how AI systems respond to adversarial, ambiguous, and norm-sensitive requests, with current work spanning research integrity, model disparities, and trustworthy AI for science.
More broadly, I am interested in understanding why failures emerge in increasingly capable AI systems and developing methods that help them behave reliably under real-world constraints. My work draws on natural language processing, red teaming, machine learning, and alignment methods, and I am particularly interested in research that connects rigorous evaluation with practical safeguards.
My path to computer science has been interdisciplinary. Before starting my Ph.D., I earned a Master’s in Physics, a Master’s in Data Science, and a postgraduate diploma in Quantitative Life Sciences, experiences that exposed me to different ways of approaching scientific problems. This background continues to shape my research and my interest in developing AI systems that can support scientific work reliably and responsibly.
Recent News
Research & Publications
My research focuses on AI safety, evaluation, alignment, and trustworthy AI for scientific and other high-stakes settings.
Selected Research
SciIntBench: Measuring LLM Compliance with Research Integrity Norms Under Adversarial Framing
An adversarial benchmark for evaluating whether large language models uphold research integrity norms under overt, covert, and benign framing.
Meguimtsop, A. D. M., Pacheco, M. L., & Acuna, D. E.
arXiv preprint, 2026
The Disparate Beliefs of Large Language Models About Science and Scientists
Examines disparities in how large language models represent and reason about science and scientists.
Meguimtsop, A. D. M., Ojukwu, C. E., Taechoyotin, P., Chávez-Ruelas, C., Burke, R., Clauset, A., & Acuna, D. E.
Manuscript under review at Proceedings of the National Academy of Sciences (PNAS)
Do LLMs Know When Science Has Been Retracted? Evaluating Retraction Awareness for Trustworthy Science Automation
Evaluates retraction awareness in large language models as a component of trustworthy AI-assisted scientific workflows.
Taechoyotin, P., Tian, Y., Meguimtsop, A. D. M., & Acuna, D. E.
Manuscript
Audio Penalty: Evaluating Clinical Safety of Multimodal LLMs Across Text and Speech in Low-Resource Languages
Examines the clinical safety of multimodal large language models across text and speech in low-resource language settings.
Oduwole, M., Abdullahi, T., Olatunji, T., Katuka, G. A., Mgonzo, M., Okocha, C., Oko-Odion, T., Ezema, K., Meguimtsop, A. D. M., & Ismaila, L. E.
ACL Rolling Review
Selected Talks & Presentations

How Willing Are LLMs to Commit Scientific Fraud? A Study of 16 Commercial and Open Models
Presented findings from our evaluation of large language models' responses to requests involving scientific misconduct and research integrity.
Boulder, Colorado

How Willing Are LLMs to Commit Scientific Fraud? A Study of 16 Commercial and Open Models
Presented a research poster on the safety and reliability of large language models when assisting with requests that may violate scientific integrity norms.
Columbia University, New York City

Disparities in Large Language Models for and About Science
Presented research examining disparities in large language models in scientific and academic contexts.
Atlanta, Georgia
Selected Experience
Research, teaching, and mentoring experiences spanning AI safety, machine learning, data science, and interdisciplinary scientific computing.
Graduate Research Assistant
- •Conduct research on AI safety, alignment, reliability, and disparities in large language models, with a focus on scientific and other high-stakes workflows.
- •Develop benchmarks and evaluation frameworks to assess the safety, reliability, and research-integrity behavior of frontier LLMs and agentic AI systems.
- •Evaluate commercial and open-weight models using red teaming, adversarial prompting, and large-scale model evaluation pipelines.
- •Investigate model monitoring, mechanistic interpretability, and scalable oversight methods for understanding and mitigating failures in LLMs and agentic systems.
Graduate Teaching Assistant
- •Teaching Assistant for Machine Learning, Intro to Data Science with Probability & Statistics, and Intro to Computational Thinking.
- •Support students through office hours, teaching sessions, project mentoring, and technical guidance in machine learning, data science, statistics, and Python programming.
- •Grade quizzes, assignments, projects, and exams while helping students strengthen problem-solving and computational skills.
Tutor
- •Mentored two AIMS students on research essay projects in collaboration with their supervisors.
- •Provided research guidance on low-resource automatic speech recognition and Chichewa machine translation using traditional machine translation systems and large language models.
Professional Involvement
Leadership, mentoring, academic service, and community engagement within and beyond CU Boulder.
- ●Selected to serve as a Program Committee member for the 41st AAAI Conference on Artificial Intelligence (AAAI-27), reviewing submissions and contributing to the conference peer-review process.
- ●Lead a graduate student and postdoc STEM community, organizing student-driven panels, faculty conversations, and professional development events.
- ●Recruit speakers and coordinate campus discussions on topics including AI ethics, equity in AI, and issues identified by the graduate research community.
- ●Coordinated professional development programming for undergraduate STEM researchers, working with graduate mentors, speakers, and panelists.
- ●Mentored two undergraduate research interns throughout the 10-week summer research program.
- ●Represent Computer Science graduate students, advocate for departmental needs, and contribute to college-wide initiatives supporting the graduate student community.
- ●Selected for a leadership development fellowship combining individualized coaching, experiential learning, professional development, and an interdisciplinary cohort experience.
Partnerships for Informal Science Education in the Community (PISEC), University of Colorado Boulder
- ●Support inquiry-based STEM learning with K–12 students through hands-on activities and mentorship.
- ●Help foster scientific curiosity, confidence, science identity, and a sense of belonging among participating students.
- ●Led a graduate community initiative focused on uplifting the achievements of Black women in STEM and Education through peer support, community building, and leadership programming.
Additional Academic & Professional Service
- ●Serve as a peer reviewer for Humanities and Social Sciences Communications, the International Conference on Computational Social Science (IC2S2), and Deep Learning Indaba.
- ●Serve as a poster and science judge for undergraduate research and pre-college STEM competitions, including the Undergraduate Research Expo, Eco-Innovation Challenge, Buckeye Science & Engineering Fair, Ohio Academy of Science Virtual Science Day, and Texas DECA.