Almene De Meran Meguimtsop

Almene De Meran Meguimtsop

AI Safety & Alignment • LLM Evaluation & Red Teaming • AI for Science

About Me

Almene De Meran Meguimtsop

I am a Ph.D. student in Computer Science at the University of Colorado Boulder, advised by Prof. Daniel Acuña. My research focuses on AI safety and alignment, particularly how to evaluate and improve the reliability of large language models and agentic AI systems in scientific and other high-stakes settings. I develop benchmarks and evaluation methods to study how AI systems respond to adversarial, ambiguous, and norm-sensitive requests, with current work spanning research integrity, model disparities, and trustworthy AI for science.

More broadly, I am interested in understanding why failures emerge in increasingly capable AI systems and developing methods that help them behave reliably under real-world constraints. My work draws on natural language processing, red teaming, machine learning, and alignment methods, and I am particularly interested in research that connects rigorous evaluation with practical safeguards.

My path to computer science has been interdisciplinary. Before starting my Ph.D., I earned a Master’s in Physics, a Master’s in Data Science, and a postgraduate diploma in Quantitative Life Sciences, experiences that exposed me to different ways of approaching scientific problems. This background continues to shape my research and my interest in developing AI systems that can support scientific work reliably and responsibly.

Recent News

[09/2026] I am serving as Lead of CU Café, a graduate student and postdoc community at CU Boulder focused on mentoring, professional development, and community building.

🔬 [06–07/2026] Presented How Willing Are LLMs to Commit Scientific Fraud? A Study of 16 Commercial and Open Models as a talk at ICSSI 2026 and as a poster at the Machine Learning Summer School 2026 at Columbia University.

📝 [05/2026] Our new preprint SciIntBench: Measuring LLM Compliance with Research Integrity Norms Under Adversarial Framing is now available on arXiv and is under review at ACL Rolling Review. Project Page · Code & Data

🤖 [05/2026] Our paper The Disparate Beliefs of Large Language Models About Science and Scientists is under review at the Proceedings of the National Academy of Sciences.

🌱 [04/2026] Nominated as the Professional Development Chair for CU Boulder's SMART Program, coordinating professional development programming and mentoring undergraduate STEM researchers throughout the summer research program.

🎓 [04/2026] Selected as the Computer Science representative on CU Boulder's Graduate Student Advisory Board for the College of Engineering and Applied Science.

🏆 [03/2026] Selected as a 2026–2027 Newton Leadership Fellow at the University of Colorado Boulder.

🎤 [05/2025] Presented Disparities in Large Language Models for and About Science at the Atlanta Conference on Science and Innovation Policy.

Research & Publications

My research focuses on AI safety, evaluation, alignment, and trustworthy AI for scientific and other high-stakes settings.

Selected Research

Featured ResearchPreprint · Under review at ACL Rolling Review

SciIntBench: Measuring LLM Compliance with Research Integrity Norms Under Adversarial Framing

An adversarial benchmark for evaluating whether large language models uphold research integrity norms under overt, covert, and benign framing.

Meguimtsop, A. D. M., Pacheco, M. L., & Acuna, D. E.

arXiv preprint, 2026

Under review

The Disparate Beliefs of Large Language Models About Science and Scientists

Examines disparities in how large language models represent and reason about science and scientists.

Meguimtsop, A. D. M., Ojukwu, C. E., Taechoyotin, P., Chávez-Ruelas, C., Burke, R., Clauset, A., & Acuna, D. E.

Manuscript under review at Proceedings of the National Academy of Sciences (PNAS)

Under resubmission

Do LLMs Know When Science Has Been Retracted? Evaluating Retraction Awareness for Trustworthy Science Automation

Evaluates retraction awareness in large language models as a component of trustworthy AI-assisted scientific workflows.

Taechoyotin, P., Tian, Y., Meguimtsop, A. D. M., & Acuna, D. E.

Manuscript

Under review

Audio Penalty: Evaluating Clinical Safety of Multimodal LLMs Across Text and Speech in Low-Resource Languages

Examines the clinical safety of multimodal large language models across text and speech in low-resource language settings.

Oduwole, M., Abdullahi, T., Olatunji, T., Katuka, G. A., Mgonzo, M., Okocha, C., Oko-Odion, T., Ezema, K., Meguimtsop, A. D. M., & Ismaila, L. E.

ACL Rolling Review

Selected Talks & Presentations

How Willing Are LLMs to Commit Scientific Fraud? A Study of 16 Commercial and Open Models
Lightning Talk

How Willing Are LLMs to Commit Scientific Fraud? A Study of 16 Commercial and Open Models

Presented findings from our evaluation of large language models' responses to requests involving scientific misconduct and research integrity.

5th International Conference on the Science of Science and Innovation (ICSSI 2026)
Boulder, Colorado
July 2026
How Willing Are LLMs to Commit Scientific Fraud? A Study of 16 Commercial and Open Models
Poster Presentation

How Willing Are LLMs to Commit Scientific Fraud? A Study of 16 Commercial and Open Models

Presented a research poster on the safety and reliability of large language models when assisting with requests that may violate scientific integrity norms.

Machine Learning Summer School (MLSS) 2026
Columbia University, New York City
June 2026
Disparities in Large Language Models for and About Science
Conference Talk

Disparities in Large Language Models for and About Science

Presented research examining disparities in large language models in scientific and academic contexts.

Atlanta Conference on Science and Innovation Policy (ATLC 2025)
Atlanta, Georgia
May 2025

Selected Experience

Research, teaching, and mentoring experiences spanning AI safety, machine learning, data science, and interdisciplinary scientific computing.

Research
August 2024 – Present

Graduate Research Assistant

Science of Science & Computational Discovery Lab, University of Colorado Boulder

  • Conduct research on AI safety, alignment, reliability, and disparities in large language models, with a focus on scientific and other high-stakes workflows.
  • Develop benchmarks and evaluation frameworks to assess the safety, reliability, and research-integrity behavior of frontier LLMs and agentic AI systems.
  • Evaluate commercial and open-weight models using red teaming, adversarial prompting, and large-scale model evaluation pipelines.
  • Investigate model monitoring, mechanistic interpretability, and scalable oversight methods for understanding and mitigating failures in LLMs and agentic systems.
Teaching
Fall 2024 · Spring 2026 · Fall 2026

Graduate Teaching Assistant

Department of Computer Science, University of Colorado Boulder

  • Teaching Assistant for Machine Learning, Intro to Data Science with Probability & Statistics, and Intro to Computational Thinking.
  • Support students through office hours, teaching sessions, project mentoring, and technical guidance in machine learning, data science, statistics, and Python programming.
  • Grade quizzes, assignments, projects, and exams while helping students strengthen problem-solving and computational skills.
Research Mentoring
Spring 2026

Tutor

African Institute for Mathematical Sciences (AIMS)

  • Mentored two AIMS students on research essay projects in collaboration with their supervisors.
  • Provided research guidance on low-resource automatic speech recognition and Chichewa machine translation using traditional machine translation systems and large language models.
Teaching
November 2020 – May 2022

Teaching Assistant, Physics

University of Dschang

  • Delivered tutorials in optoelectronics to first-year Master's students and supported assessment and grading.
  • Led undergraduate physics laboratory sessions, guided students through experiments, and graded assignments.

Professional Involvement

Leadership, mentoring, academic service, and community engagement within and beyond CU Boulder.

Program Committee Member
September 2026
  • Selected to serve as a Program Committee member for the 41st AAAI Conference on Artificial Intelligence (AAAI-27), reviewing submissions and contributing to the conference peer-review process.
Lead
Fall 2026 – Spring 2027
  • Lead a graduate student and postdoc STEM community, organizing student-driven panels, faculty conversations, and professional development events.
  • Recruit speakers and coordinate campus discussions on topics including AI ethics, equity in AI, and issues identified by the graduate research community.
Professional Development Chair
Summer 2026
  • Coordinated professional development programming for undergraduate STEM researchers, working with graduate mentors, speakers, and panelists.
  • Mentored two undergraduate research interns throughout the 10-week summer research program.
Computer Science Representative
2026 – 2027
  • Represent Computer Science graduate students, advocate for departmental needs, and contribute to college-wide initiatives supporting the graduate student community.
Newton Leadership Fellow
2026 – 2027
  • Selected for a leadership development fellowship combining individualized coaching, experiential learning, professional development, and an interdisciplinary cohort experience.
University Educator (UE) / Mentor
Spring 2026 – Present
  • Support inquiry-based STEM learning with K–12 students through hands-on activities and mentorship.
  • Help foster scientific curiosity, confidence, science identity, and a sense of belonging among participating students.
Lead
Fall 2025 – Spring 2026
  • Led a graduate community initiative focused on uplifting the achievements of Black women in STEM and Education through peer support, community building, and leadership programming.

Additional Academic & Professional Service

Peer Reviewer & Science Judge
2025 – Present
  • Serve as a peer reviewer for Humanities and Social Sciences Communications, the International Conference on Computational Social Science (IC2S2), and Deep Learning Indaba.
  • Serve as a poster and science judge for undergraduate research and pre-college STEM competitions, including the Undergraduate Research Expo, Eco-Innovation Challenge, Buckeye Science & Engineering Fair, Ohio Academy of Science Virtual Science Day, and Texas DECA.