About Me

I am an M.Sc. student in Artificial Intelligence at the University of Tehran, working at the intersection of large language models, reinforcement learning, and human preference alignment. My research focuses on instruction following and theory of mind in LLMs.

I have two peer-reviewed papers: one published at ACL 2026 and one accepted at Computers in Human Behavior.

I have industry experience as an NLP Engineer for two years. I have also a research intern experience at ESSEC Business School in France on dynamic curriculum learning, and a research assistant in the NLP Lab at University of Tehran, working on human preference alignment and role-playing in LLMs.

Feel free to reach out if you find my work interesting or would like to collaborate.

News

Education

M.Sc. in Computer Engineering – Artificial Intelligence

Sep 2023 – Present

University of Tehran, Tehran, Iran

GPA: 17.95 / 20

Thesis: The Impact of Human Preference Alignment on Large Language Models Role-Playing

B.Sc. in Electrical Engineering

Sep 2018 – Dec 2022

Isfahan University of Technology, Isfahan, Iran

GPA: 15.28 / 20

Final Project: Android app for diagnosing musculoskeletal abnormalities

Publications

(*) indicates first author

MDP-GRPO: Stabilized Group Relative Policy Optimization for Multi-Constraint Instruction Following (*)

ACL 2026 Main Conference

aclanthology.org/2026.acl-long.1982

  • Identified and formalized three failure modes of GRPO under discrete, low-dispersion reward settings: low-variance amplification, mean-centering blindness, and zero-variance collapse
  • Proposed MDP-GRPO, a stabilized RL framework for multi-constraint instruction following with verifiable rewards, improving training stability and achieving higher constraint satisfaction

Evaluating Role-Prompted LLM Responses on Autism-Related Theory-of-Mind Tasks (*)

Accepted at Computers in Human Behavior (Q1 Journal, Impact Factor: 12.2)

Preprint — papers.ssrn.com

  • Proposed a DSM-5–grounded evaluation framework for assessing theory-of-mind reasoning in LLMs under autism-spectrum–inspired cognitive constraints
  • Analyzed how prompting strategies and alignment mechanisms affect the separability and consistency of simulated cognitive profiles across multiple ToM task types and statistical properties

Notable Experiences

Natural Language Processing Engineer

Mar 2025 – Mar 2026

MCI Lab, R&D Department — MCI, Tehran, Iran

  • Developed and deployed resource-efficient LLM pipelines using supervised fine-tuning and reinforcement learning to improve instruction-following performance, including vLLM-based inference and serving
  • Used MLflow for experiment tracking

Research Intern, ODCL Project

May 2026 – Present

ESSEC Business School, Cergy, France

Supervisors: Prof. Emiliano Traversi, Dr. Pegah Alizadeh, Dr. Massinissa Hamidi

  • Developed a dynamic curriculum learning framework that adapts task ordering during training via learnable edge-preference parameters, formulated as a graph-based optimization problem

Research Assistant, NLP Lab

Dec 2024 – Present

University of Tehran, Tehran, Iran

Supervisors: Prof. Heshaam Faili, Dr. M.J. Dousti

  • Conducted research on human preference alignment and role-playing in LLMs, contributing to experimental validation and peer-reviewed manuscripts

Research Assistant, Cognitive Science Lab

Spring & Summer 2025

University of Tehran, Tehran, Iran

Supervisor: Dr. Abdol-Hossein Vahabie

  • Investigated theory-of-mind and social reasoning capabilities of LLMs using cognitively inspired experimental designs grounded in psychological and DSM-5 frameworks

Teaching Assistant

Natural Language Processing

Fall 2025

University of Tehran, Iran — Teacher: Prof. Heshaam Faili

nlp-ut.github.io

  • Designed a course project and hands-on workshop on fine-tuning LLMs using LoRA/PEFT, covering data preparation, tokenizer analysis, training, and standardized evaluation

Large Language Models

Spring 2025

University of Tehran, Iran — Teachers: Dr. M.J. Dousti, Dr. Yadollah Yaghoobzadeh

  • Developed assignments on in-context learning and human preference alignment, introducing students to prompt engineering and reinforcement learning from human feedback

Deep Learning

Fall 2024, Spring 2025

University of Tehran, Iran — Teacher: Dr. Ahmad Kalhor

  • Designed assignments on image captioning, focusing on CNN–RNN architectures and attention mechanisms

Projects

Selected Additional Projects

Scholarships & Academic Service

Skills

AI & ML

Machine Learning Deep Learning Statistical Inference Reinforcement Learning

NLP & LLMs

PEFT (LoRA, QLoRA) RLHF RAG Natural Language Inference Agentic LLMs

Frameworks & Systems

PyTorch TensorFlow vLLM Distributed Training MLflow Optuna

Programming & Tools

Python C++ JavaScript Git LaTeX Docker