Portrait of Pengfei Zhang

Pengfei Zhang

PhD Student in Computer Science, University of California, Irvine

Pengfei Zhang is a PhD candidate in Computer Science at the University of California, Irvine. His research sits at the intersection of human understanding, multimodal large language models, and multimodal agentic reinforcement learning.

His work spans two main directions. He builds VLM-based multimodal agents for human-centric image and video understanding, generation and editing, covering motion representation, action tokenization and identity-preserving multi-turn editing. He also works on large speech language models and the joint understanding of text, video, speech and audio, including post-training methods that align models to low-resource speeches.

Pengfei Zhang is applying for 2027 New Grad roles with F1 OPT. Expected graduation: Dec 2026, Mar 2027 or Jun 2027.

News

Research Interests

Human Understanding · Multimodal Large Language Models · Multimodal Agentic RL

Selected Projects and Publications

I. VLM-based Multimodal Agents for Human-Centric Image/Video Understanding, Generation, and Editing

Understanding, generation and editing of human-centric images and videos with complex motions, as essential personalized functions of multimodal LLMs and multimodal agents.

Techniques: VLM-based Multimodal Agent Orchestration, MLLM Post-Training (DPO, GRPO, GDPO)

VLM-based Multimodal Agents for Human-Centric Image/Video Understanding, Generation, and Editing
Identity Decay in Multi-Turn Human-Centric Image Editing: A Benchmark and a Multimodal Agentic Correction Framework

I.e Identity Decay in Multi-Turn Human-Centric Image Editing: A Benchmark and a Multimodal Agentic Correction Framework

Under Review

Introduces a benchmark measuring identity and fidelity decay in multi-turn human-centric image editing, and a correction pipeline in which a multimodal agent orchestrates identity correction using pretrained human segmentation and parsing models (Sapiens for body, DML_CSR for face) together with deterministic image operators such as TPS warp and high-frequency preservation transfer, flattening the degradation curve.

II. Speech Language Models and Agents in Low-Resource Settings

Joint understanding of multimedia signals (texts, videos, speeches, audios, etc), especially in low-resource settings.

Techniques: MLLM Post-Training (DPO, GRPO, GDPO), Chain-of-Thought (CoT) Distillation

Speech Language Models and Agents in Low-Resource Settings
AP-GRPO: Anchor-Gated Phonetic Alignment with Policy Optimization for Pathological Speech Reconstruction

II.f AP-GRPO: Anchor-Gated Phonetic Alignment with Policy Optimization for Pathological Speech Reconstruction

EMNLP 2026 (accepted)

A GRPO framework with a phonetic reward that aligns speech language models to the original speech signal through audible-anchor preservation and inter-anchor phonetic compatibility, combining an anchor-gated reward that matches anchors in clear regions with an inter-anchor phonetic alignment reward that checks whether recovered content is phonetically supported by the corrupted span.

AudioLens: Multi-Perspective Speech Clustering with Audio-Language Models

II.c AudioLens: Multi-Perspective Speech Clustering with Audio-Language Models

Under Review

Introduces audio multi-perspective clustering, where a model partitions speech recordings according to a natural-language perspective while inferring both the number of clusters and their assignments. Contributes AudioLens-Bench, a benchmark evaluating cross-perspective generalization, and AudioLens-R1, a large audio-language model trained with reasoning distillation and direct preference optimization.

Hallucination Rates of Speech Language Models on Pathological Speech Recognition: A Pilot Multi-Cohort Audit

II.a Hallucination Rates of Speech Language Models on Pathological Speech Recognition: A Pilot Multi-Cohort Audit

BSN 2025

An evaluation of popular speech language models on pathological speech recognition across hallucination types: insertion, fluent completion, entity fabrication and semantic substitution.

III. Personalized LLM Agents on Healthcare

Personalized conversational healthcare agents, tailored to individual patients.

Techniques: Steering Vectors for LLM Personalization, Retrieval-Augmented Generation (RAG)

Personalized LLM Agents on Healthcare
Can the Environment Speak for Itself? T2-GRPO: Turn-Trajectory Group Relative Policy Optimization for Caregiver Agents

III.c Can the Environment Speak for Itself? T2-GRPO: Turn-Trajectory Group Relative Policy Optimization for Caregiver Agents

Under Review

A turn-trajectory optimization scheme for long-horizon caregiver agents that derives dense turn-level rewards from patient-state transitions, normalizes them separately from trajectory-level outcomes via centered-rank group-relative advantages, and adds a binary safety veto to enforce hard behavioral constraints.

Working Experience

Production Projects

Awards