R2I-Bench: Benchmarking Reasoning-Driven Text-to-Image Generation
EMNLP 2025 Outstanding Paper Award Nomination & Oral (Top 0.55%) Evaluation
Selected Research (* indicates equal contribution)
R2I-Bench: Benchmarking Reasoning-Driven Text-to-Image GenerationEMNLP 2025 Outstanding Paper Award Nomination & Oral (Top 0.55%) Evaluation
GUI Agents for Continual Game GenerationFindings of EMNLP 2026 Self-Evolution
MVI-Bench: A Comprehensive Benchmark for Evaluating Robustness to Misleading Visual Inputs in LVLMsICML 2026 Evaluation
MemMark: State-Evolution Attribution Watermarking for Agent Long-Term Memory SystemsFindings of EMNLP 2026 Self-Evolution
STRIDE: Strategic Trajectory Reasoning via Discriminative Estimation for Verifiable Reinforcement LearningPreprint 2026 Post-Training
Sell More, Play Less: Benchmarking LLM Realistic Selling SkillEMNLP 2026 (main) Evaluation
EvoFlow: Evolving Diverse Agentic Workflows On The FlyPreprint 2025 Self-Evolution
SuperFlow: Training Flow Matching Models with RL on the FlyFindings of AACL 2026 Post-Training
|