HiTune project banner

QuantiPhy: A Quantitative Benchmark Evaluating Physical Reasoning Abilities of Vision-Language Models

Advised by Fei-Fei Li and Ehsan Adeli | Stanford University

QuantiPhy is the first benchmark that asks vision–language models to do physics with numerical accuracy. Across 3,300+ video–text instances, we show that today’s VLMs often sound plausible but fail quantitatively on physical reasoning tasks—they rely more on memorized world knowledge from pretraining than on the actual video and text inputs.

CVPR 2026 · Read paper →
Real-Time Watermarking for VR Motion

Real-Time Watermarking for VR Motion

Lead Researcher, supervised by Prof. Dianna Xu and Prof. Aline Normoyle

Developed a real-time watermark embedding system using Quantization Index Modulation (QIM), optimized for VR motion data with angular velocity-based filtering. Simulated various real-world distortions—including frame rate resampling, partial frame loss, and motion blending—to test watermark robustness. Evaluated performance using both text and image watermarks, comparing accuracy, visual impact, and resilience under different attack types.

View slides →
visitors