Hong T. Nguyen

USC. Los Angeles. 3237879669. Live a life you will never regret.

hong.jpg

RTH 318, USC

Los Angeles, CA 90089

I am a Ph.D. student in Electrical and Computer Engineering at the University of Southern California (USC), expected to graduate in December 2027. I am a Research Assistant in the Signal Analysis and Interpretation Laboratory (SAIL), advised by Prof. Shrikanth Narayanan. Previously, I received my Bachelor’s degree from Hanoi University of Science and Technology (HUST), supervised by Prof. Chuyen Nguyen Thanh.

My research focuses on human-centered video understanding and temporal dynamics, in the context of medical imaging, clinical trials, and seamless dyadic interaction. I develop explainable multimodal models of human behavior across video and speech, including video world models and diffusion models for real-time MRI of the vocal tract (Arti-JEPA, Speech2rtMRI), motion-centric video encoders (MOOSE), and severity representation learning for medical images (ConPro). I work with multidisciplinary teams spanning law enforcement, healthcare, linguistics, and psychology. Feel free to reach out for research collaboration at hongn@usc.edu.

Beyond technical expertise, I am also boardgame lover and hiker. Yet, I have only visited 4/63 National Parks in the US.

news

Sep 09, 2026 New preprint: Arti-JEPA adapts a video world model to real-time MRI of the vocal tract for speech-production analysis. Read it on arXiv :rocket:
May 03, 2026 Our paper on interpretable modeling of articulatory temporal dynamics from real-time MRI for phoneme recognition appears at ICASSP 2026.
Jun 01, 2025 Preprint of MOOSE, a motion-centric video encoder that pays attention to temporal dynamics via optical flow, is out on arXiv.
Apr 06, 2025 My paper Speech2rtMRI, the first speech-guided diffusion model for real-time MRI video of the vocal tract, appears at ICASSP 2025 :sparkles:
Oct 24, 2024 This website is on air...

selected publications

  1. arXiv
    Arti-JEPA: Adapting Video World Model to Real-Time MRI of the Vocal Tract for Speech-Production Analysis
    Hong Nguyen, Sean Foley, Christina Hagedorn, and 4 more authors
    arXiv preprint arXiv:2609.09757, 2026
  2. ICASSP
    Interpretable Modeling of Articulatory Temporal Dynamics from Real-Time MRI for Phoneme Recognition
    Jay Park, Hong Nguyen, Sean Foley, and 4 more authors
    In ICASSP 2026 - 2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2026
  3. arXiv
    MOOSE: Pay Attention to Temporal Dynamics for Video Understanding via Optical Flows
    Hong Nguyen, Dung Tran, Hieu Hoang, and 2 more authors
    arXiv preprint arXiv:2506.01119, 2025
  4. ICASSP
    Speech2rtMRI: Speech-Guided Diffusion Model for Real-time MRI Video of the Vocal Tract during Speech
    Hong Nguyen, Sean Foley, Kevin Huang, and 3 more authors
    In ICASSP 2025 - 2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2025
  5. CVPRW
    ConPro: Learning Severity Representation for Medical Images using Contrastive Learning and Preference Optimization
    Hong Nguyen, Hoang Nguyen, Melinda Chang, and 3 more authors
    In 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), 2024