Thomas Deng Headshot

THOMAS DENG

Robotics • Autonomous Vehicles • Machine Learning Research

About Me

Hello! I am a Coterminal Student at Stanford studying Computer Science. I work on robotics, autonomous vehicles, and machine learning at the Interactive Perception and Robot Learning Lab as part of the Stanford AI Lab, and also at the Stanford Robotics Center. I build and research systems that combine learning-based architectures with classical robotics, focusing on understanding and improving world and VLA models and learning-based perception.

Research and Publications

LIBERO-JEPA

Contrastive Action Grounding for Object-Centric World Models

Accepted at the Robotics: Science and Systems (RSS) SemRob Workshop.

Developed an object-centric world model for robotic manipulation using VideoSAUR slot representations and a JEPA-style transformer predictor. Discovered that standard world model objectives learn nearly action-invariant dynamics despite explicit access to robot actions. Proposed a contrastive action loss that improves action grounding and long-horizon semantic prediction on LIBERO-100.

Lunar Autonomy Challenge

Full Stack Navigation, Mapping, and Planning for the Lunar Autonomy Challenge

1st place in the NASA Lunar Autonomy Challenge · Published in ION GNSS+.

AI-driven lunar rover navigation pipeline for the NASA Lunar Autonomy Challenge. Integrated LangSAM for rock segmentation, Depth Anything V2 for depth estimation, stereo disparity for geometry, and arc-based Dubins trajectories for safe local planning and waypoint following within NAVLab’s full autonomy stack.

GRPO VLA Fine Tuning

Reinforcement Learning for Vision-Language-Action Models

Exploring stability and performance of Group Relative Policy Optimization (GRPO) for VLA models. Implemented robotics RL infrastructure and developed methods for improving performance across LIBERO manipulation benchmarks.

Safe Robot Steering

Safe Robot Steering

Analyzing GRPO RL fine-tuning effects on SmolVLA across LIBERO tasks. Implemented mechanistic interpretability techniques for VLA models.

Egocentric Human-to-Robot Hand Retargeting for Imitation Learning Training

Pipeline that converts egocentric human videos into robot manipulation data: Detectron for hand detection, ViTPose for 2D keypoints, HaMeR for 3D hand reconstruction into MANO parameters, then dexterous retargeting with inverse kinematics onto a robot hand. The resulting trajectories are warped and randomized with Real2Render to generate large-scale robotic data for imitation learning. (Randomization and trajectory warping were handled within the Real2Render repo.)

NAVLab Research

Stanford NAVLab Autonomous Vehicles Research

Path-planning and simulation research using Unreal Engine's AirSim and Carla.

Fun Projects

ClaudeGen Robo

ClaudeGen: Autonomous Skill Generation for Dexterous Manipulation

Developed an autonomous data generation pipeline where Claude writes, debugs, and iteratively improves robot motion scripts in NVIDIA Isaac Sim. By closing the loop between LLM-generated code and simulation feedback, the system produces large-scale manipulation demonstrations without human teleoperation.

MIT THINK

MIT THINK Finalist Project

Deployment of a collectively optimized swarm of autonomous micro UAVs for accurate 3D reconstruction of destroyed buildings during disaster recovery.

FireSense

FireSense

Wildfire detection system using CNNs on an autonomous drone.

DroneDome

DroneDome

Neural network-powered drone UI visualized with Swift.

PropStop

PropStop

Naive Bayes classifier for NBA player point predictions, built as a Stanford CS109 project.

NLP Research

Task-Specific Optimizations for GPT-2

Stanford CS224N final project investigating parameter-efficient fine-tuning with Low-Rank Adaptation (LoRA) on GPT-2 to reduce training costs without sacrificing accuracy.

Akeer Foundation

The Akeer Foundation

Akeer Foundation: My friends and I are founding a K-12 school in South Sudan. I believe that education and opportunity are inalienable rights, and every step toward universal access sets a precedent of hope for the hundreds of millions who never had these liberties.

Personal

I enjoy playing basketball and soccer. On the road to win an IM championship. As a Philly native, I am a lifelong 76ers fan, unfortunately... Also, current Warehaus dorm ping pong champion!

I have also dabbled in music production! Here is a track I made:

Contact

Email: tdeng23@stanford.edu

GitHub: github.com/thomasdeng2027