Despite the importance of tactile sensing for reliable manipulation,
most existing Vision-Language-Action (VLA) datasets remain vision-only,
and those that do incorporate tactile information typically lack the joint
combination of task diversity, language conditioning, and action trajectories.
Furthermore, existing teleoperation pipelines rarely provide haptic feedback to the operator,
despite its established role in demonstration quality and manipulation stability.
In this work, we present HapTile,
a contact-grounded visuotactile manipulation dataset that advances beyond
vision-only trajectory datasets by embedding physical interaction sensing
at two levels: fingertip tactile feedback at the robot end-effector, and
haptic-informed demonstrations at the teleoperator side. The data collection
platform integrates haptic feedback directly into the teleoperation
controller, enabling the operator to perceive contact interactions in real time.
It is built around a standard and reproducible robotic system equipped with
custom-designed fingertip tactile sensors.
The dataset comprises everyday manipulation tasks spanning a broad range of
contact-rich skills, including pick-and-place, folding, pressing, stacking,
and other routine activities. Each task is paired with language instructions
that condition the policy on the manipulation objective, together with
synchronized visuotactile observations and action trajectories. In addition,
we provide a benchmarking study on contact-rich policy learning using
two baseline models to evaluate the effectiveness of the proposed
contact-grounded dataset.