bioRxiv Science⌕ Search

Biology subjects

Vuong, Q. C.

Publications and source records attributed to Vuong, Q. C..

2 recordsLinked to original sources

Seeing Just Enough: The Contribution of Hands, Objects and Visual Features to Egocentric Action Recognition

Humans recognize everyday actions without conscious effort despite challenges such as poor viewing conditions and visual similarity between actions. Yet the visual features contributing to action recognition remain unclear. To address this, we combined semantic modelling and feature reduction methods to identify critical features for recognizing actions from challenging egocentric perspectives. We first identified egocentric action videos from home environments that a motion-focused action classification network could correctly classify (Easy videos) or not (Hard videos). In Experiment 1, participants (N=136) labelled the action and object in the videos. Using a language model framework, we derived human ground truth labels for each video and quantified its recognition consistency based on semantic similarity. Participants recognized actions and objects in Easy videos more consistently than in Hard videos. In Experiment 2, we recursively reduced the Easy and Hard videos with high recognition consistency to extract minimal recognizable configurations (MIRCs), in which any further spatial or temporal reductions disrupted recognition. The data was collected using a large-scale online study (N=4360). We extracted information related to the hand, objects, scene background and visual features (e.g., orientation or motion signals) from the 474 MIRCs. Binary classification showed that recognition was disrupted when regions containing the manipulated object and strong orientation signals were removed, while temporal reduction by frame-scrambling disrupted recognition in 73% of MIRCs. The active hand had some marginal contribution. Our results highlight the importance of both mid- and high-level information for egocentric action recognition and link hierarchical feature theories with naturalistic human perception.

neuroscience↗

ASAP: An automatic sustained attention prediction method for infants and toddlers using wearable device signals

Sustained attention (SA) is a critical cognitive ability that emerges in infancy. The recent development of wearable technology for infants enables the collection of large-scale multimodal data in the natural environment, including physiological signals. To capitalize on these new technologies, psychologists need methods to efficiently extract valid and robust SA measures from large datasets. In this study, we present an innovative automatic sustained attention prediction (ASAP) method that harnesses electrocardiogram (ECG) and accelerometer (Acc) signals recorded with wearable sensors from 75 infants (6-, 9-, 12-, 24- and 36-months). Infants undertook various naturalistic tasks similar to those encountered in their natural environment, including free play with their caregivers. Annotated SA was validated by fixation signals from eye-tracking. ASAP was trained on temporal and spectral features derived from the ECG and Acc signals to detect attention periods, and tested against human-coded SA. ASAPs performance is similar across all age groups, demonstrating its suitability for studying development. We also investigated the relationship between attention periods and low-level perceptual features (visual saliency, visual clutter) extracted from the egocentric videos recorded during caregiver-infant free play. Saliency increased during attention vs inattention periods and decreased with age for attention (but not inattention) periods. Crucially, there was no observable difference in results from ASAP attention detection relative to the human-coded attention. Our results demonstrate that ASAP is a powerful tool for detecting infant SA elicited in natural environments. Alongside the available wearable sensors, ASAP provides unprecedented opportunities for studying infant development in the wild.

neuroscience↗