bioRxiv Science⌕ Search

Biology subjects

Raviv, H.

Publications and source records attributed to Raviv, H..

2 recordsLinked to original sources

The Richness of Experience: Agent-Based Learning Reveals theMechanisms of Early Language Acquisition

A longstanding challenge in developmental science is to understand how children learn language from naturalistic everyday input. To study this process, we leveraged the First 1,000 Days (1kD) dataset, which provides longitudinal, ultra-dense daily audiovisual recordings for individual children in their home environments. This unusually detailed, child-specific record of early experience enabled us to pair each childs rich language input with a cognitively grounded learning agent, linking naturalistic experience ("nurture") to internal learning mechanisms ("nature"). Trained incrementally on each childs input without prior linguistic knowledge, the learning agent discovered speech units corresponding to the English phoneme inventory and acquired thousands of words, closely mirroring individual developmental trajectories. Learning generalized across children while preserving individual differences in rate and timing. Interestingly, learning relied not only on linguistic input but also on the rehearsal of past experiences at the end of each training day. These findings demonstrate that everyday environments provide sufficient structure for language acquisition and establish a unified mechanistic framework for studying development in real-world contexts.

neuroscience↗

The First 1,000 Days (1kD) Project - Collecting and Analyzing an Ultra-Dense Naturalistic Dataset of Human Baby Development

Human development unfolds in continuous, multimodal environments across seconds, days, and years, yet most developmental datasets capture sparse, context-limited samples of everyday life. We introduce the First 1,000 Days (1kD) Project, an initiative designed to collect ultra-dense, longitudinal, child-centered data that capture developmental trajectories within their full ecological context. Fifteen U.S. homes with 17 infants were recorded 12-14 hours per day over a median of 944 days, yielding [~]1.18 million hours of raw audiovisual data. We present an end-to-end framework for large-scale longitudinal naturalistic measurement and a scalable analysis pipeline of the collected data. In a case study, we describe how we utilized our pipeline to isolate child-centered speech, resulting in the collection of 2,000 to 6,000 hours of transcribed speech for each infant. We demonstrate that dense sampling within the home environment reveals a stable, household-specific lexical structure, which sparse sampling methods consistently fail to capture. The 1kD project offers a blueprint for teams aiming to collect and analyze natural behavior at scale in real-world settings.

neuroscience↗