- Source
- arXiv
- Published
- Runtime
- 0:00
- Snippets
- 5
A conversation between
Human-Centric Intelligence in the Era of Foundation Models: A Survey
§02
Snippets
-
Human-centric intelligence in the foundation-model era integrates six levels: visual appearance, spatial geometry, kinematic dynamics, interaction modeling, world simulation, and embodied agency.
A unified taxonomy connects previously isolated methods and reveals how to build AI that truly understands human behavior at scale.
-
Human-centric models must view humans as observable subjects, dynamic actors with kinematic behavior, and situated agents embedded in world simulation and embodied agency.
This hierarchy moves beyond static recognition to capture how humans interact with environments and each other—essential for embodied AI and robotics.
-
The field needs human-centric data families, computational architectures, and training/inference strategies specifically designed for foundation-model-scale human intelligence.
Borrowing scaling and transfer principles from language/vision models can unlock more capable and generalizable human understanding systems.
-
The survey organizes datasets, benchmarks, and evaluation metrics across all six human-centric levels to enable standardized, multi-level assessment.
Unified evaluation frameworks reveal gaps and synergies between levels, guiding development of more comprehensive human-understanding systems.
-
Human-centric intelligence must be scalable, trustworthy, physically grounded, and deployable—balancing generalization with real-world reliability and safety.
This reframes human-centric AI as an engineering and ethics challenge, not just a perception problem, advancing human-AI collaboration.
§03
Synthesis
Human-Centric Intelligence in the Foundation Model Era: A Survey
Foundation models have transformed AI by scaling up to billions of parameters and learning general-purpose representations across modalities. Yet human-centric intelligence—understanding people's appearance, actions, interactions, and agency—has lagged behind in harnessing these advances. This survey argues that the field remains fragmented across specialized tasks and research silos, missing an opportunity to unify human understanding under a coherent conceptual framework aligned with how foundation models operate.
The Core Framework
The authors propose a six-level taxonomy for organizing human-centric intelligence. Rather than treating each sub-problem separately, they structure human understanding vertically: starting with observable attributes (visual appearance and spatial geometry), moving through behavior (kinematic dynamics and interaction modeling), and culminating in agency (world simulation and embodied reasoning). This taxonomy is designed to expose methodological and conceptual connections that typically remain hidden when researchers work in isolation on pose estimation, action recognition, or human-object interaction.
The key insight is that foundation models excel at learning transferable, general representations across massive heterogeneous datasets. Human-centric tasks should be reconsidered through the same lens: can we build unified models that understand humans across visual, kinematic, and interactive dimensions, rather than training separate specialists for each task?
Methodological Foundations and Organization
The survey covers three pillars that enable this unification. First, it catalogs human-centric data families—what types of human observations exist and how they're annotated. Second, it reviews computational architectures: transformer-based models, vision-language fusion, and multi-modal encoders that can handle the diverse signals needed to understand humans. Third, it discusses training and inference optimization strategies—techniques like parameter-efficient fine-tuning, knowledge distillation, and efficient inference that make foundation models practical for deployment.
The authors then systematically review methods at each taxonomic level, connecting representative approaches to the broader landscape. They organize associated datasets, benchmarks, and metrics, making it easier to identify what's been solved and where bottlenecks remain.
Why It Matters
Three practical motivations emerge. Scalability: foundation models scale gracefully; human-centric AI should too. Trustworthiness: as these systems are deployed for surveillance, health monitoring, and autonomous systems, they must be reliable and interpretable. Physical grounding: humans exist in physical space and interact with objects; models must respect these constraints rather than learning shallow correlations.
The fragmentation problem is concrete. A researcher building a pose estimator might miss that a vision-language model pre-trained on human-object interactions could accelerate their work. A team working on embodied AI (robots that learn to act in environments) may not see connections to action recognition datasets. By mapping the conceptual and methodological landscape, the survey reveals reusable patterns across these domains.
The authors have released a systematically organized, continuously updated collection of literature and resources, positioning this as a living reference rather than a static snapshot. This matters because the field is moving rapidly—new foundation model capabilities emerge monthly, and human-centric applications expand into robotics, AR/VR, healthcare, and autonomous systems.
The survey doesn't claim breakthrough results, but rather provides clarity on what exists and where progress is needed. It is fundamentally a call to unify an otherwise scattered ecosystem around the capabilities and principles that have made foundation models successful.
Mine your own.
Lode is a workbench, not a feed. Paste a YouTube URL. The model proposes a transcript, a set of quote-grounded snippets, a synthesis essay, and the fan-out. You decide what stays.