Sharing the Human World with AI: Relationship-Scoped Multimodal Answerability Through Embodied Co-Experience
Abstract
Large language models can discuss a family recipe, a quiet street, crowding, and fatigue fluently, but those expressions are not organized by the unfolding conjunction of heat, sound, movement, hesitation, physiological load, and later consequence in one person’s life. We propose relationship-scoped multimodal answerability: a voluntary process in which an AI companion integrates authorized environmental, behavioral, physiological, interactional, and linguistic streams over time. Modality-specific encoders, temporal fusion, partner calibration, editable episodic memory, and interaction heads form a revisable representation of how situations unfold for this partner. Language supplies context and repair without becoming a compulsory translation layer; AI-initiated inquiry is optional. The design specifies a conceptual personal-guardian role for confidentiality, effective repair and exit, and external-action control, keeping governance at the boundary and ordinary interaction comparatively free from routine pre-screening. This does not solve Harnad’s symbol grounding problem by definition. It creates a testable, world-involving capacity whose status as partner-hosted grounding remains an open hypothesis.