Skip to content
Book Open access

Sharing the Human World with AI: Relationship-Scoped Multimodal Answerability Through Embodied Co-Experience

Oct 2026 · Companion Publication of the 28th International Conference on Multimodal Interaction · 1 citation · 42 references

Abstract

Large language models can discuss a family recipe, a quiet street, crowding, and fatigue fluently, but those expressions are not organized by the unfolding conjunction of heat, sound, movement, hesitation, physiological load, and later consequence in one person’s life. We propose relationship-scoped multimodal answerability: a voluntary process in which an AI companion integrates authorized environmental, behavioral, physiological, interactional, and linguistic streams over time. Modality-specific encoders, temporal fusion, partner calibration, editable episodic memory, and interaction heads form a revisable representation of how situations unfold for this partner. Language supplies context and repair without becoming a compulsory translation layer; AI-initiated inquiry is optional. The design specifies a conceptual personal-guardian role for confidentiality, effective repair and exit, and external-action control, keeping governance at the boundary and ordinary interaction comparatively free from routine pre-screening. This does not solve Harnad’s symbol grounding problem by definition. It creates a testable, world-involving capacity whose status as partner-hosted grounding remains an open hypothesis.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.