A JavaScript client supports typed input and microphone capture, lets users confirm or edit transcribed speech, renders conversational turns, plays the generated response video, and packages text plus videos for download.
Multimodal
AI Agent
A research-grounded digital-human prototype that connects persona-conditioned conversation, speech, generated voice, lip-synchronized video, and compressed long-session memory.
Digital legacy becomes harder when it can speak back.
Traditional digital legacies - photos, files, messages, and social accounts - preserve records of a person. Generative AI introduces a different possibility: an agent that can produce new content, respond to people and its environment, and potentially evolve after the represented person has died.
That interactivity may support remembrance, companionship, and family heritage, but it also creates difficult questions that static archives do not: Who authors and governs the agent? Which appearance, memories, values, and knowledge count as the “same person”? How should the system evolve without inventing an identity? When does continued presence support the living, and when does it become intrusive?
Research gap
Existing products and studies often begin with bereaved people creating an agent of someone already deceased. This project instead foregrounded the perspective of people whose own identities and data would be represented, examining what they would want to encode, permit, restrict, and leave behind.
Prototype scenario
An authenticated user can type or speak to a persona-controlled digital human. The system transcribes speech, generates a context-aware reply, synthesizes it in a reference voice, produces a lip-synchronized video response, and lets the user retain or clear the resulting conversation artifacts.
A browser experience orchestrating language and GPU media services.
The prototype separates interaction, application orchestration, and model execution. The browser manages login, text chat, microphone recording, transcription confirmation, video playback, history download, and session clearing. Flask coordinates requests and conversation state, while server-side AI services handle language, speech, voice cloning, and video generation.
Flask authenticates prototype accounts, routes each turn, maintains persona and memory state, invokes speech and video models, and organizes generated artifacts by user and response sequence.
Two conversation pipelines
The user types a message in the conversation interface.
The route associates the turn with the active prototype identity and conversation.
The persona prompt, compacted history, and recent turns form the model context.
The language service returns a persona-conditioned response within the configured token budget.
The client adds the response to the visible history and re-enables input.
MediaRecorder captures an OGG audio blob from the microphone.
The server converts audio and returns text that the user may correct before generation.
The confirmed message enters the same persona and compressed-history pipeline.
XTTS v2 clones the reference voice; Wav2Lip aligns the generated speech with a reference face video.
Flask returns the reply and per-user MP4 path for immediate playback.
Wav2Lip is the active lip-synchronization path in the implemented Flask application. The repository also contains MuseTalk as an explored alternative, but it is not presented here as part of the active request path.
Preserve identity instructions without letting history grow forever.
Each conversation begins with a configurable system prompt that defines the represented persona, language, tone, and relationship. Every user and assistant message is tokenized and appended to the active context.
The application records the token length of every message and compares the running total with the model’s configured context and response allowance.
When the conversation approaches the limit, the oldest eligible segment is summarized with a BART model instead of dropping all prior context.
The persona instruction remains first; its compact history summary is followed by recent unsummarized turns before the next GPT request.
This is a pragmatic prototype memory mechanism, not a claim of faithful autobiographical memory. Long-term identity consistency would require provenance, consent-aware retrieval, editable memory, and stronger separation between source evidence and model-generated content.
A future production system could ground each response in an identity database that separates enduring traits from time-specific experiences and multimodal representations.
Relatively stable traits, beliefs, communication style, preferences, and user-defined boundaries.
Time-bounded stories linked to people, places, roles, artifacts, and the age at which they occurred.
Consent-approved photos, recordings, and visual or vocal references associated with a specific life stage.
Every record carries provenance, access scope, temporal context, edit history, and permission to generate from it.
Write-back rule: model-generated statements should never become autobiographical facts automatically. New memories require provenance and explicit human review before entering the identity store.
Text, voice, generated video, and history in one interaction.
The interaction artifact corresponds to implemented browser behaviors and Flask routes: account entry, text conversation, microphone-driven video mode, downloadable history, and explicit session clearing.
A working agent made the human risks impossible to ignore.
Building the prototype showed that a text-to-video digital human was technically possible, but technical capability alone could not answer whether such an agent should exist or how it should behave. A convincing simulation can intensify uncanny-valley effects, misrepresent a person, prolong grief, expose intimate data, or act beyond what the represented person and their family intended. The project therefore shifted from “how can we build it?” to “what would responsible design require?”
The CHI study interviewed people about agents representing their own identities after death - a perspective often missing when systems are designed primarily for bereaved users.
Study and analysis
Remote Mandarin interviews examined attitudes, differences from traditional digital legacy, expectations across the agent life cycle, interaction design, and ethical, legal, social, and technical concerns. Two researchers coded the data iteratively, reconciled interpretations in weekly meetings, and continued recruitment until thematic saturation.
-
VALUE & ACCEPTANCE
Potential support depends on personal beliefs and family context.
Participants saw possibilities for remembrance, guidance, and family heritage, but no single form or use was universally desirable.
-
IDENTITY
Resemblance involves appearance, knowledge, thinking, and evolution.
A visually convincing avatar is insufficient if its values, memories, behavior, or later changes contradict the represented person.
-
LIFE CYCLE
Control must extend beyond a single conversation.
Encoding, access, updating, transfer, preservation, and deletion all require explicit authorship and permission decisions.
-
BOUNDARIES
Interactivity can support the living or become intrusive.
Bidirectional conversation, proactive behavior, and embodied presence need controls for timing, intensity, context, and disengagement.
-
RISK
Failure is emotional, social, ethical, and technical.
Concerns included mental health, privacy, security, reputation, ownership, inequality, and effects on family relationships.
Build for identity continuity and bounded presence.
-
Appearance and expression
Support meaningful life-stage representations and audience-sensitive presentation without treating visual realism as identity itself.
-
Knowledge, thinking, and values
Bound external AI knowledge so it enriches interaction without contradicting the person’s experiences, beliefs, or core traits.
-
Managed evolution
Use permissioned updates and version control to balance adaptation with stable identity markers.
-
Bidirectional interaction
Give the living control over interaction intensity and include clear ways to pause or disengage.
-
Proactivity
Make timing, frequency, context, and type of unsolicited interaction explicitly configurable.
-
Physical presence
Prefer flexible virtual forms; use physical embodiment only for clear needs and keep it easy to adapt or remove.
A strong foundation, not a production-ready afterlife service.
The interview sample was grounded primarily in Chinese cultural contexts, where family continuity, filial relationships, and beliefs about death shape expectations. Broader cultural work is needed. The engineering implementation is also a research prototype: its in-memory global conversation state and file-based user artifacts are suitable for demonstrating a pipeline, not secure multi-user production.
Separate verified source memories, user edits, summaries, and generated content so every response can expose where its identity claims came from.
Support represented-person permissions, family roles, access scopes, succession rules, revocation, export, and deletion across the agent life cycle.
Move state into authenticated per-user sessions and durable stores, isolate GPU media jobs in a queue, and return progress asynchronously.
Measure latency, transcription recovery, audiovisual fidelity, identity consistency, emotional safety, and longitudinal effects with appropriate safeguards.
Project lead across research framing and system implementation.
- Led the project’s conceptual framing, research execution, synthesis, and paper writing as first author.
- Developed the Flask application, authentication flow, browser interaction, API routes, session controls, and per-user generated-media organization.
- Implemented persona-conditioned GPT conversation and token-aware history compression.
- Integrated microphone capture, speech recognition, XTTS v2 reference-voice synthesis, GPU-based Wav2Lip generation, FFmpeg media processing, and browser playback.
- Translated interview findings into concrete requirements around identity consistency, bounded interaction, governance, and future system safety.