CASE STUDY / 03GENERATIVE AI · MULTIMODAL SYSTEMS · HUMAN-CENTERED AI

Multimodal
AI Agent

A research-grounded digital-human prototype that connects persona-conditioned conversation, speech, generated voice, lip-synchronized video, and compressed long-session memory.

RoleAI Software Engineer / Research Intern
OrganizationHong Kong University of Science and Technology
StackFlask · GPT · XTTS v2 · Wav2Lip
OutcomeCHI Best Paper HM · Top 5% · China Daily—HK Edition
Multimodal AI agent connecting conversation, voice, memory, persona, and generated video
01 / THE PROBLEM

Digital legacy becomes harder when it can speak back.

Traditional digital legacies - photos, files, messages, and social accounts - preserve records of a person. Generative AI introduces a different possibility: an agent that can produce new content, respond to people and its environment, and potentially evolve after the represented person has died.

That interactivity may support remembrance, companionship, and family heritage, but it also creates difficult questions that static archives do not: Who authors and governs the agent? Which appearance, memories, values, and knowledge count as the “same person”? How should the system evolve without inventing an identity? When does continued presence support the living, and when does it become intrusive?

Research gap

Existing products and studies often begin with bereaved people creating an agent of someone already deceased. This project instead foregrounded the perspective of people whose own identities and data would be represented, examining what they would want to encode, permit, restrict, and leave behind.

Prototype scenario

An authenticated user can type or speak to a persona-controlled digital human. The system transcribes speech, generates a context-aware reply, synthesizes it in a reference voice, produces a lip-synchronized video response, and lets the user retain or clear the resulting conversation artifacts.

02 / SYSTEM ARCHITECTURE

A browser experience orchestrating language and GPU media services.

The prototype separates interaction, application orchestration, and model execution. The browser manages login, text chat, microphone recording, transcription confirmation, video playback, history download, and session clearing. Flask coordinates requests and conversation state, while server-side AI services handle language, speech, voice cloning, and video generation.

JavaScriptMediaRecorderFlaskJSON / multipartOpenAI GPTGoogle Speech RecognitionXTTS v2Wav2LipPyTorch / CUDAFFmpeg
CLIENT / BROWSERConversation and media interface

A JavaScript client supports typed input and microphone capture, lets users confirm or edit transcribed speech, renders conversational turns, plays the generated response video, and packages text plus videos for download.

APPLICATION APIHTTPS⇄JSON requestsMultipart audioMedia paths
SERVER / GPU COMPUTEFlask orchestration + model services

Flask authenticates prototype accounts, routes each turn, maintains persona and memory state, invokes speech and video models, and organizes generated artifacts by user and response sequence.

ConversationSTTVoice cloningLip syncMedia store

Two conversation pipelines

PIPELINE 01 / TEXT CONVERSATIONKeep a lightweight conversational path for fast interaction.
BROWSERSubmit text

The user types a message in the conversation interface.

→
FLASKAttach session

The route associates the turn with the active prototype identity and conversation.

→
MEMORYBuild context

The persona prompt, compacted history, and recent turns form the model context.

→
GPTGenerate reply

The language service returns a persona-conditioned response within the configured token budget.

→
BROWSERRender turn

The client adds the response to the visible history and re-enables input.

PIPELINE 02 / VOICE-TO-VIDEOTransform one spoken turn into an editable, voiced, visual response.
BROWSERRecord speech

MediaRecorder captures an OGG audio blob from the microphone.

→
STTTranscribe & confirm

The server converts audio and returns text that the user may correct before generation.

→
GPT + MEMORYGenerate reply

The confirmed message enters the same persona and compressed-history pipeline.

→
GPU MEDIASynthesize & animate

XTTS v2 clones the reference voice; Wav2Lip aligns the generated speech with a reference face video.

→
BROWSERPlay response

Flask returns the reply and per-user MP4 path for immediate playback.

Wav2Lip is the active lip-synchronization path in the implemented Flask application. The repository also contains MuseTalk as an explored alternative, but it is not presented here as part of the active request path.

03 / PERSONA & MEMORY

Preserve identity instructions without letting history grow forever.

Each conversation begins with a configurable system prompt that defines the represented persona, language, tone, and relationship. Every user and assistant message is tokenized and appended to the active context.

01 / MONITORTrack the context budget

The application records the token length of every message and compares the running total with the model’s configured context and response allowance.

02 / COMPRESSSummarize older turns

When the conversation approaches the limit, the oldest eligible segment is summarized with a BART model instead of dropping all prior context.

03 / REBUILDRetain identity and recency

The persona instruction remains first; its compact history summary is followed by recent unsummarized turns before the next GPT request.

This is a pragmatic prototype memory mechanism, not a claim of faithful autobiographical memory. Long-term identity consistency would require provenance, consent-aware retrieval, editable memory, and stronger separation between source evidence and model-generated content.

FUTURE CONCEPT — NOT IMPLEMENTED IN THE CURRENT PROTOTYPEThe prototype above still uses persona instructions plus compressed conversation summaries. Everything below is a proposed next iteration.
PROPOSED RAG ARCHITECTURE / FUTURE WORKRetrieve identity evidence instead of asking a summary to represent an entire life.

A future production system could ground each response in an identity database that separates enduring traits from time-specific experiences and multimodal representations.

CORE IDENTITYPersonality, values, worldview

Relatively stable traits, beliefs, communication style, preferences, and user-defined boundaries.

LIFE TIMELINEEvents, relationships, life stages

Time-bounded stories linked to people, places, roles, artifacts, and the age at which they occurred.

MULTIMODAL IDENTITYAppearance and voice by age

Consent-approved photos, recordings, and visual or vocal references associated with a specific life stage.

GOVERNANCE METADATASource, consent, audience, confidence

Every record carries provenance, access scope, temporal context, edit history, and permission to generate from it.

01 / QUERYInterpret the conversation
→
02 / RETRIEVESemantic + metadata search
→
03 / RERANKRelevance, life stage, permission
→
04 / GROUNDBuild an evidence bundle
→
05 / GENERATERespond within identity boundaries

Write-back rule: model-generated statements should never become autobiographical facts automatically. New memories require provenance and explicit human review before entering the identity store.

WORKING PROTOTYPE

Text, voice, generated video, and history in one interaction.

The interaction artifact corresponds to implemented browser behaviors and Flask routes: account entry, text conversation, microphone-driven video mode, downloadable history, and explicit session clearing.

Implemented digital-human browser client with an AI-generated privacy-safe video stand-in and English conversation controls
Implemented browser client. Offline replay of the real HTML/CSS interface with English controls and a clearly labeled AI-generated, privacy-safe video stand-in; no GPU inference is required to render the client.
Digital human prototype flow showing login, text conversation, voice-to-video interaction, and downloadable history
Interaction design artifact. Authenticate, converse by text or voice, review the generated response, and retain or clear session artifacts.
04 / RESEARCH PIVOT

A working agent made the human risks impossible to ignore.

Building the prototype showed that a text-to-video digital human was technically possible, but technical capability alone could not answer whether such an agent should exist or how it should behave. A convincing simulation can intensify uncanny-valley effects, misrepresent a person, prolong grief, expose intimate data, or act beyond what the represented person and their family intended. The project therefore shifted from “how can we build it?” to “what would responsible design require?”

The CHI study interviewed people about agents representing their own identities after death - a perspective often missing when systems are designed primarily for bereaved users.

18Participants
60 min.Interviews
20-79Age range

Study and analysis

Remote Mandarin interviews examined attitudes, differences from traditional digital legacy, expectations across the agent life cycle, interaction design, and ethical, legal, social, and technical concerns. Two researchers coded the data iteratively, reconciled interpretations in weekly meetings, and continued recruitment until thematic saturation.

  • VALUE & ACCEPTANCE
    Potential support depends on personal beliefs and family context.

    Participants saw possibilities for remembrance, guidance, and family heritage, but no single form or use was universally desirable.

  • IDENTITY
    Resemblance involves appearance, knowledge, thinking, and evolution.

    A visually convincing avatar is insufficient if its values, memories, behavior, or later changes contradict the represented person.

  • LIFE CYCLE
    Control must extend beyond a single conversation.

    Encoding, access, updating, transfer, preservation, and deletion all require explicit authorship and permission decisions.

  • BOUNDARIES
    Interactivity can support the living or become intrusive.

    Bidirectional conversation, proactive behavior, and embodied presence need controls for timing, intensity, context, and disengagement.

  • RISK
    Failure is emotional, social, ethical, and technical.

    Concerns included mental health, privacy, security, reputation, ownership, inequality, and effects on family relationships.

05 / DESIGN IMPLICATIONS

Build for identity continuity and bounded presence.

01 / MAINTAINING IDENTITY CONSISTENCYKeep the agent recognizable as the represented person.
  • Appearance and expression

    Support meaningful life-stage representations and audience-sensitive presentation without treating visual realism as identity itself.

  • Knowledge, thinking, and values

    Bound external AI knowledge so it enriches interaction without contradicting the person’s experiences, beliefs, or core traits.

  • Managed evolution

    Use permissioned updates and version control to balance adaptation with stable identity markers.

02 / BALANCING INTRUSIVENESS & SUPPORTLet continued presence remain supportive rather than burdensome.
  • Bidirectional interaction

    Give the living control over interaction intensity and include clear ways to pause or disengage.

  • Proactivity

    Make timing, frequency, context, and type of unsolicited interaction explicitly configurable.

  • Physical presence

    Prefer flexible virtual forms; use physical embodiment only for clear needs and keep it easy to adapt or remove.

06 / LIMITATIONS & NEXT STEPS

A strong foundation, not a production-ready afterlife service.

The interview sample was grounded primarily in Chinese cultural contexts, where family continuity, filial relationships, and beliefs about death shape expectations. Broader cultural work is needed. The engineering implementation is also a research prototype: its in-memory global conversation state and file-based user artifacts are suitable for demonstrating a pipeline, not secure multi-user production.

01Provenance-aware memory

Separate verified source memories, user edits, summaries, and generated content so every response can expose where its identity claims came from.

02Consent and stewardship

Support represented-person permissions, family roles, access scopes, succession rules, revocation, export, and deletion across the agent life cycle.

03Production architecture

Move state into authenticated per-user sessions and durable stores, isolate GPU media jobs in a queue, and return progress asynchronously.

04Evaluate the experience

Measure latency, transcription recovery, audiovisual fidelity, identity consistency, emotional safety, and longitudinal effects with appropriate safeguards.

07 / MY SCOPE & COLLABORATION

Project lead across research framing and system implementation.

  • Led the project’s conceptual framing, research execution, synthesis, and paper writing as first author.
  • Developed the Flask application, authentication flow, browser interaction, API routes, session controls, and per-user generated-media organization.
  • Implemented persona-conditioned GPT conversation and token-aware history compression.
  • Integrated microphone capture, speech recognition, XTTS v2 reference-voice synthesis, GPU-based Wav2Lip generation, FFmpeg media processing, and browser playback.
  • Translated interview findings into concrete requirements around identity consistency, bounded interaction, governance, and future system safety.
4Authors
4Institutions
Ying LeiSimon Fraser University
Shuai MaAalto University
Yuling SunFudan University
Xiaojuan MaHong Kong University of Science and Technology
NEXT CASE STUDY / 04SafeQA →