CASE STUDY / 01 WEAROS · FEW-SHOT ML · CLIENT–SERVER

WatchGuardian

A personalized, just-in-time behavior intervention system that learns a user-defined action from only ten examples and delivers support on a round WearOS watch.

Role AI Software Engineer / Research Intern
Organization University of Washington
Platform Google Pixel Watch · WearOS
Outcome Journal paper · ACM Transactions on Computing for Healthcare · 2026
WatchGuardian workflow from user-defined actions and wearable data recording through augmentation, model customization, and smartwatch intervention
01 / THE PROBLEM

Personal interventions beyond predefined behaviors.

Just-in-time interventions can help people interrupt health-related behaviors at the moment support is needed. Yet most wearable systems target a predefined set of common behaviors, while individuals may want help with personal undesirable actions such as nail biting, skin picking, or lip tearing. Some of these are body-focused repetitive behaviors that can affect physical, mental, and social well-being.

These actions—and the movements that reveal them—vary substantially from person to person. Supporting user-defined targets therefore requires a system that can learn the behavior each user cares about from minimal setup data, recognize it during daily life, and intervene at the right moment without demanding a large, continuously labeled personal dataset.

User scenario

A user defines an action they want to reduce, records a small set of examples on the watch, and receives a just-in-time prompt when the personalized model recognizes that action during daily life.

02 / SYSTEM ARCHITECTURE

One wearable client, two server-side data pipelines.

WatchGuardian separates responsibilities across three boundaries. The Pixel Watch owns sensing and user interaction; a persistent bidirectional socket carries sensor batches and results; Python/PyTorch services prepare data, customize models on GPU compute, store personalized artifacts, and run live inference.

KotlinJetpack ComposeWearOSSensorManagerCoroutinesTCP socketsJSONPythonPyTorchA100 GPU cluster
CLIENT / PIXEL WATCH Kotlin Wear OS application

Jetpack Compose supports action selection, guided recording, training status, live detection, and intervention screens. SensorManager captures tri-axial linear acceleration on a background thread, while coroutines stream data without blocking the interface.

DATA TRANSPORT TCP SOCKET ⇄ Length-framed JSON Sensor batches → ← Result + confidence
SERVER / GPU COMPUTE Python + PyTorch services

Server services prepare normalized windows, train a personalized model on the A100 GPU cluster, store the best artifact by user and action, and serve smoothed live predictions.

Data preparationGPU customizationModel storeLive inference

Two data pipelines

MODE 01 / CUSTOMIZATIONRecord examples and create a personal detector.
WATCHDefine & record

The user selects or names an action and follows guided few-shot recording rounds.

→
TRANSPORTSend examples

Labeled accelerometer batches and action context travel to the customization service.

→
SERVERPrepare data

The service associates the stream with the user and action, then builds model-ready windows.

→
GPUCustomize & validate

A GPU job adapts the detector and selects the strongest checkpoint.

→
ARTIFACTStore model

The personalized model is saved for that user-action pair.

MODE 02 / LIVE INTERVENTIONRecognize the action and return support in real time.
WATCHCapture live signal

The wearable continuously collects accelerometer data while intervention mode is active.

→
TRANSPORTStream batches

A separate live session sends sensor batches and receives compact prediction responses.

→
SERVERBuild rolling window

The inference service converts the latest stream into the model's input window.

→
MODELInfer & smooth

The saved personal detector runs inference; consecutive-positive filtering suppresses transient predictions.

→
WATCHPrompt user

A confirmed result triggers vibration, an on-watch reminder, and cooldown behavior.

WatchGuardian interface for choosing a predefined or self-defined action and recording few-shot examples
Customization client flow: define the target and record a small set of guided examples.
WatchGuardian live mode detecting a selected action and delivering a just-in-time intervention
Intervention client flow: run the personalized detector and deliver a vibration plus an on-watch reminder.

Customization and live intervention use separate server endpoints so long-running training work does not share the latency-sensitive inference path. Model weights remain server-side; the watch receives only status updates or compact prediction results.

03 / MODEL CUSTOMIZATION

A reusable feature extractor, personalized from a few examples.

The learning pipeline separates reusable representation learning from per-user adaptation. Stages 1 and 2 create a shared feature extractor before customization; Stage 3 turns a small personal recording set into a detector for the user's newly defined action.

STAGE 1Self-supervised foundation

A five-layer ResNet is adopted from a model pre-trained on more than 700,000 person-days of UK Biobank wearable data.

STAGE 2Fine-grained adaptation

All layers are fine-tuned on public hand-activity datasets plus self-collected negative daily-activity data.

STAGE 3Personal few-shot learning

On the GPU cluster, user recordings become 5-second windows; six augmentations and positive-negative synthesis expand them approximately 143 times before a lightweight classification head is trained.

WatchGuardian three-stage learning pipeline: self-supervised pre-training, fine-tuning on hand activities, and few-shot personalization with augmentation and synthesis
Three-stage few-shot model-customization pipeline. Only Stage 3 repeats when a user adds a new target action.
04 / VALIDATION

Validate the detector, then the intervention.

The evaluation follows the product logic: first test whether the few-shot model can recognize personal actions, then test whether its just-in-time prompts help people reduce the action they most want to change.

Five predetermined actions and twelve categories of self-defined personal actions collected for WatchGuardian evaluation
Evaluation action set - five predetermined actions plus 12 distinct categories contributed through participants' self-defined actions, for 17 action types in total.
01 / OFFLINE MODEL EVALUATIONCan it learn a new action from only a few examples?

Twenty-six participants each recorded the five predetermined actions and one self-defined action. Their recordings formed the dataset used to test different numbers of shots and customized actions. With ten examples of one new target action, the pipeline achieved 87.7% action-level accuracy and an 87.3% F1 score.

26Data contributors
87.7%Accuracy at 10 shots
87.3%F1 at 10 shots
02 / REAL-TIME INTERVENTION STUDYDoes timely detection change behavior?

In the follow-up within-subject study, 21 valid participants selected the one action, among the six they had recorded, for which they felt the strongest need for intervention. Their personalized model was trained from ten shots and compared with a rule-based baseline that issued reminders every ten minutes.

64%Reduction with WatchGuardian
35%Reduction with rule-based prompts
+29 ppWatchGuardian advantage

WatchGuardian reduced target-action duration to 36% of each participant's pre-intervention level. The improvement over the rule-based condition remained significant after controlling for the number of notifications.

05 / DISCUSSION

What this prototype reveals - and what comes next.

  • PERCEPTION
    Timely AI feedback changes more than behavior.

    Some participants perceived WatchGuardian's vibration as stronger or believed they performed the action more often, even when objective measurements indicated otherwise. Accurate, well-timed feedback can heighten self-awareness, but designers must also consider anxiety, perceived surveillance, and emotional interpretation.

  • HUMAN-AI RELATIONSHIP
    Intervention should feel collaborative, not punitive.

    Participants imagined the system as a watchdog, playful companion, coach, or mind reader. Future systems should let people configure reminder style and intensity, correct predictions, revise goals, and understand why an intervention was triggered.

  • CONTEXT & ADAPTATION
    The right moment depends on what the person is doing.

    Prompts may support behavior change yet interrupt focused work. A more mature JITAI should adapt to task engagement, learn from user feedback, and evolve its strategy rather than repeat the same alert pattern indefinitely.

  • LONG-TERM SUPPORT
    Combine immediate intervention with reflective self-tracking.

    WatchGuardian focuses on in-the-moment support. Historical trends, progress visualizations, changing goals, and longitudinal patterns could help people reflect on improvement and sustain behavior change beyond individual reminders.

  • LIMITATIONS & FUTURE WORK
    Validate beyond this accelerometer-only proof of concept.

    The current system depends on server-side processing, stable connectivity, one target action, dominant-wrist sensing, and a short controlled study. Future work should explore secure or on-device inference, additional sensing modalities, multiple simultaneous actions, broader populations, varied postures and contexts, and longer in-the-wild deployments.

06 / MY SCOPE & COLLABORATION

Project leadership within a cross-disciplinary team.

My contribution

  • Initiated and led the research exploration, system design, evaluation, and paper writing.
  • Independently built the WearOS experience for accelerometer sensing, sample capture, and intervention prompts.
  • Implemented the client–server ML workflow for processing, customization, training, and inference.
  • Connected model behavior to a usable real-time interaction rather than a notebook-only evaluation.
  • Worked with cross-institution collaborators while owning the project’s core research and engineering direction.
16Authors
8Institutions

Publication team

Grouped by affiliations listed in the publication. This shows the breadth of the collaboration without assigning individual contribution claims.

Simon Fraser UniversityYing Lei
Columbia UniversityYancheng Cao, Will Ke Wang, Chunhua Weng, Randy Auerbach, Lena Mamykina, Xuhai Xu
Stanford UniversityYuanzhe Dong
The Ohio State UniversityChangchang Yin, Weidan Cao, Ping Zhang
Nationwide Children's HospitalJingzhen Yang
Northeastern UniversityBingsheng Yao, Dakuo Wang
Weill Cornell MedicineYifan Peng
University of WashingtonYuntao Wang
SYSTEM DEMO Watch the personalized intervention workflow.
NEXT CASE STUDY / 02 FamilyCanvas →