CASE STUDY / 04RAG · NLP · STORYTELLING AGENT

SafeQA

A human-centered, retrieval-augmented pipeline that turns moments in familiar stories into grounded questions about real-world child safety.

RoleAI Software Engineer / Research Intern
OrganizationEast China Normal University
StackInformation retrieval · LLM generation (141 knowledge entries · 157 QA pairs)
OutcomeBachelor Thesis · 2023
A fairy-tale princess pauses before accepting an apple from a mysterious stranger
Generated scene illustrating the project’s core interaction: use a familiar story moment to prompt reflection about a related real-world safety decision.
01 / THE PROBLEM

Make abstract safety rules concrete, memorable, and easier to revisit.

Child safety education is important, yet abstract rules can be difficult for young children to understand and remember. Parents may also lack the time, specialized knowledge, or confidence to provide safety instruction consistently. SafeQA investigates whether familiar story events can become situated prompts for discussing how a child should respond in a comparable real-life situation.

STORY MOMENTA stranger offers Snow White an apple.
→
SAFETY CONNECTIONShould you accept food from someone you do not know?
→
REAL-LIFE ACTIONRefuse politely and ask a trusted adult.

What the formative study changed

Seven parents of children aged 5-9 participated in semi-structured interviews. The findings shaped three requirements: integrate safety learning into recurring everyday activities; anchor abstract guidance in simple, imaginable situations; and use an agent to assist, rather than assume unlimited parent time or expertise.

01Repeat in context

Safety learning should recur naturally instead of depending only on occasional formal instruction.

02Show a situation

Concrete story scenes give children an accessible mental model for otherwise abstract risks.

03Support caregivers

Automation can help parents surface relevant material while preserving adult involvement.

02 / RESEARCH PIPELINE

Move from caregiver needs to grounded generation - and validate each layer separately.

01 / DISCOVERFormative study

Interview seven parents about current practices, constraints, and design needs.

→
02 / STRUCTUREBuild the data foundation

Create 141 structured safety entries and annotate 157 story-knowledge-QA pairs.

→
03 / RETRIEVEMatch safety knowledge

Rank curated knowledge against each story passage.

→
04 / GENERATEProduce grounded QA

Condition GPT-3 on the story, retrieved knowledge, and selected examples.

→
05 / EVALUATETest model and experience

Run offline module tests, then compare complete outputs with parents.

The formative study and final parent evaluation involved the same seven participants. The project therefore treats their feedback as an exploratory design signal, not population-level evidence.

03 / TECHNICAL PIPELINE

Retrieve explicit safety knowledge before asking the language model to generate.

The architecture separates factual grounding from language generation. A learned ranker connects a story passage to curated safety entries; a frozen GPT-3 model then receives that evidence through a structured few-shot prompt. This design makes the intended lesson explicit instead of relying only on knowledge implicit in the language model.

PythonPyTorchSentence-BERTDual-tower MLPGradient reversalBPR lossGPT-3
SafeQA pipeline from story passage through safety knowledge retrieval and grounded QA generation
The implemented end-to-end pipeline: explicit curated knowledge supervises retrieval; selected evidence and demonstrations ground frozen GPT-3 generation.
MODULE A / RETRIEVAL

Learn a story-to-knowledge ranking.

Sentence-BERT embeds story sections and safety entries. Separate MLP towers learn task-specific representations. A gradient-reversal discriminator discourages overfitting to the small set of individual stories, while pairwise BPR loss teaches relevant knowledge to rank above sampled negatives.

MODULE B / GENERATION

Build a safety-aware few-shot prompt.

The prompt combines the task definition, child-safety objective, story-to-real-life guidance, retrieved knowledge, and up to two selected demonstrations. Examples first match the knowledge label and are then selected or supplemented by story similarity under the model’s token limit.

IMPLEMENTED MODEL DETAILS

The ranker and prompt are inspectable parts of the system.

These thesis figures expose the two mechanisms behind the result: cross-story knowledge retrieval and knowledge-conditioned few-shot generation.

Neural knowledge retrieval architecture
Retrieval model. Sentence-BERT preprocessing, dual-tower representation learning, gradient-reversal regularization, and pairwise BPR ranking.
GPT-3 few-shot prompt construction
Generation prompt. Fixed safety instructions, retrieved evidence, selected story-knowledge-QA demonstrations, and the target story input.
04 / ALGORITHM VALIDATION

Test retrieval and generation as two distinct technical claims.

Knowledge retrieval

Leave-one-out cross-validation tested whether the ranker generalized to a held-out story. It outperformed TF-IDF and Sentence-BERT cosine similarity across Hit@K, NDCG, precision, and recall. Removing adversarial learning reduced every reported metric, providing an ablation check on that design choice.

Model Hit@1 Hit@6 NDCG Precision Recall
TF-IDF 0.000 0.259 0.124 0.043 0.046
Sentence-BERT + cosine 0.500 0.667 0.565 0.178 0.221
NCF, no adversarial loss 0.467 0.800 0.642 0.339 0.412
NCF + adversarial loss 0.533 0.833 0.670 0.372 0.466

Question-answer generation

Across all 157 annotated pairs, leave-one-out evaluation compared the knowledge-grounded few-shot configuration with a zero-shot GPT-3 baseline. Few-shot prompting improved both lexical overlap and semantic similarity without updating GPT-3 parameters.

ZERO-SHOT BASELINE0.617ROUGE-L F10.860BERTScore F1
→
GROUNDED FEW-SHOT0.811ROUGE-L F10.957BERTScore F1
05 / PARENT VALIDATION

Evaluate the complete outputs as educational material, not only as NLP scores.

The same seven parents reviewed five story cases. Each case contained three QA pairs from the proposed pipeline and three from a zero-shot baseline. Questionnaires measured preference, technical quality, and child-safety value; follow-up interviews examined usefulness, knowledge matching, comparison with existing education, adoption intent, and future interaction needs.

6 / 7Parents preferred the proposed outputs overall
> 4 / 5Average ratings across reported technical dimensions
> 4 / 5Average ratings across child-safety dimensions
72%Expressed interest in future use
Parent ratings for safety knowledge expertise, comprehensiveness, real-life relevance, and universality
The proposed pipeline was rated above the baseline for safety-knowledge expertise, comprehensiveness, real-life relevance, and applicability across situations.

What parents added beyond the scores

Parents saw story-grounded QA as a useful supplement for sparking interest and reinforcing safety awareness in daily life. They did not consider a single generated exchange a substitute for systematic professional education.

They also wanted richer back-and-forth interaction, child-only and parent-child modes, and more visual or embodied presentation. These findings define product directions rather than capabilities already implemented in the thesis prototype.

06 / LIMITATIONS & NEXT STEPS

Promising evidence, with deliberately narrow claims.

01Broaden stakeholders

The exploratory sample included seven parents recruited through network and snowball sampling. Future studies should include more families, children, educators, and child-safety professionals.

02Strengthen the knowledge base

Expand and professionally validate the 141-entry corpus and the 157-pair dataset before relying on generated material in higher-stakes settings.

03Move beyond one turn

Develop guided multi-turn interaction, including a child-facing mode and a co-reading mode in which caregivers select or adapt suggested prompts.

04Preserve adult oversight

Position the agent as a contextual supplement, expose retrieved evidence, and give caregivers control rather than presenting generated QA as professional safety instruction.

07 / MY SCOPE & COLLABORATION

Led the thesis across formative research, data, ML, evaluation, and prototyping.

  • Conducted the formative study and translated caregiver findings into the storytelling-agent concept and system requirements.
  • Built the 141-entry structured safety corpus and annotated 157 story-knowledge-QA samples derived from FairytaleQA.
  • Designed and implemented the Sentence-BERT, dual-tower MLP, adversarial-learning, and BPR retrieval pipeline in PyTorch.
  • Designed the GPT-3 prompt, knowledge-aware demonstration selection, and grounded few-shot QA generation workflow.
  • Ran offline retrieval and generation experiments, ablation testing, parent questionnaires, interviews, and mixed-method analysis.
  • Implemented the supporting web prototype and authored the bachelor thesis.
1Student project lead
1Faculty supervisor
7Parent participants
PROJECT LEADYing Lei

Research design, interviews, corpus and dataset construction, model and prototype implementation, experiments, analysis, and thesis writing.

FACULTY SUPERVISORYuling Sun

Provided domain and research-method guidance, critical feedback, and academic supervision throughout the thesis.

NEXT CASE STUDY / 05Decision Intelligence →