LIVE SYSTEM / INDEPENDENT BUILDAGENTIC RAG · VECTOR SEARCH · FUNCTION CALLING

Portfolio
Agentic RAG

A live technical concierge that retrieves verified evidence before it answers—and says when the portfolio cannot support a claim.

Ask about a project, compare systems, or assess a job fit.
RoleIndependent AI Software Engineer
RuntimeGoogle Cloud Run
RetrievalHybrid lexical + vector
StatusLive in this portfolio
01 / THE PROBLEM

A recruiter needs evidence—not a general-purpose chatbot.

The agent must answer exact facts, compare systems, and map a job description to demonstrated experience. It must also recognize when a request—such as solving an unrelated coding problem—does not belong to the portfolio at all. Fluency is secondary to scope, traceability, predictable cost, and honest claim boundaries.

AnswerVerified portfolio Q&A

Ground project and experience claims in source-linked evidence.

AssessRole-fit mapping

Map requirements to proof and expose unsupported gaps.

DeclineOff-domain work

Reject general coding and assistant tasks before paid services run.

02 / SCOPE & ROUTING

Decide whether the request deserves an AI call.

A local preflight gate runs before embeddings, vector search, or the Responses API. It protects the recruiter experience and the API budget without pretending that retrieval planning is a domain classifier.

Usage policy: blocked requests still count toward the IP and conversational limits, but they never trigger retrieval or generation.

03 / KNOWLEDGE

Turn project pages into evidence the agent can cite.

Each system is decomposed into section-level chunks for its problem, architecture, algorithms, validation, contribution, and limitations. Every chunk keeps its project ID, source URL, keywords, and corpus version.

01Authoritative pages

Reviewed project narratives and verified recruiter facts.

→
02Evidence chunks

Small, section-specific claims with source metadata.

→
03Dual indexes

Local lexical corpus + Firestore vector embeddings.

→
04Clickable proof

The UI exposes the page supporting each answer.

Versioned corpusSection metadataSource URLsVerified FAQs
04 / AGENT ORCHESTRATION

One bounded loop connects planning, tools, and evidence.

For an in-scope complex request, deterministic planning first constrains the relevant projects. The model then requests evidence through a strict, read-only function. Backend code validates the arguments, executes retrieval, and returns structured tool output. A second and final model call synthesizes the answer with tool use disabled.

Why bounded?Two model calls at most · one retrieval stage · no open-ended tool loop
LEXICALExact terms

Project names, frameworks, metrics, and named technologies.

+
VECTORSemantic similarity

Query embedding → Firestore cosine KNN over verified chunks.

→
FUSIONRank + coverage

Fuse rankings and preserve planned project coverage.

05 / TWO WORKFLOWS

The same evidence engine serves two recruiter tasks.

PORTFOLIO Q&AQuestion → project scope → evidence → answer

Handles technical details, ownership, comparisons, validation, and limitations without searching irrelevant projects.

JOB-FIT ASSESSMENTJD → requirements → evidence map → gaps

Maps requirements to demonstrated proof, recommends the relevant résumé, and labels unsupported skills instead of inferring them.

06 / PRODUCTION CONTROLS

Failures, abuse, and cost are part of reliability.

Cost & abuseLayered limits

Local scope gate, IP request window, conversational cap, daily budget, and organization hard limit.

Retrieval failureVector → lexical

Embedding or Firestore failure degrades to the local lexical corpus.

Generation failureDeterministic fallback

Job fit can still return supported evidence and explicit gaps.

Persistence failureLocal continuity

Session counts fall back to memory; operational events fall back to JSONL.

Least privilegeRead-only AI tool

The model can search evidence; it cannot mutate portfolio data.

PrivacyMetadata, not messages

Operational logs retain workflow and evidence IDs, not recruiter content.

07 / EVALUATION

Regression tests cover routes—not only final prose.

The suite tests FAQ routing, off-domain code rejection, workflow detection, retrieval scope, cross-project coverage, source exposure, résumé selection, unsupported claims, and dependency fallbacks. The next evaluation layer measures retrieval Recall@K, citation correctness, refusal precision, unsupported-claim rate, latency, and estimated cost.

Scope accuracyRetrieval relevanceCitation correctnessGrounded task successCost per answer
BACK TO CASE STUDY / 01WatchGuardian →