Case study · Oct 2026

How I built makunike.com: an AI twin you can call or message, which drives the page as it talks and gets smarter every time it can't answer.

4
channels, one brain: text, voice calls, MCP, A2A
~0.6s
typed Jev decisions before the LLM writes
10
Jev use cases around the model
0
invented facts allowed
  1. 01
    Ask
    A call, a message, or another agent
  2. 02
    Route
    Jev: intent, tools, interview mode, guards, lead
  3. 03
    Act
    Page navigates instantly; LLM streams with tools
  4. 04
    Ground
    pgvector search over FAQs and documents
  5. 05
    Check
    Jev grounding check after the answer
  6. 06
    Learn
    Gaps and call QA → inbox → FAQs

The brief

Portfolios are static, but hiring conversations aren't. A recruiter wants to know whether I have shipped mobile apps. A hiring manager wants a production incident in STAR form. Another company's sourcing agent wants structured answers it can compare across candidates. A PDF CV can't handle any of that.

The goal for makunike.com was a digital version of me: a site that holds a real interview by voice or text, grounded in my own material. It should change what is on screen as it answers, get smarter as I add knowledge, and be reachable by other AI agents as well as by people.

Constraints I set:

  • Never invent. Every claim about my career has to come from material I curated. When the twin doesn't know something, it says so and tells me.
  • Feel instant. The page should react before the language model has finished thinking.
  • Grow without deploys. Adding documents and answers has to be a CMS task, not a code change.
  • Open to agents. The twin is a first-class MCP server and A2A agent, with guard rails.
  • Cheap to run. Serverless on Vercel and pay-per-use AI, with no always-on servers.

What I built

1. A generative UI that the agent drives

The twin doesn't just reply in a chat bubble. It works the website like a presenter working slides. The model has two UI tools:

  • navigate opens a page, for example the projects page filtered to mobile, a single project, the experience timeline, skills, contact, the agents page or this case study.
  • show_panel composes an on-screen panel with a title, a short body, bullets, stats and related project cards. Interview answers use it to show a STAR breakdown.

The chat API streams newline-delimited JSON events (text, status, ui, done). The browser applies ui events as they arrive: Next.js router navigation for navigate, and a generated panel for show_panel. Ask "What mobile apps have you built?" and the page switches to my mobile work while the answer is still being written.

2. A phone call, not a chat widget

Voice works like a phone call. Tap Call and you get a full call screen: a pre-call notice that the call is recorded, a South African-style ringback tone, then a connected screen with a timer, live captions, mute, and a Show page button. The avatar's rings move with the twin's voice, and a meter moves with yours. When the twin navigates or shows a panel mid-call, the call steps aside into a floating green pill so you can see what it's showing, and one tap brings you back.

The call is true speech-to-speech on Azure OpenAI's realtime model (gpt-realtime-2.1) over WebRTC. That means low latency, natural turn-taking, and the caller can interrupt. The same tools work on a call as in chat: the twin searches the knowledge base, opens pages and draws panels while it talks.

3. Every call recorded and scored for QA

Calls are recorded for quality assurance, and callers are told so before they dial and again in the twin's greeting.

  • In the browser, the caller's microphone and the twin's voice are mixed into one track and recorded in five-second chunks. Each chunk streams to a private Azure Blob container as the call happens, so a dropped tab still leaves a recording.
  • A transcript with timestamps, page changes and knowledge searches is saved after every turn.
  • Uploads are authorised by a per-call HMAC token, size- and time-capped, and rate-limited.
  • When the call ends, Jev scores it with a QA scorecard: overall quality, whether the recording was disclosed, answer quality, groundedness, professionalism, the main issue, the outcome, and whether I should follow up. Jev also scores the caller as a potential lead.
  • I get an email with the scorecard, the transcript and a private seven-day link to the recording. The CMS lists calls with the weakest QA score first, with an audio player beside each transcript.

4. Retrieval-augmented answers

Everything the twin knows about me lives in a knowledge base of documents and FAQs:

  • Documents (my master career profile, CVs, write-ups) are chunked at about 1,200 characters with 200 characters of overlap. They are embedded with text-embedding-3-small and stored in Postgres with pgvector, behind an HNSW index.
  • FAQs are exact answers in my own words, embedded together with their question. One SQL function, match_knowledge, searches FAQs and document chunks together and ranks them by cosine similarity, so a curated FAQ can outrank a paragraph from a document.
  • The model calls a search_knowledge tool when it needs facts, then answers using what it found. Until the database is connected, the site falls back to an in-memory index over the bundled profile, so it can never answer from nothing.

5. A fast decision layer around the LLM (TypeSafe Jev)

Large models are slow and expensive for small decisions. I put TypeSafe's Jev model, which returns typed answers in about half a second, around the LLM to make those decisions:

DecisionHow Jev is usedEffect
Intent routingchoice over 10 intents, plus category and projectHigh-confidence questions navigate the page before the LLM starts writing
Tool selectionnoul (needs knowledge?)The LLM only gets the tools it needs, which cuts tokens and latency
Interview modechoice: behavioural, technical depth, system design, motivation, strengths and weaknessesThe prompt switches to STAR or system-design structure, with a matching on-screen panel
Safety guardsnoul for sensitive data and prompt injectionSalary or address requests are declined politely; role-override attempts are ignored
Lead scoringchoice for visitor type, plus an ordinal score for hiring intentHot leads get a gentle invitation to get in touch, and I get one email with an AI summary
Grounding checknoul: answered? grounded?Weak or unsupported answers go to my CMS inbox even when retrieval looked fine
Inbox triageordinal score for priority, choice for topicUnanswered questions are ranked by how much they matter to hiring, and low-value ones don't trigger email
FAQ reviewchoice for duplicates, noul for contradictionsSaving an FAQ flags overlaps or conflicting facts with existing answers
Agent vettingnoul for legitimate and abusive requestsMCP and A2A requests that try to extract prompts, scrape or misuse the twin are refused before reaching the LLM
Call QAordinal score for overall quality, noul checks, choice for issue and outcomeEvery recorded call gets a scorecard, and the weakest calls surface first for review

The router fails open: if Jev is slow or down, the twin runs with all tools and default behaviour. Visitors never see a broken experience because a classifier failed.

The two judgements a turn needs before the model starts (intent and lead) run in parallel. The judgements that look back at the answer (grounding and lead capture) run after the response has finished streaming, using Next.js after(), so they add no visible latency.

6. A twin that gets smarter

The knowledge base improves in a loop:

  1. A visitor asks something the twin can't answer well. Either retrieval is weak (best match below 0.42), or the grounding check says the answer was evasive or unsupported.
  2. The question is logged with a reason, a priority and a topic. If it is worth answering, I get an email.
  3. In the CMS inbox, questions are ranked by priority. I write an answer and it becomes a published FAQ immediately, by voice, chat, MCP and A2A.
  4. Jev checks the new FAQ against similar existing ones and flags duplicates or contradictions.

The CMS is deliberately minimal: magic-link sign-in limited to an allowlist of admin emails, documents (pasted text or uploaded .md, .txt or .pdf, with originals kept in a private Azure Blob container), FAQs, the ranked inbox, leads, and 30-day analytics of what visitors ask about.

7. Built for agents too

The twin is reachable by other AI agents:

  • MCP server at /api/mcp (Streamable HTTP) with five tools: ask_tawanda, search_knowledge, list_projects, get_project and get_profile.
  • A2A agent with an agent card at /.well-known/agent-card.json and JSON-RPC message/send at /api/a2a.

A recruiting agent can interview me without a human in the loop, and gets the same grounded answers people get. Both endpoints are rate-limited per client, vetted by Jev, and can require a bearer key.

Architecture

Browser (voice / text)
   │  NDJSON stream: text · status · ui(navigate | show_panel)
   ▼
Next.js 15 route handlers on Vercel (Node runtime)
   ├─ Jev router ──────── intent · tools · interview mode · guards · lead score   (~0.6 s, parallel)
   ├─ gpt-6.1-sol ─────── Responses API streaming tool loop (≤5 steps)
   │     └─ search_knowledge → embeddings → Supabase pgvector (FAQs + chunks)
   └─ after() ─────────── grounding check · lead capture · inbox triage · email (ACS)

Voice call (WebRTC) ⇄ gpt-realtime-2.1 · tools via data channel · mixed audio → private Blob · Jev call QA
CMS (/admin) → documents · FAQs (+ Jev review) · ranked inbox · leads · calls (QA) · analytics
Agents → /api/mcp (MCP) · /api/a2a (A2A) → rate limit → Jev vet → same RAG pipeline
Observability → LangSmith traces for every model and retrieval call

Stack: Next.js 15 (App Router) and React 19, Tailwind CSS v4, Azure OpenAI (gpt-6.1-sol via the Responses API, gpt-realtime-2.1 for voice calls, text-embedding-3-small for embeddings), Supabase Postgres with pgvector and row-level security, TypeSafe Jev, the Model Context Protocol (mcp-handler), A2A JSON-RPC, Azure Blob Storage, Azure Communication Services email, LangSmith, and Vercel.

Design

The design takes its cues from Apple:

  • Restraint. The palette is black, greys and a single royal purple (#7C4DCC), used only for actions. A lighter purple is used for text links so they stay legible on black.
  • Flat surfaces. There are no gradients, glows or decorative noise.
  • Type. Headlines are large, tight and semibold, with the second phrase in grey ("Interview me. Literally."). Body copy is generous and quiet. Tiles are borderless, with soft 28px corners.
  • Navigation. A thin translucent nav bar sits at the top. On phones it turns into a full-screen menu.

Motion is used to guide attention, not to decorate:

  • headline words settle in one after another;
  • sections fade and rise as they scroll into view;
  • the hero eases back as you scroll past it (CSS scroll-driven animation, no JavaScript);
  • stats count up;
  • pages cross-fade when you, or the twin, navigate;
  • cards lift slightly on hover.

Everything uses one easing curve and switches off for visitors who prefer reduced motion.

The call screen follows the same rules. It looks like a monochrome phone call, with system green and red for the call buttons, thin rings that react to the twin's voice, and nothing that glows.

Engineering decisions and trade-offs

  • Tools for UI, not free-form HTML. The model can only call typed tools with enums for pages, categories and project slugs. It can't render arbitrary markup, so it can't break the design or inject content, and a bad call degrades to "nothing happens".
  • A classifier before the LLM. The model alone was accurate but slow to start moving the page. A typed classifier gives instant, deterministic reactions and makes the LLM's job smaller.
  • Saying "I don't know" is a feature. Honest gaps are routed to me as tasks. A twin that bluffs would be worse than no twin at all in an interview.
  • Fail open everywhere. Classifier, tracing, email, triage and storage failures are all logged and never surfaced to visitors.
  • Record in the browser, not on a media server. Mixing both voices client-side with Web Audio, and streaming chunks to Blob storage, avoids running a media server. The trade-off is that the recording depends on the caller's browser, so chunks are uploaded continuously and the transcript is saved after every turn.
  • Serverless-first. No long-running workers. Post-answer work uses after(). Rate limits are in-memory per instance: a deliberate speed bump, not a wall, until traffic justifies a shared store.
  • Privacy by construction. Private details such as salary, address and identity documents never enter the knowledge base. The persona and a Jev guard also decline those questions.

Results

  • One knowledge base serves four channels: text, voice calls, MCP and A2A.
  • Every voice call is recorded, transcribed and scored automatically, so quality review starts with the calls that need it.
  • Intent-matched navigation fires within about a second of a question, before the first answer token.
  • In testing, the twin declined to invent a production-incident story it had no source for, logged the gap and pointed the recruiter to me. That is exactly the behaviour I want in an interview.
  • The whole system was designed, built and deployed in a single working session with an AI pair, Claude. It is a working example of how I use AI to deliver: scoped, verified, and kept under human control.

What's next

  • Evaluation sets: replaying recorded interview questions against each knowledge-base change.
  • Calendar booking for hot leads, straight from the conversation.