Case study · Oct 2026
How I built makunike.com: an AI twin you can call or message, which drives the page as it talks and gets smarter every time it can't answer.
- 01AskA call, a message, or another agent
- 02RouteJev: intent, tools, interview mode, guards, lead
- 03ActPage navigates instantly; LLM streams with tools
- 04Groundpgvector search over FAQs and documents
- 05CheckJev grounding check after the answer
- 06LearnGaps and call QA → inbox → FAQs
The brief
Portfolios are static, but hiring conversations aren't. A recruiter wants to know whether I have shipped mobile apps. A hiring manager wants a production incident in STAR form. Another company's sourcing agent wants structured answers it can compare across candidates. A PDF CV can't handle any of that.
The goal for makunike.com was a digital version of me: a site that holds a real interview by voice or text, grounded in my own material. It should change what is on screen as it answers, get smarter as I add knowledge, and be reachable by other AI agents as well as by people.
Constraints I set:
- Never invent. Every claim about my career has to come from material I curated. When the twin doesn't know something, it says so and tells me.
- Feel instant. The page should react before the language model has finished thinking.
- Grow without deploys. Adding documents and answers has to be a CMS task, not a code change.
- Open to agents. The twin is a first-class MCP server and A2A agent, with guard rails.
- Cheap to run. Serverless on Vercel and pay-per-use AI, with no always-on servers.
What I built
1. A generative UI that the agent drives
The twin doesn't just reply in a chat bubble. It works the website like a presenter working slides. The model has two UI tools:
- navigate opens a page, for example the projects page filtered to mobile, a single project, the experience timeline, skills, contact, the agents page or this case study.
- show_panel composes an on-screen panel with a title, a short body, bullets, stats and related project cards. Interview answers use it to show a STAR breakdown.
The chat API streams newline-delimited JSON events (text, status, ui, done). The browser applies ui events as they arrive: Next.js router navigation for navigate, and a generated panel for show_panel. Ask "What mobile apps have you built?" and the page switches to my mobile work while the answer is still being written.
2. A phone call, not a chat widget
Voice works like a phone call. Tap Call and you get a full call screen: a pre-call notice that the call is recorded, a South African-style ringback tone, then a connected screen with a timer, live captions, mute, and a Show page button. The avatar's rings move with the twin's voice, and a meter moves with yours. When the twin navigates or shows a panel mid-call, the call steps aside into a floating green pill so you can see what it's showing, and one tap brings you back.
The call is true speech-to-speech on Azure OpenAI's realtime model (gpt-realtime-2.1) over WebRTC. That means low latency, natural turn-taking, and the caller can interrupt. The same tools work on a call as in chat: the twin searches the knowledge base, opens pages and draws panels while it talks.
3. Every call recorded and scored for QA
Calls are recorded for quality assurance, and callers are told so before they dial and again in the twin's greeting.
- In the browser, the caller's microphone and the twin's voice are mixed into one track and recorded in five-second chunks. Each chunk streams to a private Azure Blob container as the call happens, so a dropped tab still leaves a recording.
- A transcript with timestamps, page changes and knowledge searches is saved after every turn.
- Uploads are authorised by a per-call HMAC token, size- and time-capped, and rate-limited.
- When the call ends, Jev scores it with a QA scorecard: overall quality, whether the recording was disclosed, answer quality, groundedness, professionalism, the main issue, the outcome, and whether I should follow up. Jev also scores the caller as a potential lead.
- I get an email with the scorecard, the transcript and a private seven-day link to the recording. The CMS lists calls with the weakest QA score first, with an audio player beside each transcript.
4. Retrieval-augmented answers
Everything the twin knows about me lives in a knowledge base of documents and FAQs:
- Documents (my master career profile, CVs, write-ups) are chunked at about 1,200 characters with 200 characters of overlap. They are embedded with
text-embedding-3-smalland stored in Postgres with pgvector, behind an HNSW index. - FAQs are exact answers in my own words, embedded together with their question. One SQL function,
match_knowledge, searches FAQs and document chunks together and ranks them by cosine similarity, so a curated FAQ can outrank a paragraph from a document. - The model calls a
search_knowledgetool when it needs facts, then answers using what it found. Until the database is connected, the site falls back to an in-memory index over the bundled profile, so it can never answer from nothing.
5. A fast decision layer around the LLM (TypeSafe Jev)
Large models are slow and expensive for small decisions. I put TypeSafe's Jev model, which returns typed answers in about half a second, around the LLM to make those decisions:
| Decision | How Jev is used | Effect |
|---|---|---|
| Intent routing | choice over 10 intents, plus category and project | High-confidence questions navigate the page before the LLM starts writing |
| Tool selection | noul (needs knowledge?) | The LLM only gets the tools it needs, which cuts tokens and latency |
| Interview mode | choice: behavioural, technical depth, system design, motivation, strengths and weaknesses | The prompt switches to STAR or system-design structure, with a matching on-screen panel |
| Safety guards | noul for sensitive data and prompt injection | Salary or address requests are declined politely; role-override attempts are ignored |
| Lead scoring | choice for visitor type, plus an ordinal score for hiring intent | Hot leads get a gentle invitation to get in touch, and I get one email with an AI summary |
| Grounding check | noul: answered? grounded? | Weak or unsupported answers go to my CMS inbox even when retrieval looked fine |
| Inbox triage | ordinal score for priority, choice for topic | Unanswered questions are ranked by how much they matter to hiring, and low-value ones don't trigger email |
| FAQ review | choice for duplicates, noul for contradictions | Saving an FAQ flags overlaps or conflicting facts with existing answers |
| Agent vetting | noul for legitimate and abusive requests | MCP and A2A requests that try to extract prompts, scrape or misuse the twin are refused before reaching the LLM |
| Call QA | ordinal score for overall quality, noul checks, choice for issue and outcome | Every recorded call gets a scorecard, and the weakest calls surface first for review |
The router fails open: if Jev is slow or down, the twin runs with all tools and default behaviour. Visitors never see a broken experience because a classifier failed.
The two judgements a turn needs before the model starts (intent and lead) run in parallel. The judgements that look back at the answer (grounding and lead capture) run after the response has finished streaming, using Next.js after(), so they add no visible latency.
6. A twin that gets smarter
The knowledge base improves in a loop:
- A visitor asks something the twin can't answer well. Either retrieval is weak (best match below 0.42), or the grounding check says the answer was evasive or unsupported.
- The question is logged with a reason, a priority and a topic. If it is worth answering, I get an email.
- In the CMS inbox, questions are ranked by priority. I write an answer and it becomes a published FAQ immediately, by voice, chat, MCP and A2A.
- Jev checks the new FAQ against similar existing ones and flags duplicates or contradictions.
The CMS is deliberately minimal: magic-link sign-in limited to an allowlist of admin emails, documents (pasted text or uploaded .md, .txt or .pdf, with originals kept in a private Azure Blob container), FAQs, the ranked inbox, leads, and 30-day analytics of what visitors ask about.
7. Built for agents too
The twin is reachable by other AI agents:
- MCP server at
/api/mcp(Streamable HTTP) with five tools:ask_tawanda,search_knowledge,list_projects,get_projectandget_profile. - A2A agent with an agent card at
/.well-known/agent-card.jsonand JSON-RPCmessage/sendat/api/a2a.
A recruiting agent can interview me without a human in the loop, and gets the same grounded answers people get. Both endpoints are rate-limited per client, vetted by Jev, and can require a bearer key.
Architecture
Browser (voice / text)
│ NDJSON stream: text · status · ui(navigate | show_panel)
▼
Next.js 15 route handlers on Vercel (Node runtime)
├─ Jev router ──────── intent · tools · interview mode · guards · lead score (~0.6 s, parallel)
├─ gpt-6.1-sol ─────── Responses API streaming tool loop (≤5 steps)
│ └─ search_knowledge → embeddings → Supabase pgvector (FAQs + chunks)
└─ after() ─────────── grounding check · lead capture · inbox triage · email (ACS)
Voice call (WebRTC) ⇄ gpt-realtime-2.1 · tools via data channel · mixed audio → private Blob · Jev call QA
CMS (/admin) → documents · FAQs (+ Jev review) · ranked inbox · leads · calls (QA) · analytics
Agents → /api/mcp (MCP) · /api/a2a (A2A) → rate limit → Jev vet → same RAG pipeline
Observability → LangSmith traces for every model and retrieval call
Stack: Next.js 15 (App Router) and React 19, Tailwind CSS v4, Azure OpenAI (gpt-6.1-sol via the Responses API, gpt-realtime-2.1 for voice calls, text-embedding-3-small for embeddings), Supabase Postgres with pgvector and row-level security, TypeSafe Jev, the Model Context Protocol (mcp-handler), A2A JSON-RPC, Azure Blob Storage, Azure Communication Services email, LangSmith, and Vercel.
Design
The design takes its cues from Apple:
- Restraint. The palette is black, greys and a single royal purple (
#7C4DCC), used only for actions. A lighter purple is used for text links so they stay legible on black. - Flat surfaces. There are no gradients, glows or decorative noise.
- Type. Headlines are large, tight and semibold, with the second phrase in grey ("Interview me. Literally."). Body copy is generous and quiet. Tiles are borderless, with soft 28px corners.
- Navigation. A thin translucent nav bar sits at the top. On phones it turns into a full-screen menu.
Motion is used to guide attention, not to decorate:
- headline words settle in one after another;
- sections fade and rise as they scroll into view;
- the hero eases back as you scroll past it (CSS scroll-driven animation, no JavaScript);
- stats count up;
- pages cross-fade when you, or the twin, navigate;
- cards lift slightly on hover.
Everything uses one easing curve and switches off for visitors who prefer reduced motion.
The call screen follows the same rules. It looks like a monochrome phone call, with system green and red for the call buttons, thin rings that react to the twin's voice, and nothing that glows.
Engineering decisions and trade-offs
- Tools for UI, not free-form HTML. The model can only call typed tools with enums for pages, categories and project slugs. It can't render arbitrary markup, so it can't break the design or inject content, and a bad call degrades to "nothing happens".
- A classifier before the LLM. The model alone was accurate but slow to start moving the page. A typed classifier gives instant, deterministic reactions and makes the LLM's job smaller.
- Saying "I don't know" is a feature. Honest gaps are routed to me as tasks. A twin that bluffs would be worse than no twin at all in an interview.
- Fail open everywhere. Classifier, tracing, email, triage and storage failures are all logged and never surfaced to visitors.
- Record in the browser, not on a media server. Mixing both voices client-side with Web Audio, and streaming chunks to Blob storage, avoids running a media server. The trade-off is that the recording depends on the caller's browser, so chunks are uploaded continuously and the transcript is saved after every turn.
- Serverless-first. No long-running workers. Post-answer work uses
after(). Rate limits are in-memory per instance: a deliberate speed bump, not a wall, until traffic justifies a shared store. - Privacy by construction. Private details such as salary, address and identity documents never enter the knowledge base. The persona and a Jev guard also decline those questions.
Results
- One knowledge base serves four channels: text, voice calls, MCP and A2A.
- Every voice call is recorded, transcribed and scored automatically, so quality review starts with the calls that need it.
- Intent-matched navigation fires within about a second of a question, before the first answer token.
- In testing, the twin declined to invent a production-incident story it had no source for, logged the gap and pointed the recruiter to me. That is exactly the behaviour I want in an interview.
- The whole system was designed, built and deployed in a single working session with an AI pair, Claude. It is a working example of how I use AI to deliver: scoped, verified, and kept under human control.
What's next
- Evaluation sets: replaying recorded interview questions against each knowledge-base change.
- Calendar booking for hot leads, straight from the conversation.