Live · Oct 2026 · Designer & Engineer

This website: a portfolio you can interview by phone call or chat

Makunike.com AI Twin home page on desktop

Overview

A portfolio that answers back. A RAG-grounded AI twin speaks for me by voice or text and drives the site as it talks. A TypeSafe Jev layer makes sub-second decisions around the LLM, a minimal CMS grows its knowledge, and MCP and A2A servers let other agents interview it.

The problem

Portfolios are static, but hiring conversations aren't. A recruiter wants to know whether I've shipped mobile apps, a hiring manager wants a production story in STAR form, and a sourcing agent wants structured answers it can compare. A PDF CV can't do any of that.

The solution

A digital version of me that holds a real interview by phone-style voice call or text, grounded only in material I curated. It changes what's on screen as it answers, logs what it can't answer so I can teach it, records and QA-scores every call, and is reachable by other AI agents over MCP and A2A.

Makunike.com AI Twin home page on a phone

What I built

  • Generative UI: the model calls typed navigate and show_panel tools, streamed to the browser as NDJSON events
  • Phone-call voice on gpt-realtime-2.1 over WebRTC, with every call recorded, transcribed and QA-scored
  • TypeSafe Jev in ten places: routing, tool selection, interview mode, guards, lead scoring, grounding, triage, FAQ review, agent vetting and call QA
  • Self-improving RAG: weak or ungrounded answers land in a ranked CMS inbox and become FAQs
  • MCP and A2A servers with request vetting and rate limits for agent-to-agent interviews

Technical challenges

What got in the way. And how we solved it.

  1. The newest model rejected tools on Chat Completions

    Challenge

    Moving chat to gpt-6.1-sol broke the tool loop: GPT-6.x doesn't accept function tools on the Chat Completions API.

    How we solved it

    Rebuilt the chat loop on the Responses API: streamed output, previous_response_id to chain turns and function_call_output to return tool results, with low reasoning effort to keep it fast.

  2. Moving the page before the model has thought

    Challenge

    The LLM was accurate but took too long to decide to change the page, so the UI felt laggy.

    How we solved it

    Put a typed Jev classifier in front of it. High-confidence intents navigate instantly, in parallel with lead scoring, while the LLM streams with only the tools it needs.

  3. Recording calls without a media server

    Challenge

    Voice runs peer-to-peer over WebRTC, so there's no server in the audio path to record from, and serverless functions can't hold a long-lived stream.

    How we solved it

    Mix the caller's microphone and the twin's voice in the browser with Web Audio, record five-second chunks and upload each one to an Azure append blob, authorised by a per-call HMAC token. Transcripts save after every turn, and a beacon closes the call if the tab is closed.

  4. Letting a model change the UI safely

    Challenge

    Free-form generated HTML could break the design or inject content.

    How we solved it

    The model can only call typed tools with enums for pages, categories and project slugs. A bad call degrades to nothing happening.

Solved ahead of time

Designed in, not bolted on.

Fail open everywhere

Classifier, tracing, email, triage and storage failures are logged and never shown to visitors. If Jev is down, the twin runs with all tools and default behaviour.

Privacy by construction

Salary, address and identity documents never enter the knowledge base, and the persona and a Jev guard decline those questions.

Abuse from other agents

MCP and A2A requests are rate-limited per client and vetted by Jev before they reach the LLM, and can require a bearer key.

Recording consent

Callers are told the call is recorded before they dial and again in the twin's greeting.

Never answering from nothing

Without a database, the site falls back to an in-memory index over the bundled profile.

Infrastructure

The choices. And why.

Vercel serverless (Next.js route handlers)
No always-on servers. Post-answer work runs in after(), so it adds no visible latency.
Azure OpenAI
One provider for chat (Responses API), realtime voice and embeddings.
Supabase Postgres + pgvector (HNSW)
FAQs and document chunks are searched in one SQL function, so a curated FAQ can outrank a document paragraph. Row-level security keeps every table server-only.
TypeSafe Jev
Typed decisions in about half a second, cheaper and faster than asking the LLM.
Azure Blob Storage
Private, append-friendly storage for call recordings and CMS uploads, shared via short-lived SAS links.
In-memory rate limits
A deliberate speed bump, not a wall, until traffic justifies a shared store.

What success looks like

The twin never invents a fact about my career: when it doesn't know, it says so and I'm notified. Intent-matched navigation fires within about a second, before the first answer token. Every call is recorded, transcribed and scored so review starts with the calls that need it, and one knowledge base serves text, voice, MCP and A2A.

How we improve it

Three feedback loops drive changes. Unanswered or ungrounded questions are triaged by priority into the CMS inbox and turned into FAQs. Every call gets a Jev QA scorecard, weakest first. LangSmith traces every model and retrieval call. Next up: evaluation sets that replay recorded interview questions against each knowledge-base change.

Ask the twin