What if you could clone yourself for meetings? Not a creepy deepfake, but a transparent AI copilot that represents you when you can't be there—and seamlessly hands off to the real you when you can.
You have three meetings scheduled at the same time. One is a status update you don't really need to attend but should probably monitor. Another is a sales call where your presence matters but most of the talking is done by your team. The third actually requires your brain.
What if instead you had an AI avatar that could attend meetings on your behalf—with actual video presence—and knew when to summon specialist capabilities for specific tasks?
The core model isn't "a team of bots in the meeting." It's one intelligent avatar that knows who to call.
You invite a single bot to the meeting. It joins with video presence—an actual avatar in the meeting grid, not a chat sidebar. Participants see a face, hear a voice, and interact naturally. Under the hood, this avatar is backed by a coordinator called Maestro that understands context and can summon specialist capabilities on demand.
Someone asks about scheduling → Maestro summons Tempo, who has access to everyone's calendars and can propose times in real-time.
The conversation shifts to social media strategy → Maestro summons The Algorithm, who can pull live engagement data and draft tweets on the fly.
A question comes up about recent emails → Gatekeeper gets summoned, surfaces relevant threads, and hands back.
Someone asks "what's the market rate for X?" → Radar does real-time web research and reports back.
The more ambitious vision: an avatar that represents you specifically.
Here's what sets this apart from a generic bot: you can take over at any moment.
You speak, the avatar speaks your words in real-time
You type, the avatar vocalizes what you typed
And crucially: transparency built in. The system makes clear when it's pulling from your knowledge base versus when it's actually you. Maybe a subtle visual indicator. Maybe the avatar explicitly says "David mentioned in his notes that..." vs. "David says..."
This isn't about tricking people. It's about extending your presence honestly.
Three things have converged to make this feasible:
OpenAI's Realtime API delivers speech-to-text, LLM reasoning, and text-to-speech in a single streaming pipeline. Response times are now 300-600ms—conversational speed. A year ago this required stitching together three separate systems with compounding latency.
Services like Recall.ai and LiveKit have solved the hard problems of getting AI into video calls. You can join Google Meet, Zoom, or Teams programmatically, capture audio, display video, and stream responses back. The plumbing works.
LiveKit's agent dispatch lets you spin up a fresh agent instance by name into an existing room. Combined with Redis for coordination state, you get clean handoffs: the current agent finishes speaking, signals exit, and the new specialist takes over within seconds. No multi-agent chaos—just sequential, controlled transitions.

Here's how the single-avatar, multi-specialist model works:
This is the "summoning" pattern—Maestro summons Tempo for a calendar question the same way you'd turn to a colleague and say "hey, you handle this one." The meeting participants see one continuous avatar presence. The architecture underneath is swapping agents.
Let's be honest about current state versus aspirations.
Here's the thing architects need to understand: cloud avatar rendering is slow.
Current end-to-end latency breakdown:

100-300ms
200-500ms
100-300ms
500-2000ms
That's not conversational. It's awkward.
The avatar rendering is the bottleneck. Every cloud avatar service we tested (HeyGen, Hedra) has this problem. The video frames have to be generated on their servers and streamed back. Physics and network latency impose a floor.
The path forward is probably client-side avatar rendering using WebGL and something like Ready Player Me. That could get the total latency under 500ms. But it's a significant engineering lift.
For now, there's also an audio-only mode that skips avatar rendering entirely. Total latency drops to 300-800ms—genuinely conversational. The trade-off is no video presence, just a static image with an audio visualizer.
This series will walk through the architecture in detail:
Deep dive into getting AI into video calls. Recall.ai integration, LiveKit room management, OpenAI Realtime API, and avatar rendering. We'll trace a single utterance from human speech to AI response and back.
How the summoning pattern works in code. LiveKit agent dispatch, Redis coordination, the Maestro-specialist handoff lifecycle, and why one-at-a-time beats a crowd.
Building the personal knowledge base, the takeover UX, and how to signal transparency between AI and human speaking.
One avatar seat with specialist capabilities behind it, plus a personal avatar that represents you
Agents swap in and out via dispatch, not multiple bots fighting for airtime
Clear signaling when AI is speaking versus when you are
Recall.ai, LiveKit, OpenAI Realtime make the core pipeline possible today
Cloud avatar rendering adds 0.5-2 seconds, pushing total latency to 1-3+ seconds
The same tool infrastructure that powers chat agents powers the meeting avatar
This is Part 1 of a 4-part series on building a meeting copilot. Part 2: The Audio-Video Pipeline →
Next up: the technical deep dive into getting an AI's voice and face into a Google Meet call.
Building a Meeting Copilot: The Vision