How The Double Is Built: A Complete Look at Our Tech Stack

Channels, agent orchestration, intent, retrieval, and Echo. A layer by layer look at how The Double works and where your knowledge lives.

Sahith KrishnaSahith Krishna15 min read
How The Double Is Built: A Complete Look at Our Tech Stack

This is how The Double is built, written so that anyone can follow it. We are publishing it because we think people have a right to understand what happens to their knowledge and their voice inside the system they are trusting with them.

Going from a rough prototype to a working product taught us more than we expected. We assumed the hard part would be the model. Attach an LLM, give it some documents, put a conversation in front of it, and you would have a Double. That turned out to be a small part of the work.

The biggest problem we had to solve was intent, meaning understanding what a person is actually trying to do when they send a message rather than assuming every message is a question. Alongside that, privacy, security, and guardrails became our highest priority rather than something to handle later, because a system that speaks in someone's voice using their knowledge has to be built for control from the start.

Two decisions shaped everything else. We build on open-source foundations rather than training base models from scratch, and for the parts that matter most, particularly voice, we developed our own model and self-host it. That gives us control over performance, over where processing happens, and over which outside companies are involved in handling user data at all.

Here is the whole thing, layer by layer.

The channel layer

The channel layer is where a Double is exposed to people, and the constraint we set early was that no channel would be a separate agent.

Every Double gets a page of its own that can be shared as a link. The same Double can sit on a website as a corner widget, a centre bar, an inline embed, or a card integration.

Where a conversation happens is separate from how it happens. Any of these places can be typed to or spoken to, and somebody can move between the two without the conversation restarting. Presenting a deck or a PDF works the same way, available wherever the Double already is rather than being a mode you have to switch into.

Behind all of them sits one configuration. The same knowledge, the same voice, the same persona and rules, the same enabled skills. Greeting messages, suggested replies, retrieval, and lead capture are set once rather than place by place.

That rule costs a little up front and pays for itself every time we add somewhere new. A new channel becomes another way into an existing Double rather than another agent to set up, which is what stops one Double behaving like three different ones depending on where somebody found it.

The agent layer

The agent layer sits between the person having the conversation and the models underneath. It decides what should happen once a message arrives.

Specifically, it handles:

  • Understanding the user's intent
  • Routing the request to the correct model or system
  • Searching the user's knowledge
  • Selecting and executing skills
  • Continuing multi-step workflows
  • Applying the owner's rules and permissions before taking any action
  • Returning the response through the surface the conversation started in

That coordination role is what we call agent orchestration, and it is the part that turns a set of separate capabilities into something that behaves coherently.

The intent engine

The intent engine works out what a person is trying to achieve. The difference this makes is easy to see with an example. If someone types "can you send me the deck," a system that treats everything as a question will describe the deck instead of sending it. The answer is relevant, and the person still does not have the file.

So before generating anything, the intent engine identifies what kind of request has arrived. That might be asking a question, retrieving knowledge, opening or presenting content, sharing a document, capturing a lead, sending an email, starting an interview, booking a meeting, or triggering a skill or workflow.

This is one of the main differences between a Double and a simple chatbot. The system does not treat every incoming message as a request for text generation.

We tune and run this ourselves for a practical reason. Intent runs on every single turn of every conversation, so it sits directly in the speed and cost path of the entire product. A narrow model doing one job well is a better fit for that than a general purpose one.

Skills and workflows

Skills are the individual actions a Double can perform. Presenting a document, conducting an interview, supporting a training experience, capturing a lead, sending an email, sharing content, sharing a file or link, and booking a meeting.

Workflows connect several actions into a controlled sequence. A typical one runs like this:

  1. Understand what a visitor needs
  2. Ask the relevant questions
  3. Capture the required information
  4. Confirm it back to them
  5. Save the lead
  6. Trigger a follow-up action

The action system is built around defined and controlled capabilities rather than open-ended autonomy. A Double does what its owner has enabled, and nothing else.

Knowledge and retrieval

The Double uses Retrieval-Augmented Generation, or RAG. Your knowledge is stored separately from the language model and brought into a conversation only when it is relevant to what is being asked.

Knowledge sources can include PDFs, documents, custom text, websites, images, YouTube videos, LinkedIn, X, and other social content.

When a question arrives, the system:

  1. Understands what is being asked
  2. Searches the user's knowledge
  3. Finds the most relevant information
  4. Adds that information to the model's context
  5. Generates an answer grounded in what was retrieved

Your knowledge is never built into the model itself. We do not retrain a model on your documents, and the model does not memorise them. That distinction has three practical consequences that matter to you. Add a source and your Double can use it immediately, with no retraining and no waiting. Update a source and the old version stops being used. Delete a source and your Double stops knowing it straight away. If your knowledge were trained into a model's weights instead, none of those three would be true, and there would be no reliable way to take something back once it had gone in.

How knowledge is processed

Adding a source is not the same as the Double being able to use it. In between sits an ingestion pipeline:

  1. The user adds a document, website, text, image, video, or social source
  2. The content is uploaded or fetched
  3. Its text or usable information is extracted
  4. The content is divided into smaller searchable sections
  5. Embeddings are generated for each section
  6. The processed knowledge is added to vector memory
  7. Those sections can then be retrieved during conversations

Embeddings are numerical representations of meaning, which is what makes the search work on sense rather than on exact wording.

Vector memory

Vector memory is where those processed sections live. It is what allows someone to ask a question in their own phrasing and still reach the right passage of a document they have never opened, because they do not need to guess the words you used.

This is a separate component from our other storage. MongoDB holds structured product data. File storage holds the original files. Vector memory holds the searchable representations used for retrieval. Three different systems doing three different jobs.

The language models

The language model layer handles understanding messages, generating responses, following instructions, supporting intent detection, reasoning across requests, and supporting agent actions.

We do not train our own base language models, and we have no plans to. Training one from scratch takes enormous amounts of data, compute, and time, and it would not make The Double better at the things that actually matter for representing a person. Strong open-source models already exist and are improving quickly. The work that makes a Double good is in choosing them, adapting them, and building the system around them.

We build on open-source models and run them on infrastructure chosen for inference speed, which is what makes conversation at a natural pace possible. We separate models by job rather than using one for everything, with one handling written conversation, one supporting spoken interaction, and one handling reasoning, tool selection, and workflow decisions.

The architecture is deliberately not tied to a single model or provider. Different models can be used for different parts of the product based on speed, quality, cost, and infrastructure requirements, and keeping that flexibility is a design goal rather than an accident.

Echo, our voice model

Echo is our voice model, and it is the piece we have put the most engineering into. It handles the whole spoken path, covering speech recognition, speech generation, voice cloning, streaming audio, and live response generation.

Echo is built on top of open-source foundations and developed into our own model from there. We are not claiming to have invented text to speech. What we built is the version of it that a Double needs, tuned specifically for cloning an individual voice from a short sample, for naturalness and expressiveness, for streaming, for low latency, and for staying reliable across a long real-time conversation.

We self-host Echo rather than calling anyone else's voice API. That is more work, and it is the reason we can say exactly where your voice is processed.

Speed is the constant pressure in this layer. In a written chat, a short pause reads as the system working. In a spoken conversation, silence reads as something being broken, and people start talking over it. Everything in the voice path has to complete inside the window where a person still believes they are in a conversation, which is also why turn-taking and interruption handling matter so much. A voice system that cannot be interrupted does not feel like a conversation, it feels like a recording that will not stop.

How AWS and NVIDIA fit together

These two get confused for alternatives, so it is worth being clear. They do different jobs.

NVIDIA provides the GPU computing layer, meaning the accelerated hardware needed to run demanding voice models and hold multiple live conversations at once. Our voice models are self-hosted on NVIDIA GPU infrastructure.

AWS provides cloud application infrastructure, backend services, and file and media storage.

We accepted that extra work for two reasons. Voice samples are the most sensitive thing a user gives us. And the part of the product that most defines it should not sit behind another company's service, subject to their pricing, their limits, and their decisions about what our users' voices may be used for.

What the model sees on every message

This is the part most people find surprising, so it is worth spelling out.

The underlying model does not remember you between requests. It has no ongoing awareness of who you are, what was said earlier, or what happened yesterday. For every single message, our application assembles the context and hands it over.

That context includes:

  • Platform instructions
  • The owner's persona instructions
  • The owner's rules
  • Recent conversation messages
  • Knowledge retrieved for this specific question
  • The actions currently available
  • The current message

A Double's memory therefore lives in the surrounding application, database, knowledge, and retrieval systems rather than permanently inside the base language model. This is not a limitation we are working around. It is what makes the system inspectable, updatable, and controllable, because everything the model knows in a given moment is something that was put there deliberately.

Where your data lives

Structured product and user data is stored in MongoDB. This includes user accounts, Double configurations, knowledge-source information, persona and behaviour settings, conversation settings, skills and workflows, contacts and leads, usage information, and references to uploaded files.

Files and media are stored on AWS. This includes documents, PDFs, images, audio, voice samples, presentations, and other uploaded media. MongoDB is not where original media lives.

An example makes the relationship between the three storage systems clear. When you upload a PDF:

  • The file itself is stored on AWS
  • MongoDB records who owns it and which Double it belongs to
  • The document is processed and its knowledge is added to vector memory
  • The RAG layer retrieves the relevant sections during conversations

Data is encrypted both while moving between systems and while stored.

User rules and guardrails

Rules and guardrails are a separate control layer around the Double. This layer covers persona instructions, tone, formality, humour, response length, goals, behavioural rules, restrictions, action permissions, and fallback behaviour for situations the Double was not built for.

The simplest way to put it is this. Knowledge tells a Double what you know. Rules tell it how to behave.

We keep these deliberately separate. Knowledge is treated as material to draw on, never as instructions to follow, so an uploaded document cannot quietly change how your Double behaves. These rules are supplied as part of the context for every request, they run alongside system-level guardrails that cannot be overridden, and permissions are checked before an action is carried out rather than after it has been decided.

Security and infrastructure control

Part of our security approach is simply controlling more of our own infrastructure. By self-hosting adapted open-source models, we reduce our dependence on external AI providers and gain control over where inference takes place, how models are deployed, how voice processing is handled, how data flows, how long it is kept, and how our components communicate.

It is worth saying plainly that self-hosting on its own does not make a system secure. The accurate claim is narrower and more useful. Self-hosting gives us greater control over the processing path, model deployment, and which outside providers are involved in handling user data at all.

Security itself is built on top of that control, and includes encryption in transit and at rest, access controls, authentication, tenant separation, secure file storage, infrastructure permissions, logging and monitoring, key and secret management, and data-retention controls.

Usage and capacity

The product tracks usage across voice minutes, text replies, knowledge sources, training words, voice clones, and concurrent calls. These records support plan limits, usage visibility, billing, and internal monitoring.

Concurrent voice calls are the interesting one technically, because that number is tied directly to available GPU capacity. Every live voice conversation occupies real hardware, which is a constraint text conversations simply do not have, and it is one of the reasons the voice layer gets so much of our attention.

How one conversation moves through the system

Putting it all together, a single exchange looks like this:

  1. A visitor types or speaks through a Double Page or website agent
  2. The request reaches our backend
  3. The intent engine determines what the visitor is trying to do
  4. The system applies the Double's persona, rules, and guardrails
  5. When knowledge is required, the RAG layer searches vector memory
  6. Relevant information is retrieved from the owner's processed knowledge
  7. A language model generates the response or supports the next decision
  8. For voice, Echo transcribes the incoming speech and generates the spoken reply in the owner's cloned voice
  9. When an action is required, the agent layer triggers the relevant skill or workflow
  10. Structured data is maintained in MongoDB and files remain on AWS

All of that has to finish before a person speaking out loud notices it happening, and that single constraint has shaped more of our architecture than anything else.

What is coming next

Three pieces of work are underway.

A Double that can attend meetings. Booking, rescheduling, and cancelling meetings already work today as a skill. What is coming is attendance. Your Double joins the call, takes its own notes, and can answer on your behalf in your voice, so the meeting still has the benefit of your knowledge at moments when you could not be in the room. Anything it captures there can be added back into your Double's knowledge afterwards, if you choose to keep it, so the next conversation starts from what the last one produced.

More places to reach a Double. WhatsApp, Slack, and Google Chat are the channels people ask for most, because they are where conversations already happen. Each will be another way to reach the same Double, carrying the same knowledge, rules, personality, and voice rather than becoming a separate system with its own configuration.

Bring your own key. We are building support for BYOK, so that users can connect their own model keys and run their Double against infrastructure they control directly. This is the furthest extension of the principle already running through everything above, which is that the person should hold as much control over their own data as we are able to give them.

The part we care about most

Everything in this post is engineering detail, and all of it exists to serve three things. Your knowledge, your guardrails, and your security.

Your knowledge stays yours. It is stored separately from any model, it is not used to train models, you can see what is in it, and you can remove any part of it at any time and have that take effect immediately.

Your guardrails stay yours. What your Double says, what it declines to say, which actions it may take, and the point at which it should stop and hand back to you are all your decisions, and all of them can be changed whenever you want.

Your security is not a later phase of the roadmap. It is the reason we fine-tune and self-host our own voice models rather than sending voice samples to another company, the reason the rules layer is kept separate from the knowledge layer, and the reason permissions are checked before an action rather than after it.

A Double carries someone's name, knowledge, and voice. Every other decision in this stack exists to make sure it carries them the way that person intended.

Sahith Krishna

Sahith Krishna

Founder & CEO