---
title: "How We Train an AI Double to Sound Like You"
description: "Voice, response time, knowledge, persona, and boundaries. The five things that decide how close a Double gets to the person it belongs to."
url: https://thedouble.ai/blog/how-close-is-your-double
---

1.  [Home](https://thedouble.ai/)
3.  [Blog](https://thedouble.ai/blog)

[News](https://thedouble.ai/blog/category/news)

# How We Train an AI Double to Get Close to the Real Person

Voice, response time, knowledge, persona, and boundaries. The five things that decide how close a Double gets to the person it belongs to.

[Sahith Krishna](https://thedouble.ai/authors/sahith-krishna)3 August 20267 min read

Almost everyone asks us the same question when they first hear about The Double. How close is it? Close in voice, close in the way I answer, close in the way I would actually say something to somebody who asked me that.

It is a fair question, and it is the one we built the company around. A Double is only really a Double when it gets close to the person it belongs to. Anything short of that is a generic assistant wearing someone's name.

Getting close is not a single step. It comes from five things working together, each one closing a little more of the distance:

-   **Voice**, so the interaction feels familiar
-   **Response time**, so it feels like a conversation
-   **Knowledge**, so there is something worth saying
-   **Persona**, so it answers the way you would
-   **Boundaries**, so it stays trustworthy

Here is how we build each one.

## **Voice: why we fine-tuned an open-source model**

Voice is where most of our engineering has gone, because it is where the distance shows up fastest. Somebody will accept a slightly imperfect answer. They will not accept a voice that is not yours.

A person's voice carries their pace, their warmth, the small gaps where they stop to think, and the breath before a longer sentence. Those are the things people recognise in each other, and the first things a generic voice loses.

The existing options did not get us there:

-   **Default voices.** Many systems ship with voices that are already trained. They sound good, and they belong to nobody in particular.
-   **Long recordings.** Where cloning is offered, it has usually meant hours of clean audio, which is more than most founders, coaches, or consultants will ever sit down and record.
-   **No cloning at all.** A great many products simply do not let users clone their own voice.

So we took an open-source voice model and fine-tuned it heavily for one job, which is cloning a specific person well from a small amount of audio, and we self-host the result.

We did not build a voice model of our own and we do not claim to have. We adapted an existing one for a purpose it was not shaped around, and that work is why a Double sounds like its owner rather than like a product.

## **Thirty seconds of audio**

Cloning a voice on The Double starts with a thirty second recording. We arrived at that number by working backwards, starting with longer samples and cutting them shorter, then shorter again, watching for the point where quality began to fall away.

Thirty seconds is short enough that nobody keeps putting it off, and long enough to give the model what it needs.

What goes into those thirty seconds matters as much as the length. You are not asked to talk freely. We give you a short passage to read, written to cover the vowels and the phonetic range the model has to hear before it can reproduce sounds you never actually made.

Inside that audio, we analyse:

-   **Frequency**, the underlying pitch of the voice
-   **Breaths**, where you take them and how audible they are
-   **Gaps**, how long you pause between words and phrases
-   **Movement**, where your pitch climbs and where it drops away

Those last two are what people genuinely recognise in each other. Two voices with a similar tone can still be told apart instantly by the way they pause.

From there a clone is usually ready in a minute or two. It should not take an afternoon to make your Double sound like you.

## **Response time**

Response time is part of the resemblance rather than a separate performance number. Somebody who waits three seconds for a reply stops believing they are in a conversation, however good the voice is.

A Double should begin answering at roughly the moment you expect a person to begin answering. Everything in the path underneath it is built to protect that.

Interruption is the other half of the problem. Real conversations are full of people cutting in, changing direction halfway through a sentence, or adding something before the other person has finished.

A Double can be interrupted while it is speaking. It stops, and it takes what you have just said into account alongside everything that came before it. Without that, you are not talking with something. You are listening to something that has not noticed you.

## **The knowledge behind the voice**

A voice clone is only a sound until there is knowledge behind it, and the novelty wears off in about two exchanges. What a Double knows is what makes the conversation worth having.

Knowledge can come from:

-   Custom text written specifically for your Double
-   PDFs, documents, and presentations
-   Websites and product pages
-   Videos
-   Social content, which can stay synced rather than uploaded once

Social is worth its own mention because it does not sit still. What you posted this month often represents your current thinking better than a document you wrote last year.

The part people underestimate is choosing what goes in. A Double should not be trained on everything that exists simply because it exists. It should be trained on the material that helps it represent you accurately.

**For a founder,** that usually means the company story, how the product actually works, pricing, real customer examples, and the way you talk about where things are heading.

**For a coach,** it means the frameworks, the exercises they actually use with clients, and the questions that come up every week.

Where the source material is outdated, vague, or contradicts itself, the Double will be too.

## **Teaching it how you respond**

Persona settings define how your Double responds, and they matter because two people can know the same thing and explain it completely differently. One is short and practical. Another works through an example first. A third asks a question back before offering any advice at all.

None of that lives in your documents, so it has to be described directly:

-   How long your answers usually run
-   How formal you are
-   Whether you use humour
-   Whether you hedge or state things plainly
-   What you always mention when a particular subject comes up
-   What you never raise at all

This is also where you decide how your Double behaves when it does not know something. Saying there is not enough information, or that this is a question for the real person, is what an actual human does in that position.

A Double that always has an answer ready is already behaving unlike its owner.

## **Knowing what not to answer**

Guardrails decide what a Double will not answer, and they are as much a part of the resemblance as anything else. We run system-level guardrails on every Double, and each owner sets their own boundaries on top of those.

Take a public speaking coach. Somebody opens a conversation with their Double and asks how to write a Python function, or wants advice on a bank loan, or raises a health problem. A general assistant will answer all three, because answering is what it does.

That is the wrong behaviour, and not mainly because the answers might be poor ones. It is wrong because the real coach would never have answered them. They would have said this is not their area, and moved the conversation back towards what they can genuinely help with.

The standard we hold a Double to is not what an AI system is capable of saying. It is what the person it represents would have said.

## **Getting closer over time**

A Double is not finished on the day it is set up, because the person behind it keeps moving. People change their minds, companies change their pricing, products ship, and opinions sharpen. A Double that represented someone accurately in March will have drifted by September if nobody has touched it.

Conversations are what show you the gaps:

-   **A question it answers weakly** is a knowledge gap you can fill directly
-   **An answer that does not sound like you** is a persona instruction to adjust
-   **A question it should not have answered** is a boundary to tighten

Owners who read back through their conversations end up with noticeably better Doubles than owners who set one up and walk away. The improvements are almost always small and specific rather than dramatic.

## **What closeness actually means**

We are not trying to build a copy of a mind. That is neither realistic nor what anybody actually needs.

What we are building is a version of someone close enough to be genuinely useful, in their voice, with their knowledge, responding the way they would respond, and staying inside the limits they have set.

No single layer gets you there. Take away any one of the five and the distance opens up again.

The point of getting close was never to make the person unnecessary. It is to make more of them available.

[Sahith Krishna](https://thedouble.ai/authors/sahith-krishna)

Founder & CEO

## Keep reading.

[All posts](https://thedouble.ai/blog)

[

News

### Introducing The Double

Knowledge moved from elders to books to search engines. The person kept getting left behind. The Double is our attempt to bring them back.

03 Aug 2026

](https://thedouble.ai/blog/activate-yourself)

[

News

### What Is an AI Double, and How Is It Different From a Chatbot?

A chatbot answers on behalf of a product. A Double represents a person, in their voice and their knowledge. Here are the six layers behind it.

03 Aug 2026

](https://thedouble.ai/blog/what-is-an-ai-double)

[

News

### What an AI Double Can Actually Do Today

Answer questions, present your deck, share files, qualify people, book meetings, update your CRM. All inside a conversation, in your own voice.

03 Aug 2026

](https://thedouble.ai/blog/ai-double-capabilities)
