Is It Safe to Upload Writing Samples to AI? What Happens to Them
Is it safe to upload writing samples to AI? What ChatGPT, Claude and Gemini keep, what API terms say, and how MyWritingTwin scrubs, stores and deletes them.
By Emmanuel
TL;DR: Is it safe to upload writing samples to AI? Safer than most people fear and riskier than most tools admit, depending on which door you use. The same company runs two rule sets: consumer chat apps can use your chats to improve models or show them to reviewers unless you opt out, while developer APIs from OpenAI, Anthropic and Google (paid tier) state that inputs are not used for training by default. Your samples are still the only place your voice lives, because model capability and user-data conditioning are separate layers and no lab can supply the second. So strip the facts, keep the patterns, upload once to extract a Style Profile, and deploy the profile instead of re-pasting raw emails into every chat. Below: what each vendor says, what MyWritingTwin.com does with your samples at each step, and a six-point checklist for any tool.
You have ten emails that sound exactly like you. Three name a client. One quotes a salary. One is a resignation you never sent. Someone tells you to paste all ten into ChatGPT so it can learn your style, and you hesitate.
The hesitation is correct. It is also the reason most people never get AI to write like them: they either hand over everything or hand over nothing, and both choices are wrong.
Key takeaways: uploading writing samples to AI
- The door matters more than the vendor.OpenAI, Anthropic and Google each run a consumer app and a developer API under different data rules. Ask which one a tool uses.Facts and patterns are separable.Your voice is in sentence length, qualifier placement, transitions and punctuation. Client names and deal figures carry none of it, so strip them before uploading.Nothing sticks in a chat window.Samples pasted into a conversation shape that conversation only. You leak them again every session and the model still starts from zero next time.Extract once, deploy the profile.A Style Profile holds measurements and characteristic phrasing, not your documents. It is the thing worth carrying into every AI tool.MyWritingTwin scrubs before it stores.Samples are anonymized on save, held under row-level security, analyzed through developer APIs, never shared, and deleted with your account.
Is It Safe to Upload Writing Samples to AI?
It depends on which door you walk through, and the gap between doors is wider than the gap between vendors. A personal ChatGPT account and an application built on the OpenAI API belong to the same company and operate under different rules about training, human review and retention. The safety question has no vendor-level answer, only a door-level one.
Three separate questions hide inside "is it safe":
- Will the text be used to train or improve a model, or be read by a person? This is the one people mean, and it is the one where consumer and API rules differ most.
- Who can read it later, for how long, and can you delete it? Retention windows and deletion rights are written down. They are rarely read.
- Does the thing you get back contain the thing you put in? If the output is a profile, does it quote your email about the merger? If it is a Custom GPT, did you upload the corpus as a knowledge file that anyone with the link can now query?
The fear behind the first question is usually misplaced in one specific way. A model does not "learn you" from ten pasted emails. Inside a chat, samples sit in the context window and influence that conversation only; open a new one and the model starts from its training distribution again. We covered why in why better models still don't know your voice. The real exposure is the mundane one: the text sits in someone's logs, possibly gets sampled for review, and on some consumer plans is eligible for training. That is a data-handling question, and data-handling questions have documented answers.
Where Do Your Samples Go When You Paste Them Into ChatGPT, Claude or Gemini?
Into one of two rule sets, and the vendor's own documentation tells you which. The consumer apps default to some form of "may be used to improve our models" with an opt-out. The developer APIs default to "not used for training" with an opt-in. Here is what each vendor states, as of this writing.
| Door | What the vendor says about your text |
|---|---|
| ChatGPT, personal plans | The setting that allows OpenAI to use your conversations to improve models is on by default for personal accounts and can be turned off in Data Controls. Temporary chats are excluded. |
| OpenAI API | "Data sent to the OpenAI API is not used to train or improve OpenAI models" since March 1, 2023, unless you explicitly opt in. Abuse-monitoring logs are retained for up to 30 days by default. |
| Claude Free, Pro, Max | Anthropic states it will use chats to improve models if you choose to allow it via a setting, or if a conversation is flagged for safety review. |
| Anthropic API and Claude for Work | "By default, we will not use your inputs or outputs from our commercial products" to train models. |
| Gemini app | Google's privacy notice says human reviewers may read some data and asks you not to enter confidential information you would not want a reviewer to see or Google to use to improve its services. The Keep Activity setting controls this. |
| Gemini API, paid quota | Google does not use prompts or responses to improve its products, and logs them for a limited period only to detect abuse. |
| Gemini API, unpaid quota and AI Studio | Google uses submitted content to provide, improve and develop its products, and human reviewers may read and annotate it. |
Read the table by row pairs and the pattern is obvious. OpenAI appears twice with opposite defaults. Anthropic appears twice with opposite defaults. Google appears three times, and the line runs between paid and unpaid rather than between app and API. "Does Gemini train on my data" has no answer. "Which Gemini door is this tool using" does.
Two consequences follow for anyone about to paste writing samples.
The consumer app is the worst place to do it, and the most common. It is the door with training on by default, the door with human review, and the door where nothing persists between sessions anyway. The ChatGPT memory feature stores facts you state, not measurements of how you write, so the samples buy you one conversation of imitation and then evaporate. You pay the exposure every time and keep nothing.
A tool built on the API inherits the API's terms. That is the structural reason a dedicated profile builder can be safer than the chat app it feeds into: the analysis runs through the commercial door, and what you carry into the consumer app afterwards is the output, not the input.
Why Do You Need to Upload Samples at All?
Because your voice exists nowhere else. Model capability and user-data conditioning are separate layers. Labs improve the first with every release and cannot supply the second, because your writing was never in the training data and no model update adds it.
The obvious workaround is to describe yourself instead of uploading anything. Type "professional yet friendly, concise, warm" into custom instructions and skip the samples entirely. We tested this at length in tone adjectives are not your voice, and the result is always the same: "professional yet friendly" describes thousands of people, so the model returns the average of all of them. AI does not know your tone; it knows words that describe tone. A voice lives in sentence length, qualifier placement, transitions and punctuation, and those four things are only observable in text you actually wrote.
So samples are the price of admission. The privacy question is not whether to pay it but how much, how often and in what form. That reframing changes the answer, because a writing sample contains two different kinds of information mixed together:
| In the email | Carries your voice? | Carries risk? |
|---|---|---|
| Typical sentence of 13 words, shortest 4, longest 29 | Yes | No |
| "Probably" always before the recommendation, never after | Yes | No |
| One colon per email, asides in parentheses | Yes | No |
| Opens with the ask, closes with a first name | Yes | No |
| The client's surname | No | Yes |
| The contract value | No | Yes |
| The date of the board meeting | No | Yes |
Everything in the top half is countable, and counting it is what style extraction does. Everything in the bottom half is a fact that could be replaced with a placeholder without changing a single measurement. "[COMPANY] confirmed the [ACCOUNT_NUMBER] transfer on Tuesday" has the same sentence length, the same clause order and the same punctuation as the original. The style survives the scrub intact.
Once the patterns are extracted, they become a profile: a set of constraints and characteristic phrasing that describes how you write. That profile is an asset you carry, and the raw documents can stay where they were. The better the model gets at following nuanced instructions, the better it executes your profile too, so extraction pays off more with each model release. Re-pasting raw emails into a fresh chat pays off less each time: it leaks the same facts again and still leaves the model with nothing next session.
What Does MyWritingTwin Do With Your Writing Samples?
It scrubs them before storing them, stores them under access rules tied to your account, analyzes them through developer APIs, returns a profile that contains measurements rather than documents, and deletes them with your account. Here is each step in the order it happens, grounded in what the product actually does today rather than what a policy page could promise.
1. Anonymization on save. When you save a sample in the dashboard, the text passes through an anonymization step before it is written to the database. Last names become [LASTNAME], email addresses become [EMAIL], and phone numbers, company names, physical addresses, account numbers, dates of birth, social handles and government ID numbers each get their own placeholder. First names are kept on purpose: "Hi Sarah" and "Dear Ms. [LASTNAME]" are different registers, and the register is part of your voice. The step has language-specific rules for Japanese (family names, postal codes, My Number), French (noms de famille, SIRET, numéro de sécurité sociale) and Spanish (apellidos, DNI, NIE), because a scrub that only understands English names would leave a Japanese sample untouched.
Two limits, stated plainly because the form states them too. Samples longer than 15,000 characters skip the step and are saved as written. And the scrub is a model-based pass, not a guarantee; the sample form says it "isn't 100% perfect" and asks you to remove sensitive details yourself before saving. Samples imported through the ChatGPT app connector get a deterministic server-side email scrub at the import boundary, and the importing agent is required to anonymize names and companies before it sends anything.
2. Storage under row-level security. Samples live in a Postgres database with row-level security, which means the database itself refuses to return a row to any account other than the one that owns it. This is listed in the privacy policy as a security measure, and it is enforced at the data layer rather than in application code that could be bypassed.
3. Analysis through developer APIs. To build your Writing DNA, the corpus is sent to Google's Gemini API (the primary model) with Anthropic's Claude API as the fallback; the anonymization step itself runs on Claude. Audio samples go to OpenAI's Whisper for transcription and nothing else. All three are named in the privacy policy, and all three are reached through their developer APIs, which sit under the commercial terms in the table above, not through the consumer chat apps.
4. What comes back. The output is a Style Profile: sentence-length mean and spread, comma density, vocabulary richness, hedging habits, transitions, always-and-never vocabulary rules, and a set of key phrases pulled from your corpus (at least ten per writing mode). It contains characteristic phrasing, not your documents. Because scrubbing happens before analysis, a phrase extracted from a scrubbed sample cannot carry a surname or a deal figure. Your own samples are also returned to you, formatted, as part of the package, so the decision about whether they go anywhere else (a Custom GPT knowledge file, for instance) is yours and explicit.
5. Sharing is off, and narrow when on. A public sharing toggle exists for the free Writing DNA snapshot. When you turn it on, only the snapshot is shared; writing samples, the Master Prompt and account details remain private. Nothing about your samples is visible to anyone by default.
6. Deletion is immediate. Deleting your account deletes your samples, profiles and Writing DNA data immediately and irreversibly. The deletion routine also erases your person record in the analytics tool. Two things survive, both required by law: purchase records are kept for seven years in anonymized form (amounts, dates, payment references, with your name and email removed) under Japanese accounting rules, and server logs are retained for up to 90 days.
None of this makes uploading risk-free. It makes the risk small, bounded and written down, which is the standard any tool that asks for your writing should be held to.
How Do You Upload Writing Samples Safely to Any AI Tool?
Choose samples for their patterns, strip the facts, use the API door, read the retention line, extract once, and check the output. Six steps, and the first two cost nothing.
- Pick samples for rhythm, not for content. You need the emails whose sentence length and structure are unmistakably yours, not the ones with the most sensitive facts. Our guide to selecting writing samples covers what makes a sample useful; the short version is that a routine follow-up you wrote in ninety seconds often carries more of your voice than a proposal you polished for a week.
- Replace facts with placeholders before you paste. Names, figures, dates, company names. Use bracketed tokens. Your sentence length, clause order and punctuation do not change, so the analysis does not degrade.
- Use the API door, or fix the consumer one. Prefer a tool built on the developer API, or a workspace plan whose terms exclude training. If you must use a personal chat app, turn off the improve-the-model setting in ChatGPT, use Keep Activity off in Gemini, or use a temporary chat. Then remember that whatever you paste there will not persist anyway.
- Read the retention line, not the marketing line. Look for a number of days and a deletion rule. "Your data is secure" is a sentiment. "Deleted immediately and irreversibly when you delete your account; logs kept up to 90 days" is a commitment you can hold someone to.
- Upload once, extract, then stop uploading. The point of the exercise is a profile you deploy inside the AI you already use. After it exists, the raw corpus has no reason to enter another chat window.
- Read the profile before you deploy it. Skim the output for any name or number that slipped through the scrub. The profile is going into custom instructions or a project system prompt, so it deserves the same thirty-second check you would give a forwarded email.
Run the checklist against any tool, including this one. A profile builder that cannot answer questions three and four in writing is a consumer chat app with a nicer landing page.
Your writing samples are the only place your voice exists, and they are also the only thing you need to protect. Separate the two, hand over the patterns, keep the facts, and the trade stops being frightening.
If you want the measurement done for you, build a Style Profile at MyWritingTwin.com from samples that are scrubbed before they are stored, and deploy your Writing Twin for ChatGPT, Claude and Gemini without pasting another email into a chat window.
Make AI write like you, not like a bot
You just read how to tune AI output by hand. MyWritingTwin does it from your real writing: paste a few samples, get a voice profile that works in ChatGPT, Claude, and Gemini.
Create your free writing profileNot ready yet? Get our guide to AI voice profiles by email instead.