NotesBefore You Send

Before You Paste Company Data into ChatGPT — Four Places to Check

Whether company data may go into ChatGPT is decided by your policy and your contract, not by an article. Once that is settled, here are the four places I check before pasting, why personal data, trade secrets, credentials and system details need different handling, and what a pattern can and cannot find.

Before You Send Field Note

Written by k-wada (Legacy Tools)

この記事を日本語で読む

On this page
  1. This article does not authorise the paste
  2. The four places I check
    1. 1. The greeting and the signature
    2. 2. The quoted thread underneath
    3. 3. Client names and project names
    4. 4. Logs and configuration
  3. Before and after
  4. What survives the replacement
  5. Four kinds of information, four different answers
  6. Contract and settings change the answer — the ChatGPT case
  7. What detection can and cannot do
  8. The four steps, short
  9. Sources
A hand-drawn diagram — shaded passages in a document on the left are replaced in the middle, the text then reaches an AI chat input on the right, and four small cards stand for the four places to check

Before I paste a work email into ChatGPT, I replace the parts the model does not need with labels like [client]. There are four places I look: the greeting and signature, the quoted thread below, client and project names, and anything from a log or a config file.

All of that is the second step, though. The first one is not mine to answer for you, so it goes first.

This article does not authorise the paste

Whether your company's information may be entered into an AI chat is decided by your internal policy, your approved tool list and the contract you are under. No amount of relabelling changes that answer.

  • Is the service — and the account you are signed in with — approved for work?
  • Under that contract, how is what you type handled?
  • Does a client agreement or a team-level rule say something narrower?

If the information is not allowed in, replacing a name with [client] does not let it in. Replacement is what you do after the decision, not instead of it. Anything you are unsure about belongs with whoever owns privacy and security where you work. What follows is the routine I use on my own text, not legal advice.

The four places I check

Between pasting into the input field and pressing send, my eyes go back to the same four spots. Less a sequence than four places worth returning to.

1. The greeting and the signature

  • What sits there: names, company, department, job title, phone number, email address
  • Why it gets missed: while re-reading the request, attention stays in the body of the message
  • What I do: replace with [contact] and [client], and drop the whole signature block — it is rarely part of the question

2. The quoted thread underneath

  • What sits there: signatures from earlier messages, people from an unrelated thread, the original recipients of a forward, old attachment names
  • Why it gets missed: it is below the fold of what you just pasted
  • What I do: scroll to the very bottom of the pasted text before sending, and cut quoted material the question does not need

This is the step I forget most often. On a reply, the quoted history is frequently longer than the message itself.

3. Client names and project names

  • What sits there: company names, brand names, project code names, internal system names, team nicknames
  • Why it gets missed: they have no fixed shape, so the rule-based detection described below does not see them
  • What I do: decide first whether they are confidential at all; then replace the ones that are allowed in with [client] and [project]

4. Logs and configuration

  • What sits there: API keys, tokens, passwords, connection strings, internal URLs, hostnames, IP addresses
  • Why it gets missed: pasting an error log wholesale brings along lines that have nothing to do with the question
  • What I do: credentials do not get a label — they do not go in at all. One already pasted gets revoked and reissued

Before and after

Say I want help rewriting a reply, and I paste this:

Hi Jane Doe,

Thanks for your time earlier.

About the system migration at Acme Logistics, I would like to
set up one more call next week.

You can reach me at jane@example.com or +1 555-0100.

If the request is only "make this read better", the real name, the address and the number are not part of it.

Hi [contact],

Thanks for your time earlier.

About the system migration at [client], I would like to
set up one more call next week.

You can reach me at [email] or [phone].

Something is still in there, and what to keep depends on what you are asking for.

  • Worth keeping: the relationship — who is writing to whom, and what kind of message it is. Strip that and the rewrite comes back addressed to nobody
  • Worth generalising: "system migration" and "next week" narrow down the industry, the timing and the size of the deal. If the rewrite does not need them, "a project" and "soon" will do

My first instinct had been to strip anything that looked sensitive, and that went badly: once the names are gone, the text stops saying who is writing to whom, and so does the answer. The question I ask now is not "can I remove this word" but "does this request need this particular value?"

What survives the replacement

A label is not anonymisation. In the example above, the industry, the stage of the work, the timing of the call and the relationship between the two people are all still on the page.

Ask about the same account a few times and the fragments accumulate. A role, a date and a project description are often enough for someone in that industry to work out who is meant, even with the names gone — and the smaller the market, the faster that goes.

Four kinds of information, four different answers

"Personal information" as a single bucket puts things together that need different handling. I split them four ways.

Personal data — names, contact details, employer, the client's staff — is what replacement is actually for. A label keeps the structure of the message and drops the identity. This is also the category data protection law is about, so whether it may be entered without the person's consent depends on the purpose and the contract.

Trade secrets and confidential business information — client names, deal names, pricing, specifications, unannounced plans — are governed by your policy and your client agreements. A label does not change whether they are allowed in. Where they are allowed in, the work is to include no more of the specifics than the question needs.

Credentials — API keys, tokens, passwords, connection strings — are not a replacement target at all.

System details — internal URLs, hostnames, IP addresses, internal system names — are rarely needed as exact values. [internal URL] and [hostname] work, and so does describing the role instead of the value ("our internal auth server").

Contract and settings change the answer — the ChatGPT case

How a service handles what you type varies by service, by plan and by setting, so "AI is unsafe" is not a useful rule. Taking ChatGPT as the example, this is the default handling OpenAI describes on its own pages.

Table 1 Default handling as described on OpenAI's own pages. Plan names and terms change; the sources and the date this was checked (2 September 2026) are under "Sources" below.
What you useUsed to improve models?
Personal accounts (Free, Plus and similar)It can be, depending on the setting. Data controls turn it off
ChatGPT BusinessNot by default
ChatGPT EnterpriseNot by default
APINot by default

Three things I keep in mind when reading a table like this one.

  1. Which row you are on is not obvious from the screen. An account invited into a company workspace and a personal account being used for work look much the same. If you cannot tell, your workspace admin can
  2. Setting names move. The wording of the consumer data controls has changed before, so I check the actual settings screen rather than quoting a label
  3. "Not used for training" is not "not stored". Retention periods, deletion, and whether an admin can read a conversation are separate items in the same documentation

Japan's data protection authority, the Personal Information Protection Commission, published a notice on generative AI services on 2 June 2023 asking users to check the provider's terms and privacy policy before deciding what to enter. For businesses handling personal data it goes further: entering personal data in a prompt without consent, where that data is then processed for anything beyond producing the response, may breach the law — and businesses are asked to confirm that the provider does not use the data for machine learning. (That is a summary of a 2023 document; the original is linked under "Sources".)

The confirming is homework on the sender's side. How it applies to your organisation is a question for the people who own that decision there.

What detection can and cannot do

I build the browser extension that does this replacement, and writing the detection side is what made the line clear to me.

Table 2 A split by what rule-based detection (patterns, essentially) can work with. It is not a statement about accuracy.
CategoryExamples
Usually findable by patternemail addresses, phone numbers, postal codes, card numbers, IP addresses, API keys with a known prefix
Needs judgement from contextpersonal names, client names, project names, department names, internal system names
KeyShould not be entered at alllive API keys, tokens, passwords, connection strings

The first row is written in a format. Something sits either side of an @; digits appear in a known length; a key starts with a known prefix. Card numbers can even be checked arithmetically to see whether the digits are plausible.

But anything written outside the format is missed even on that row: spacing variants, a value broken across two lines, an in-house identifier scheme, a key with an unfamiliar prefix. It also runs the other way — an unrelated run of digits can be flagged as a phone number.

The middle row has no format at all. "Acme Logistics" and "the new membership system" are ordinary words. The rule-based detection in this extension cannot reliably tell, from context, which of them should be held back. Techniques that infer names and organisations from context do exist — named entity recognition, and the DLP products built on it — but what I can speak to is the thing I built, and there the dependable answer was for the person doing the work to register the term.

Doing it by hand every time did not last for me. Over enough emails and meeting notes, it is the same values, found and retyped. So the extension (Safe Privacy Mask) detects text in the input field and replaces it before sending. It works on the first row of the table; client and project names are custom rules you add yourself. The permissions it asks for, what it stores on the device, and whether it sends anything are listed on the transparency page, along with how to check for yourself.

It is not a substitute for a company DLP or an internal policy. It finds what can be written as a pattern, which is not the same as finding everything sensitive. What actually gets sent is still a decision made in front of the text.

One boundary worth naming: all of this is about text pasted into an input field. When a file is attached instead, the information lives inside the file — a different check, which I wrote up separately in what is stored in PDF metadata.

The four steps, short

  1. Check the policy — is this information allowed into an AI chat at all? Anything unclear goes to the people who decide that
  2. Check the environment — is this service, this account and this contract approved for work?
  3. Replace what the question does not need — walk the four places; credentials are not replaced, they are left out
  4. Re-read for context — with the labels in place, can a person, a company or a deal still be worked out?

There is no need to suspect the whole message. Only the parts that were never the point.

Sources

Table 1 reflects what those pages said on 2 September 2026. Plan names, setting names and terms change, so I re-read them about every three months, and sooner if I notice a change to a plan or to the terms. Table 2 is my own split by detection shape, not a classification from a source.

Tags: ChatGPT・personal data・credentials・before you send