NearSeal

2026-08-16

Should you upload a sensitive file to ChatGPT or another AI chatbot?

Pasting a contract into ChatGPT for a quick summary, or uploading a resume to have an AI clean up the formatting, has become as routine as attaching a file to an email. The question worth asking before that upload is the same one people have started asking about cloud drives and file-transfer links: where does that data actually go, and does it stay under the rules the provider's settings page implies? Three real, dated incidents — one from a company's own security team, two from open court filings — give a more concrete answer than a privacy-policy summary does.

What "uploading a file" to a chatbot means, by the provider's own account

OpenAI's own help center is fairly direct about the mechanics: prompts, file uploads, and images sent to ChatGPT are logged on its servers, and depending on plan and settings may be retained, used to improve models, or both. On the consumer tiers — Free, Plus, and Pro — conversations may be used to train future models by default, unless a user finds Settings → Data Controls and turns off "Improve the model for everyone"; ChatGPT Business, Enterprise, and Edu accounts are excluded from training by default. None of that is secret or unusual among AI chatbots generally — it's disclosed, and the opt-out exists. But it also means the file itself, not just a description of it, is content the provider now holds under its own retention rules — not yours.

When a lawsuit between two companies changes what "delete" means for everyone else

On May 13, 2025, Magistrate Judge Ona T. Wang, presiding over the consolidated copyright litigation between The New York Times and other news organizations against OpenAI, ordered OpenAI to "preserve and segregate all output log data that would otherwise be deleted" — including chats a user had deleted, and ones sent in Temporary Chat mode, which OpenAI's own product copy describes as not being saved. The order applied to ChatGPT's Free, Plus, and Pro tiers, not to Enterprise, Edu, or API customers under zero-data-retention agreements (National Law Review's account of the litigation; OpenAI's own explanation is at openai.com/index/response-to-nyt-data-demands). None of the users whose deleted chats got frozen were parties to that lawsuit — the freeze happened because of a copyright dispute between OpenAI and news publishers that had nothing to do with them. The going-forward preservation duty was lifted in fall 2025, but data tied to plaintiff-flagged accounts kept being retained, and on January 5, 2026 a federal district judge affirmed the underlying order: OpenAI must produce 20 million de-identified ChatGPT logs to the news-organization plaintiffs. Whatever a provider's privacy policy promises about deletion, that promise sits downstream of whatever litigation the provider happens to be in — and a user has no way to know, at the moment they upload a file, whether that will end up mattering.

A leak that needed no breach, no hack, and no lawsuit at all

In August 2025, a "make this chat discoverable" sharing option — something a user had to actively opt into by checking a box when sharing a conversation link — resulted in nearly 4,500 shared ChatGPT conversations turning up in ordinary Google Search results, some containing resumes, legal names, and mental-health disclosures. OpenAI's Chief Information Security Officer, Dane Stuckey, called it a "short-lived experiment" that "introduced too many opportunities for folks to accidentally share things they didn't intend to," and the feature was pulled within days. No credentials were stolen and no system was compromised — a checkbox behaved exactly as coded, and the result was fully readable plaintext sitting in a public search index anyway.

This isn't a 2025 problem — it goes back to 2023

The exposure risk isn't unique to file uploads specifically, and it isn't new. In March 2023, Samsung lifted an internal ban on generative-AI chat tools; within 20 days, employees had pasted confidential semiconductor source code into ChatGPT twice asking for help fixing bugs, and a third employee submitted an entire internal meeting's notes and asked it to generate minutes. Samsung banned ChatGPT and other generative-AI chatbots company-wide as a result. Nothing was "uploaded" as a file in that case — it was typed and pasted directly into the chat box — but the underlying mechanism is identical to a file upload: the provider receives and stores the plaintext content either way.

Two separate questions — only one is about encryption Should this go into a chatbot? The model needs the plaintext to summarize, extract, or answer A file-encryption tool can't change that — an encrypted file is one the model can't read at all Decide this before you upload Should the file be protected everywhere else it goes? Cloud storage, email, a USB drive, a transfer link — this is where client-side encryption like NearSeal applies Applies before or after the AI step
Whether a file's content should reach an AI chatbot, and whether the file itself should be encrypted everywhere else it travels, are two different decisions — encryption only answers the second one.

Where encryption doesn't apply — and pretending otherwise would be dishonest

No file-encryption tool, NearSeal included, can protect a file's content while an AI model is actively reading it, because summarizing, extracting, or answering questions about a file requires the model to receive it as plaintext — that's the entire mechanism, not a gap. Encrypting a file first wouldn't add protection at that step; it would just hand the chatbot a file it can't open at all, which is encryption doing exactly what it's supposed to do, just not solving this particular problem. Any product that implied otherwise would be making a claim its own cryptography contradicts.

Where NearSeal actually fits

Two real places, both separate from the upload decision itself. First: for a file you decide doesn't need to go into a chatbot at all — and deciding not to upload something is itself a valid protection — if you're keeping or archiving that file anyway, encrypting it before it sits in cloud storage, gets emailed, or gets copied to a drive keeps it out of every one of those exposure paths, independently of whatever an AI provider does or doesn't do with its own copy of anything else. Second: for the original document you still need to store or hand to someone else after you're done using an AI tool on it — that file's exposure everywhere else it goes has nothing to do with whatever retention or training policy applied to the copy an AI provider now holds, and encrypting it there is exactly NearSeal's real, shipped capability: a passphrase-derived AES-256-GCM key via PBKDF2-SHA256 in the default NearSeal format, or the interoperable age-encryption.org format (scrypt + ChaCha20-Poly1305) if the file needs to open with a standard tool later — entirely in the browser, nothing uploaded to run either one. NearSeal can't undo a preservation order, an indexing bug, or a pasted paragraph that already reached a chatbot's servers. What it can do is make sure the file's other copies — the ones sitting in storage, in an inbox, on a drive — aren't an equally readable duplicate of whatever risk that upload already carried.

Sponsored
← NearSeal

This page shows ads only if you consent.