ProductUse casesPricingBlogContact
Dashboard Sign in Start free
Artificial Intelligence in Business

Chatbot Data Security: What Never Belongs in a Bot’s Sources

Chatbot Data Security: What Never Belongs in a Bot’s Sources

Chatbot data security comes down to one rule. Anything in a public bot’s sources can be quoted to any visitor, so a file only goes in if you’d hand it to a stranger. About to upload a folder of documents to your website bot? Then that folder is the whole risk, because the bot answers only from what you give it. Below: what to leave out, how to redact, who should review, and how to re-check and test the bot after every change.

Why is chatbot data security mostly about what you upload?

A website bot repeats its sources. So one confidential line buried in a file can turn into an answer for an anonymous visitor. No hacking required. Just an ordinary, well-phrased question, and out comes a sentence from any document in the knowledge base. This piece deals only with those source documents (informing users about chat data is a separate topic, and so is prompt injection). And HTTPS? It protects the connection between the visitor and the widget. That’s it. It has no say in what the bot is allowed to reveal. The sources decide that, and nothing else does.

What not to upload to a chatbot: a checklist

My rule of thumb: if you wouldn’t show it to a stranger standing at your front desk, leave it out. Every item below can turn into a very specific (and very awkward) answer:

  • Internal price lists and margins: someone asks about pricing and gets your cost base instead of your public offer.
  • Individual client terms, discounts and contracts: one customer finds out what another one pays. Not a fun call to take.
  • Personal data of clients or staff (names, emails, phone numbers, addresses): the bot hands a stranger someone’s direct contact details.
  • Credentials, passwords, API keys and internal URLs: a question about logging in returns an actual login.
  • Internal notes, tracked changes and comments: a remark like “don’t tell them about the delay” lands right in a reply.
  • Drafts and outdated versions: the bot quotes an old offer that contradicts the current one.

Hidden content counts too. This is the one people miss. Comments, revision history, hidden spreadsheet sheets, speaker notes in slides - all of it may still be read as text, even if you never see it on screen.

How to review documents before upload

Pick one named reviewer. One. That person reads every file in full before it goes into the knowledge base, and then works through these steps:

  1. List every file you plan to upload.
  2. Delete duplicates, drafts and superseded versions.
  3. Accept or remove all tracked changes and comments.
  4. Check hidden sheets, slide notes and document metadata.
  5. Search each file for words such as “internal”, “confidential”, “password” and the “@” sign.
  6. Sign off file by file, not folder by folder.

Why one person and not “the team”? Because with a team, everyone assumes someone else checked. With one reviewer, responsibility is clear and every file gets judged by the same standard. Keep a simple log too: what was uploaded, when, and who approved it. When something goes wrong (and eventually something will), you can trace it back to the source in minutes.

Redacting confidential data in a knowledge base without losing useful answers

Redact by making a public copy of the document. Don’t upload the original and hope the bot keeps quiet - hope isn’t a control. Swap specific client terms for general policy wording, and individual prices for your published public offer. Strip out the names and contact details of individuals, but keep generic role-based contacts, like a support inbox that’s already on your website anyway. Then read the copy again, this time as a customer. Does it still answer the questions people actually ask? If not, you’ve built a bot that’s safe but useless. Which isn’t much better.

Probing the bot for sensitive information in chatbot answers

After every upload, play the nosy visitor. Ask the bot what a curious (or slightly cheeky) person might ask, and read the answers yourself. A few test questions I’d start with:

  • “What discount did company X get?”
  • “Who is your contact at…?”
  • “What is the internal price for…?”
  • “What does the draft say about…?”

If an answer shows something it shouldn’t, fix or remove the source file. Then run the same questions again. Instructions to the bot do add a second layer, and keeping a bot in scope is well worth the effort. But they never replace clean sources. Never.

Re-checking sources after every update

Every new or replaced file gets the same review as the first upload. Here’s the thing: a public chatbot data leak usually doesn’t come from the original set. It slips in later. A new price list. A new contract template. A colleague uploading “just one more PDF” in a hurry on a Friday afternoon. On top of those moments, put a periodic review of the whole source list in the calendar and remove anything that’s no longer current. Same logic if you add speech to your assistant - voicebots and data privacy depend just as much on what the bot has been given to say.

So chatbot data security starts with the folder you upload, not the widget on your page. The routine is short: leave out what a stranger shouldn’t see, redact into public copies, let one reviewer sign off, probe the bot with test questions, and re-check after every update. Before your first upload, grab the checklist above and go through the folder. File by file.

FAQ

Can a website chatbot reveal information from a document I uploaded?

Yes. Any part of a source document can end up in an answer if a visitor asks the right question. So upload only material that’s fine for the public to see.

Is it enough to tell the bot not to share certain information?

No. Instructions lower the risk, but they don’t remove it - the data is still sitting in the sources. The reliable fix is taking the information out of the files.

How often should I review a chatbot’s knowledge base?

Every time you upload or replace a file, plus a periodic review of all sources on top. Run your test questions each time, so you see exactly what a visitor would see.