The Confidentiality Leaks Hiding in Your AI Tools
Every AI tool your team touches is, by default, a potential data pipeline back to the vendor. Sometimes that data trains the next model. Sometimes it doesn't. The gap between those two outcomes is usually a setting nobody checked, a contract clause nobody read, or a habit nobody thought to flag as a habit.
This isn't a guide to any one platform. Vendors change their policies often enough that anything platform-specific goes stale within a quarter. What doesn't go stale is the set of questions you should be asking, of any tool, before you or your team puts real information into it. Below is a red, yellow, green checklist across the six places this risk actually lives.
Consumer vs. enterprise/API tier
π΄ Red: Your team is using the free, consumer-facing version of a tool for work product. Free consumer tiers are frequently the version most likely to use inputs for training by default, because the product itself is the training pipeline.
π‘ Yellow: You're on a paid individual or "pro" consumer plan. Better than free, but paid consumer tiers don't automatically mean no training. The commitments are often thinner than enterprise agreements and can change with a policy update.
π’ Green: You're on an enterprise agreement or using the API directly. These tiers typically come with contractual no-training commitments as a baseline feature, not an opt-in you have to hunt for.
Data retention settings & duration
π΄ Red: You don't know how long the vendor retains your inputs, or the answer is "indefinitely" with no clear deletion mechanism.
π‘ Yellow: There's a stated retention window, but it's long, vague, or tied to vendor discretion ("as needed for safety and abuse monitoring").
π’ Green: There's a short, specific, and controllable retention window, ideally with an option to zero it out entirely for business use.
In-product feedback signals
π΄ Red: Nobody on your team knows that a thumbs-up, a star rating, or even correcting an AI's answer can itself flag that exchange for human review or training, separate from whatever the base data policy says. This is the one people get wrong most often, because it feels like a UI interaction, not a data decision.
π‘ Yellow: Your team knows this in theory but hasn't built it into practice. Nobody's told the associate not to thumbs-up the response that happened to include a client's deal terms.
π’ Green: You've made "don't rate outputs that contain confidential information" an actual instruction, not an assumption, and it's part of onboarding for any tool that touches client or company data.
Vendor DPA / no-training addendum availability
π΄ Red: The vendor doesn't offer a data processing agreement or a no-training addendum at any price point. If a vendor can't or won't put a no-training commitment in writing, that tells you something about how central your data is to their business model.
π‘ Yellow: A DPA or addendum exists but you haven't signed it, or it's available only at a tier you're not on.
π’ Green: You have an executed DPA or no-training addendum in place, and someone has actually read it rather than assumed it says what the sales page implies.
Contractual opt-out clauses / ToS language
π΄ Red: You're relying on the general terms of service, which are unilaterally amendable by the vendor at any time, often with nothing more than a notice buried in a changelog.
π‘ Yellow: There's a specific opt-out mechanism in the terms, but it's opt-in by design, meaning the default is "in" and you have to actively find and flip the setting.
π’ Green: No-training is the contractual default for your tier, survives a policy change without your consent, and is backed by something more durable than a settings toggle you could lose in a UI redesign.
Confidentiality & trade secret exposure
π΄ Red: You've put material into a tool that would qualify as a trade secret or client confidential information, without confirming the tool's training and retention posture first. Once that information has left your control, "we didn't mean to" is not a defense, and depending on what was disclosed, it can jeopardize trade secret status entirely.
π‘ Yellow: You have internal guidance about what's off-limits for AI tools, but it's informal and not consistently enforced across the team.
π’ Green: You have a written policy about what categories of information can and can't go into which tiers of AI tools, it's actually distributed, and your client and vendor contracts contain language addressing AI tool use and confidentiality obligations directly.
Why this matters more than it looks like it does
None of these are exotic risks. They're the kind of thing that feels like a settings page until it's a breach notification, or a due diligence question you can't answer cleanly, or a client asking pointed questions about where their information went. The fix isn't avoiding AI tools. It's treating "what happens to this input" as a real question with a real answer, not an assumption you inherited from the default settings.
If you're building or reviewing contracts that touch AI tool use, whether that's a vendor agreement, an employee policy, or a client-facing confidentiality clause, that's exactly the kind of gap worth closing before it becomes a problem instead of after.