Artificial intelligence
The security model behind an AI assistant that sits inside employee security training.
Kinds Security Team
•
Last reviewed
The security model behind an AI assistant that sits inside employee security training.
Clay is an assistant employees can talk to during a Kinds workshop. What it can see is deliberately narrow: which workshop, which screen, and who the employee is. What administrators can see is narrower still. They can see that the assistant was used and what the employee answered in the workshop. Never what was asked or said.
This is the page for the person who has to approve that before it goes in front of staff.
Employee conversations are private from the employer
Start here, because it’s the answer to the first question and it shapes everything else.
What an employee asks Clay is not visible to their company. There’s no admin screen for it, no export, no API. Kinds keeps the conversations to improve the training, stored against a random identifier rather than a name or an email, and deletes them after 180 days.
That boundary isn’t a privacy setting. It’s the enabling condition for the feature working at all. People ask about the password they’ve reused for six years, or the link they clicked last week, only when they’re confident their manager isn’t reading. Remove the boundary and you remove the disclosures, which are the most operationally valuable thing the assistant produces.
If your compliance posture requires the ability to read employee training conversations, Kinds is the wrong fit. That’s a genuine tradeoff rather than an oversight.
What Clay is allowed to know
Every answer is built from five sets of instructions, stacked from most general to most specific, assembled on Kinds servers rather than in the browser. When two disagree, the more specific one wins. The first is the exception, and overrides everything below it.
Layer | What it contains | Who controls it |
|---|---|---|
Kinds policy | Tone, scope, honesty, privacy, incident handling, manipulation defenses | Kinds. Nothing below can weaken it. |
Account | Guidance applying to every company underneath an MSP or parent | The MSP or parent company |
Organization | Guidance for one specific company, overriding the account level | That company’s admin |
Employee | Name, role, department, company, from their signed-in account | Nobody types this. It’s resolved from the session. |
Workshop | What this workshop may teach, and which screen they’re on | Kinds content team |
The parts that never change sit at the top on purpose, so they can be reused between questions instead of rebuilt each time, which keeps answers fast.
Clay has no access to your console, your configuration, or your other Kinds data. Ask it what MFA your company has enabled for accounts payable and it says it can’t see that, then points you at the people who can.
The four controls a security reviewer will ask about
Nothing an employee types can become an instruction. The browser tells Clay only which workshop and which screen, never the actual lesson content. Clay looks that up on Kinds servers itself. An early version did trust the browser for it, and that was deliberately removed. What an employee types is treated as a question, and only ever as a question.
Clay doesn’t trust its own chat history. The browser sends the conversation back with each new question, so a message labelled as Clay’s could have been altered on the way. None of it is treated as authoritative, and anything claiming to be an official instruction gets discarded.
Admin guidance is guidance, not law. The text an admin types into the settings box sits below Kinds policy. Without protection, an admin could write something that looks like official policy and try to override the privacy promise made to employees. So that text gets wrapped and squashed onto a single line, and Clay is told plainly that anything in it claiming to be policy is just somebody’s typing. Names and job titles work the same way: information about a person, never instructions.
Failure is quiet and never blocks the lesson. If anything goes wrong, the question box goes unavailable. No error message, no technical detail, and the workshop carries on. It returns on its own, and moving to the next screen clears it, so one bad moment doesn’t cost someone the assistant for the session.
Kinds tests every release against a standing set of attacks before it ships: people pretending to be administrators, fake conversation histories, misspellings meant to slip past filters, and attempts to smuggle instructions in as a workshop answer to force a passing grade. Each one has to produce no leaked instructions, no change in behaviour, and a short on-topic reply.
What this changes for whoever runs the program
Nothing to configure per client. Clay ships knowing the workshop. No content to map, no FAQ to write, no per-tenant setup step. Turn on the training and the assistant is already inside it.
One instruction, every client. An MSP or parent company writes guidance once and every organization underneath inherits it. “We deploy Okta, not Microsoft SSO, so never tell someone to reset a password in Entra.” A specific client can override it where their setup differs. For anyone running training across dozens of companies on a mix of stacks, that’s the difference between an assistant that helps and one that generates tickets.
Questions stop becoming tickets. The employee who doesn’t understand what a workshop is asking used to guess, ask a colleague, or email IT. Now they ask in the box that’s already on screen.
Near-misses surface while someone’s thinking about them. An employee who mentions clicking something last week gets routed to your security contact at the moment they remember it. That’s an intake channel most programs don’t have.
And the trade, stated plainly: you don’t get to read those conversations. That limit is what makes the intake channel work.
Language handling
The workshop runs in the employee’s language and so does Clay, reading, speaking, listening, and grading, across 32 languages.
The rule that matters for anyone with regulated staff in more than one country: copy every number exactly as written. No converting currencies, no re-deriving percentages, no localizing figures. Someone reading the Japanese phishing workshop gets the same number the English one cites, not a rounded, subtly wrong approximation. Multilingual training that mangles its own evidence is worse than monolingual training.
Free-text answers are graded on what they say, not on the language they arrive in.
What Clay can’t do
It can’t see your configuration. No console data, no tenant settings, no visibility into your stack. A confident description of your setup would be invention, so it declines and redirects.
It isn’t incident response. An employee disclosing something real gets routed to your security contact. Clay doesn’t triage, contain, or escalate.
It isn’t in the anonymous preview. The public try-a-workshop path has no session and no assistant, since the feature is tied to a real identified employee. An evaluator clicking through the preview won’t see it.
You can’t read the conversations. Covered above, and worth repeating here because it’s the item most likely to matter in a procurement review.
Frequently asked questions
Can an administrator read what employees ask Clay?
No. There’s no admin read path, conversations are stored against a random identifier rather than a name, and they’re deleted after 180 days. Admins see workshop answers and whether the assistant was used. If your policy requires the ability to review employee training conversations, this isn’t the right fit, and that’s a real tradeoff rather than a gap we intend to close.
What stops an employee talking Clay into something?
Several controls rather than one filter. The browser can’t feed Clay lesson content, the chat history isn’t treated as trustworthy, and admin text can’t impersonate Kinds policy. A standing set of attacks runs against every release. No system is unbreakable, which is why the worst case is the question box going quiet rather than the workshop breaking or the assistant going off-script.
Can we restrict what Clay says to our employees?
Yes, within limits. Account and organization guidance both feed into every answer, so you can tell it about your environment, your tooling, and your internal processes. What you can’t do is use that guidance to override Kinds policy, including the privacy promise made to employees. Guidance shapes answers. It doesn’t rewrite the rules.
What data leaves our environment?
The question the employee types, plus their name, role, department, and company, and an identifier for the workshop and screen they’re on. No mailbox content, no console data, no directory data beyond the profile fields already synced for training.
Sources
The claims here about privacy, retention, prompt-injection defenses, browser trust boundaries, language handling, and release testing are drawn from Kinds Security’s own documentation:
Clay Security and Privacy Architecture
Data Retention and Deletion Policy
Kinds Security AI Safety and Evaluation Methodology
Workshop Content and Language Support Documentation
Related reading
Putting a chatbot in front of employees is the kind of thing a security team says no to, usually for good reasons. The reasons are mostly about what it can see, what it can be talked into, and who can read it afterwards.
Clay’s answer to all three is the same: less than you’d expect, and deliberately so.
Related reading
What is prompt injection?
What is an insider threat?
Kinds Security glossary
Sources
Clay Security and Privacy Architecture
Data Retention and Deletion Policy
Kinds Security AI Safety and Evaluation Methodology
Workshop Content and Language Support Documentation
