Last reviewed: September 29, 2026
This page explains the process Safely Clever follows when we review the safety of an AI tool, such as a chatbot, an AI assistant built into an app, or an AI agent. It turns the principles in our editorial policy into concrete steps, so you can judge our reviews by how they were made. We published this methodology in September 2026 and will update this page as our process develops. When we change it, we’ll note the change here.
What we assess
Every review answers the same core questions, taken from the scorecard in our AI tool safety reviews guide. We report each dimension separately instead of giving a tool a single “safe” or “unsafe” badge.
| Dimension | What we look at |
|---|---|
| Data and training | Whether chats are used for training by default, how to opt out, how long data is kept and how to delete it |
| Content safety | How clear the usage policies are, how consistently harmful requests are handled and whether harmless requests are over-refused |
| Accuracy | Whether the tool cites sources, whether those citations hold up and whether it admits uncertainty |
| Security | Two-factor authentication, and admin, SSO and audit features for teams |
| Transparency | Whether system cards, policies and change logs are published and readable |
| Agents and integrations | Whether permissions are scoped, risky actions confirmed and prompt injection risks addressed |
| Younger users | Age rules, teen protections and parental controls |
Our review process, step by step
- Choose what to review. We prioritize tools that many people use and the questions our readers ask. Companies can’t pay to be reviewed or to get a better review.
- Define the scope. Before testing we record the exact product, the plan (for example free or paid, personal or business), the platform (web, desktop or mobile app) and the date. Safety settings often differ between plans and platforms, so a review only covers what we actually tested.
- Read the official documentation. We read the provider’s privacy policy, terms of use, usage policies and help-center articles, and, where they exist, system cards and transparency reports. We note what the company says it does and link to those pages in the review.
- Check the settings hands-on. Using ordinary accounts, we walk through the real settings screens and common scenarios. These include finding and changing the training setting, starting a temporary chat, viewing and deleting memory, exporting and deleting data, connecting an integration and then revoking it, and turning on two-factor authentication. We take notes and screenshots as we go.
- Try realistic prompts. We ask questions with answers we can check, to see whether sources and citations hold up. We also try a small set of borderline but legitimate requests, such as medical or security questions, to see how the tool balances refusing harmful requests against being useful.
- Compare claims with what we observed. When the documentation and the product disagree, or a setting is hard to find or unclear, we say so in the review.
- Write it up with dates. Each review lists its findings dimension by dimension, states the date checked and the plan and platform tested, and links to its sources.
- Editorial review. Before publication, the Safely Clever editorial team checks the draft’s claims against the cited sources and our notes.
- Update and correct. AI products change often. We recheck reviews when significant changes are announced and show a “Last reviewed” date. Factual corrections are noted as described in our editorial policy.
What we don’t do
- We don’t accept payment, free upgrades or other incentives in exchange for coverage or ratings. We currently run no advertising or affiliate links. If that changes, we’ll disclose it on the pages concerned.
- We don’t use real personal data in tests. We use test content that we’re comfortable sharing with the provider.
- We don’t hack or attack services. Our checks use the product as an ordinary user would. We don’t try to break into systems, bypass security controls or run large-scale jailbreak campaigns, and we don’t publish instructions for misusing a tool.
- We don’t claim more than we tested. We can’t see inside a company’s servers, so statements about how data is stored or used internally are based on the company’s own documentation, and we attribute them that way.
Limits of our testing
A review is a dated snapshot of one product, on one plan and platform, at one point in time. Features roll out gradually and can vary by country, account type and app version, so your experience may differ. Our prompt checks are a small sample, not a benchmark. For formal security assurance, look for independent audits and certifications, or commission your own testing.
How we use AI in our work
As our editorial policy explains, AI tools may help us research, summarize long documents and draft. Every published review is checked by a human against the sources and our test notes, and we never use AI to invent findings, quotes or test results.
Who writes our reviews
Our guides and reviews are written by Ellis Marlow and reviewed by the Safely Clever editorial team. You can read more about Safely Clever.
Sources and frameworks we rely on
Alongside each provider’s own documentation, we use these public references to structure our checks:
- NIST AI Risk Management Framework (NIST)
- OWASP Top 10 for Large Language Model Applications (OWASP GenAI Security Project), for prompt injection, excessive agency and other risks in agents and integrations
- Provider documentation, for example Data Controls in ChatGPT (OpenAI), the Anthropic Privacy Center, the Gemini Apps Privacy Hub (Google) and Microsoft Copilot privacy controls
Spotted something out of date?
If a setting has moved or a review no longer matches what you see, please tell us through our contact page. Include the page, what you saw and the date.

