Shipping a large language model (LLM) feature is easy. Shipping one that stays safe when real users, messy data and motivated attackers show up is harder. Models misread instructions, invent facts, leak information they were given, and can be steered by text hidden in documents or web pages. Once an LLM can call tools, such as sending email, querying databases, running code or spending money, those failures become actions.
Guardrails are the technical and operational controls that keep an AI system within safe bounds while keeping it useful. This guide covers the full picture for developers: the main threats, the design principles that matter most, guardrails at each layer of an LLM application, evaluation and red teaming, agent-specific risks, and the frameworks that can structure your program. It is the hub for our Responsible AI for Developers section.
Last reviewed: September 29, 2026.
Start with the threat model
Before choosing tools, write down what could go wrong. For each feature, ask:
- What can the model see? User input, system prompts, retrieved documents, tool results, other users’ data.
- What can the model do? Which tools, APIs and permissions it has, and what the worst possible call would be.
- Who might attack it? Curious users, abusive users, competitors, or attackers who never touch your UI but plant content your model will later read.
- What does failure cost? Embarrassment, wrong advice, data leakage, financial loss, legal exposure, physical harm.
The OWASP Top 10 for Large Language Model Applications is a good checklist of common failure classes, including prompt injection, sensitive information disclosure, supply chain risks, data and model poisoning, improper output handling, excessive agency, system prompt leakage, weaknesses in vector stores and embeddings, misinformation, and unbounded resource consumption. MITRE ATLAS catalogs real-world adversary techniques against AI systems.
Five design principles that matter most
- Treat the model as an untrusted component. Its output is influenced by every piece of text it reads. Do not let it make security decisions on its own.
- Least privilege. Give the model, and the credentials it uses, only the minimum access needed for the task. Prefer read-only, narrowly scoped tools.
- Defense in depth. No single filter or prompt will stop every failure. Layer controls so that one miss is caught by the next.
- Deterministic checks where possible. Validate formats, allow-lists, amounts and destinations in regular code, not by asking the model to check itself.
- Humans for high-stakes steps. Require explicit confirmation for irreversible or consequential actions, and make the confirmation meaningful by showing exactly what will happen.
Guardrails layer by layer
Model and provider choice
Pick a model whose documented capabilities, safety behavior and data terms fit your use case. Read the provider’s system card and usage policies, understand data retention for API traffic, and pin model versions so behavior doesn’t change under you without testing.
System prompt and application design
Write clear instructions about scope, tone and refusals, but do not rely on the system prompt for security. Assume it can be extracted, and never put secrets, credentials or sensitive business logic in it. Narrow, well-defined tasks are far easier to secure than a general-purpose assistant with broad access.
Input handling
- Screen inputs for abuse and policy violations with moderation classifiers where appropriate.
- Enforce length, rate and cost limits to prevent resource exhaustion.
- Mark untrusted content clearly, for example with delimiters, and tell the model not to follow instructions found inside it. This helps but is not a reliable defense on its own.
Prompt injection and retrieval hygiene
Prompt injection is the defining security problem of LLM applications. Direct injection comes from the user; indirect injection hides instructions in content the model processes: web pages, emails, PDFs, tickets, code comments or retrieved documents. There is currently no complete fix, so design to limit the damage:
- Keep untrusted content away from privileged tools. A model that has just read an arbitrary web page should not be able to send email without a human check.
- Filter and attribute retrieved content. Track where each chunk came from and restrict retrieval to sources the current user is allowed to access.
- Watch for data exfiltration paths such as rendered images or links that encode data in URLs, and disable or sanitize them.
Our practical guardrails checklist goes into more detail on these controls.
Tool and agent permissions
- Give each tool the narrowest scope possible, with separate credentials per tool and per user.
- Enforce authorization in your backend, not in the prompt. The model should never be able to access data the end user couldn’t access directly.
- Allow-list destinations, commands and parameters. Cap amounts, counts and frequency.
- Run code execution in sandboxes with no network access or tightly limited access.
- Require human approval for sending, deleting, publishing, paying and changing permissions.
Output validation
- Use structured outputs, such as JSON with a schema, and validate them strictly before acting on them.
- Treat model output as untrusted when it flows into HTML, SQL, shell commands or other interpreters. Escape and parameterize it exactly as you would user input, to avoid cross-site scripting, injection and similar bugs.
- For factual features, ground answers in retrieved sources, show citations and check that cited passages actually support the answer.
- Screen outputs for sensitive data, such as personal information or secrets, before display or logging.
Monitoring and incident response
Log prompts, tool calls and outputs with appropriate privacy protections. Alert on unusual patterns such as spikes in refusals, tool errors, cost or data access. Give users an easy way to report problems, and have a runbook for disabling features, rotating credentials and notifying affected people when something goes wrong.
Evaluations and red teaming
Guardrails are only as good as your evidence that they work.
- Build an eval set that reflects real usage: typical requests, edge cases, known failure modes and policy-sensitive topics. Include expected behavior for each.
- Add adversarial tests for jailbreaks, direct and indirect prompt injection, data leakage and harmful-content requests relevant to your domain.
- Run evals on every change, whether prompt edits, model upgrades, new tools or retrieval changes, and treat regressions like failing unit tests.
- Be careful with LLM-as-judge. Using a model to grade outputs scales well but inherits model biases. Spot-check with humans and calibrate against labeled examples.
- Red team before launch with people who did not build the feature, and repeat periodically. Include attackers’ goals, not just forbidden words.
Special considerations for AI agents
Agents that plan and take multi-step actions multiply risk. Each step can compound an earlier error, and long tasks give injected instructions more chances to take effect. Keep agents on short leashes: limit the number of steps and tool calls, checkpoint progress for human review, prefer reversible actions, and design for “safe failure”, where the agent stops and asks rather than guessing. Monitor not just final outputs but the sequence of actions.
Frameworks that can structure your program
You do not need to invent governance from scratch. Useful references include:
- NIST AI Risk Management Framework (AI RMF 1.0). A voluntary framework released in January 2023, organized around four functions: Govern, Map, Measure and Manage. NIST also published a Generative AI Profile (NIST AI 600-1) in July 2024. See the NIST AI RMF page.
- ISO/IEC 42001:2023. An international standard for AI management systems that organizations can certify against.
- OWASP resources for LLM and agent security testing.
- The EU AI Act, which places legal obligations on providers and deployers of certain AI systems, including transparency duties and requirements for high-risk uses. Our AI risk and policy guide explains the timeline.
A launch-readiness checklist
- Threat model written and reviewed for each AI feature.
- No secrets in prompts; system prompt assumed public.
- Tools scoped to least privilege with backend authorization.
- Human confirmation for consequential actions.
- Structured outputs validated; model output escaped before rendering or execution.
- Eval suite, including adversarial cases, passing on the release candidate.
- Logging, alerting, abuse reporting and an incident runbook in place.
- User-facing disclosure that AI is involved and that outputs may be wrong.
Guardrails should preserve usefulness
Over-restrictive systems push users toward riskier workarounds. The goal is not to make the model refuse more; it is to make failures rare, contained and recoverable. For background on why models fail in the first place, see AI safety and alignment explained. To see how consumer tools implement their own safeguards, read our AI tool safety review framework, and to follow how frontier labs approach the same problems at scale, see how AI labs approach safety.
Frequently asked questions
What are LLM guardrails?
LLM guardrails are the technical and operational controls around a language model that keep an application safe and reliable. They include input and output filtering, permission limits on tools, output validation, human approval steps, evaluations and monitoring.
Can prompt injection be fully prevented?
Not with current techniques. Filters and careful prompting reduce it, but the reliable approach is to limit what an injected instruction could achieve: keep untrusted content away from privileged tools, enforce authorization outside the model, and require human confirmation for consequential actions.
Is a good system prompt enough to make an AI app safe?
No. System prompts guide behavior but can be overridden or extracted. Security controls such as authorization, validation and permission limits must be enforced in your application code.
How do I test an LLM application for safety?
Build an evaluation set covering normal use, edge cases and adversarial attacks such as jailbreaks and prompt injection, run it on every change, and supplement automated tests with human red teaming before launch.
Which framework should a small team use for responsible AI?
Many teams start with the OWASP Top 10 for LLM Applications for security and the NIST AI Risk Management Framework for governance, then add ISO/IEC 42001 or EU AI Act requirements if their customers or markets require them.
Sources
Primary sources used for this guide, checked on September 29, 2026. Policies and products change, so always check the latest version at the source.
- OWASP Top 10 for Large Language Model Applications, OWASP GenAI Security Project
- MITRE ATLAS, adversarial threat landscape for AI systems
- AI Risk Management Framework (AI RMF 1.0), NIST, January 2023
- Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile (NIST AI 600-1), NIST, July 2024 (PDF)
- ISO/IEC 42001:2023: AI management systems, ISO
- Regulation (EU) 2024/1689 (the AI Act), EUR-Lex, official text
Latest developer guides
- AI Guardrails for Developers: A Practical Checklist for Safer AI AppsA practical AI guardrails checklist for developers: threat modeling, prompt injection defenses, least privilege for tools, output validation, evaluation and monitoring.

