Rasa Whitepaper Enterprise Grade Security.pdf

Enterprise-Grade Security: Structuring AI Agents for Control and Scale

Purpose

To help enterprise engineering, security, and product teams understand where AI agents introduce risk and how CALM provides structural safeguards against common failure modes in LLM-powered systems.

Audience

Security engineers, platform teams, compliance officers, AI/ML architects, and technical leaders responsible for building or evaluating conversational AI systems.

Introduction

Large language models expose new security and governance challenges, especially when agents operate without structured constraints. When AI agents rely on unstructured prompts to interpret inputs, determine actions, and produce responses, small variations in wording or context can trigger behavior the system was never designed to allow. The agent becomes harder to govern, harder to debug, and more likely to violate constraints that were never explicitly enforced.

For enterprise teams, securing the infrastructure is no longer enough. The agent itself must operate within a defined structure that enforces how it understands language, how it takes action, and how those actions are validated. Systems that blend reasoning, execution, and output into a single layer lose the ability to enforce policy, prove compliance, or recover from failure with confidence.

CALM (Conversational AI with Language Models) addresses these challenges with architectural separation. Interpretation, decision-making, and execution operate through distinct components, each with its own logic, validators, and audit trail. The model helps identify user intent, but task execution follows deterministic flows into something more straightforward like 'clearly defined steps that always follow the same rules/logic.'

This white paper examines the security architecture of CALM through the lens of real-world threats. Each section addresses a different risk (from prompt injection and hallucinations to state replay, adversarial input, and governance breakdown) and explains how CALM mitigates those risks.

Prompt Injection

Unlike traditional injection attacks, prompt injection targets the structure of language itself. It works by taking advantage of the model's natural tendency to treat instructions embedded in text as meaningful. When AI agents can trigger downstream actions, the potential for misuse is significantly increased. Mitigating prompt injection requires architectural control, not reactive filtering.

CALM addresses this by separating the model's role as a language interpreter from the deterministic flow engine that governs execution. The agent understands user input through the model, but decisions about what actions to take happen through versioned logic, schema validation, and gated transitions. This shifts security from heuristics to design, ensuring that adversarial inputs cannot override intent boundaries or trigger unintended behavior.

Hallucinations and Generative Drift

Hallucinations happen when an LLM produces plausible-sounding but incorrect information. In a regulated industry, hallucinations create severe problems. The more reliable solution is to limit where the model has room to improvise. With CALM, the LLM handles language interpretation but doesn't generate core content or decide actions. Execution always follows deterministic flows that are testable and tightly controlled.

Data Privacy and LLM Scope

LLMs process everything as raw text, which makes it harder to control what they see and how they use it. Understanding what the model sees, what leaves the system, and what gets stored helps maintain control over sensitive information. CALM keeps control in the hands of the enterprise by embedding decisions into the system structure.

Command Execution and Flow Control

When LLMs decide what to do and also carry it out, it's hard to know what will happen or why. CALM limits the agent's ability to act by enforcing rules defined in code. Execution follows deterministic flows, where each step occurs only when predefined conditions are met.

Monitoring, Observability, and Incident Readiness

LLM-powered agents introduce new operational risks. CALM makes observability possible by embedding it into the structure of the agent itself. Traceability is foundational to diagnose issues and enforce controls, ensuring that teams have full visibility into how the agent behaves.

Adversarial Prompt Design and Data Poisoning

An agent can be manipulated without being hacked, simply by feeding it misleading inputs. Preventing these attacks requires defensive design at multiple levels, ensuring that structured design matters to limit the attack's scope and impact.

Multimodal and Latent Threat Surfaces

As agents increasingly support voice inputs, risks introduced by longer input paths must be managed. Mitigating these risks requires guardrails at each layer of the input pipeline.

Governance, Standards, and Policy Enforcement

Enterprises need systems that enforce policy through constraints embedded directly in the agent's architecture.

Model Updates and Behavior Drift

Version pinning reduces exposure to risk as LLM behavior can change over time due to model updates. CALM narrows the model's role, isolating the impact of model changes and assuring stability across updates.

Continuous Evaluation and Red Teaming

Agents may drift, degrade, or behave inconsistently. Continuous evaluation combined with red teaming proactively tests agent resilience under changing conditions, exposing breakdowns before real-world deployment.

Overtrust and Human Factors

Preventing overtrust requires careful design and communication at the interface and system levels. CALM addresses the issue by structuring the distinction between what the model understands from what the agent is allowed to do, ensuring that every transition is governed by clear conditions.

Trust Boundaries: Internal Versus External Systems

CALM maintains strict constraints, ensuring that external systems can't trigger actions unless explicitly allowed by internal rules.

Replay Attacks and State Integrity

Replay attacks exploit the conversational nature of agents, relying on state integrity. CALM treats flow state as a first-class security control, revalidating inputs before they are used.

Safety by Design: Decomposed Agent Architecture

Decomposed architectures separate concerns, allowing each component to handle clear roles, which enhances testability and reduces risk.

Conclusion

Enterprise-grade security requires design decisions that withstand pressures of real-world applications. By defining what actions can occur and under what conditions, CALM provides a structured, predictable framework for secure AI agents.