All insights
Agent Security

AI Agent Security Architecture: Governance Across Data, Models, and Agents

September 2026Cybersecurity7 min readExecutive Brief

AI agents represent a paradigm shift in enterprise security. Beyond answering queries, agentic systems reason, call tools, trigger workflows, generate code, and alter configurations across enterprise boundaries at high speed.

Because their actions combine reasoning, data access, and execution, traditional runtime security models are insufficient. Governing agentic AI requires an integrated, defense-in-depth framework that anchors controls where enterprise data, context, and identities reside.

Key takeaways
  • 01
    The Shift to Autonomous RiskEnterprise security teams have moved from asking “How do we grant AI access?” to “How do we ensure autonomous agents execute tasks safely and aligned with business intent?”
  • 02
    The Data-Model-Agent FrameworkWe structure agentic security across three core layers: Data Governance, Model Protection, and Agent Identity & Tool Control.
  • 03
    Prompt Injection as a Primary VectorIndirect prompt injection—where agents ingest malicious instructions embedded in external files, tickets, or web pages—requires dedicated guardrails between model intent and execution.
  • 04
    Identity Attribution & Tool GatewaysUniquely identifying agents and mediating tool invocations via Model Context Protocol (MCP) gateways prevents “confused deputy” attacks and shadow AI proliferation.

Why Agentic AI Demands a New Security Model

Traditional Security Orchestration and Access Controls rely on deterministic human action. Agents introduce non-deterministic execution paths where a single user input or event can trigger multiple downstream operational pathways:

  1. 01
    Reading and querying internal data layers.
  2. 02
    Interpreting and reasoning via the core LLM.
  3. 03
    Invoking external tools and APIs via Model Context Protocol (MCP).
  4. 04
    Generating and executing system code.
  5. 05
    Proposing and enforcing operational changes.

Critical Enterprise Control Questions

Putting AI agents into production safely requires answering six fundamental operational requirements:

  • Attribution. Can system logs explicitly distinguish agent actions from human actions?
  • Least Privilege. Can agents be bounded strictly to the specific data and tools required for their sub-task?
  • Data Loss Prevention. Can sensitive information be prevented from crossing approved jurisdictional or network boundaries?
  • Adversarial Resilience. Can the system withstand both direct and indirect prompt injection attacks?
  • Human-in-the-Loop. Can high-risk, irreversible actions mandate multi-party human approval?
  • Auditability. Is every intermediate reasoning step, tool call, and state mutation recorded in a tamper-proof trace?
Agent Security

The Data-Model-Agent Framework

Securing agentic workflows requires a unified strategy spanning three operational layers:

Data-Model-Agent framework
Agent governance layer — action & execution
  • Unique machine identities
  • Centralized MCP tool gateways
  • Multi-party approvals
  • Sandboxed code execution
  • Auditable traces
Model security layer — intelligence engine
  • Direct & indirect injection guardrails
  • Private network execution
  • Input/output guardrails
  • Model context isolation
Data foundation layer — foundation & storage
  • Role-based access control
  • Dynamic data masking
  • Zero-copy architecture
  • Regional data sovereignty
  • Encryption

Deep-Dive: Layer-by-Layer Security Architecture

1. Protect the Data: Zero-Copy & Least Privilege (Foundation). Data governance forms the baseline. AI systems expose underlying data weaknesses at scale. At Orbis, we enforce strict least-privilege policies and dynamic data masking before information reaches model contexts. Utilizing a zero-copy architecture minimizes data sprawl, preserves regional sovereignty, and ensures security policies remain bound directly to the source data rather than duplicated across secondary systems.

2. Secure the Model: Defend Against Prompt Manipulation (Intelligence Engine). Agents are vulnerable to direct prompt injection (users attempting to bypass system constraints) and indirect prompt injection (hidden malicious instructions within retrieved PDFs, web pages, or tickets). Orbis integrates advanced prompt injection guardrails directly between user intent, model reasoning, and execution, ensuring execution stays within secure enterprise boundaries without external data leakage.

3. Govern the Agent: Identity, Tools, and Isolation (Execution & Runtime). When agents utilize tools, they become autonomous actors. Orbis assigns agents distinct, auditable identities so that every query, API invocation, and tool call is attributable. Tool usage is routed through centralized Model Context Protocol (MCP) gateways, granting granular control over thousands of tools and preventing shadow AI deployments. Code-generating agents execute within isolated, sandboxed environments to restrict filesystem and network access by design.

4. Continuous Control: Posture Management & Resilience (Day 2 Operations). Post-deployment security requires continuous monitoring and rapid recovery capabilities. Orbis integrates AI Security Posture Management (AI-SPM) to detect anomalous data movement, operationalize compliance reporting, and enforce multi-party approval for high-stakes actions. System state resilience is backed by Write-Once-Read-Many (WORM) backups, point-in-time recovery, and cross-region replication.

Enterprise Production Readiness Checklist

Before moving agentic workflows from prototype to production, verify that the following controls are operational within your environment:

  • Identity & logging. Agents operate under unique machine credentials; human and agent actions are distinctly tagged in audit logs.
  • MCP gateway enforceability. All agent-to-tool connections route through a centralized gateway with least-privilege access rules.
  • Prompt injection defense. Active guardrails scan all incoming text and unverified third-party documents for indirect prompt injections.
  • Sandboxed runtime. Code-execution agents run in isolated containers with blocked outward network access and restricted storage paths.
  • Multi-party sign-off. High-risk operational commands (e.g., system configuration changes, financial transfers, data deletion) mandate explicit human approval.
  • Resilience verification. Automated recovery pipelines (including WORM storage and point-in-time rollback) are tested against agent error scenarios.