How Should Organizations Govern Autonomous AI Agents?

In an era where artificial intelligence has moved beyond mere chatbots to become active participants in business logic, the stakes for governance have never been higher. As we navigate the complexities of 2026, organizations are no longer just asking if a model can generate a coherent paragraph, but whether an autonomous agent can safely execute a million-dollar transaction or manage a pharmaceutical supply chain without human intervention. Marinela Profi, the Global Market Strategy Lead for AI Agents and Generative AI at SAS, stands at the forefront of this architectural shift. With her extensive background in data strategy and AI impact, she bridges the gap between high-level policy and the hard engineering required to make autonomous systems reliable. In this conversation, she explores why traditional deployment strategies are failing, the hidden “trust debt” accumulating in enterprise architectures, and how the most successful companies are using governance as a high-performance engine for growth rather than a restrictive set of brakes.

When software shifts from simply displaying information on a screen to independently triggering backend workflows, how does the definition of a “bad answer” change for an enterprise?

In the past, a “bad answer” was something a human could catch, critique, and discard before it left their terminal. Today, the definition has fundamentally shifted because the “blast radius” of a single error has expanded beyond the user interface. When an autonomous agent makes a mistake, it isn’t just a typo or a factual error; it is a triggered workflow, an interaction with a legacy system, or an executed action that occurs before a human even realizes something is wrong. We are seeing a shift where an error propagates through the system, potentially creating a domino effect across interconnected platforms. This means that “bad” no longer describes the quality of the content, but the impact of the consequence, turning a linguistic failure into an operational crisis.

Why does the classic Silicon Valley playbook of “moving fast and breaking things” fail so spectacularly when applied to autonomous software systems that make automated decisions?

The traditional playbook relies on the idea that the cost of failure is low—if a feature is bad, you observe user behavior, learn, and iterate in the next sprint. However, when software has the authority to act on its own, that iteration cycle is often too slow to prevent real-world damage. You cannot simply “fail fast” when an agent has the permission to alter financial records or change a patient’s treatment protocol without an immediate human checkpoint. According to our latest research, while 89% of agents in production are already taking actions rather than just assisting, more than half are operating with limited or no human approval. This is why we have to redefine what moving fast means; in 2026, speed must include the built-in ability to observe, constrain, interrupt, and recover in real-time.

Your research highlights a noticeable gap in trust between standard generative AI and these newer agentic systems. What is driving that 10% drop in confidence among users?

It is a fascinating psychological and technical divide: trust in generative AI currently sits at 76%, but it drops to 66% the moment we discuss agentic AI. This 10% gap is fueled by the loss of direct oversight and the inherent unpredictability of probabilistic systems making deterministic choices. People are naturally more anxious when the machine moves from “telling” to “doing,” especially when they feel the guardrails are invisible or non-existent. We have found that the anxiety is justified because while many organizations are deploying these agents, very few have the infrastructure to explain why a specific action was taken after the fact. This lack of transparency creates a “black box” of action that makes stakeholders hesitant to fully commit to the technology’s potential.

How do the architectural shortcuts taken during a pilot project eventually turn into massive compliance liabilities and technical debt for a company?

The shortcuts that feel inexpensive during a 2026 pilot—like failing to create a reliable lineage between data and decisions—become a nightmare when a regulator or auditor knocks on your door two years from now. If your architecture cannot reconstruct exactly which tool an agent invoked, what data it used, and what policy governed that specific moment, you have accumulated what I call “trust debt.” You might find yourself in a position where you cannot explain a decision to a customer or an auditor, and at that point, no amount of new policy-writing can fix the underlying structural failure. Only 17.5% of organizations currently report having a fully optimized data infrastructure that includes the validation and explainability required for these systems. Without that foundation, the debt eventually becomes so high that the organization is forced to halt its AI expansion entirely.

Why is a successful proof of concept often a poor indicator of whether an autonomous agent is actually ready for the pressures of enterprise production?

A proof of concept is designed to see if the AI can perform a task, which usually happens in a sanitized sandbox with predictable data and controlled permissions. Enterprise production, by contrast, is a chaotic environment where data is stale, permissions are fragmented, and tools can fail in ways the developer never anticipated. We have to move away from simple model testing and toward holistic system testing because a demo that looks impressive in a conference room doesn’t prove operational readiness. In the real world, users will challenge the AI not just to get an answer, but to probe its reasoning and push its boundaries. If the system hasn’t been tested against those stressors, it will fail the moment it encounters the complexity of a live backend environment.

There is a common design pattern of placing “one intelligent layer” over an entire enterprise. Why do you view this as a potential recipe for catastrophic failure?

On an executive roadmap, the idea of an LLM or an agent orchestrating everything across the enterprise looks incredibly efficient and visionary. However, intelligence without strict boundaries is a massive liability because a probabilistic model should never be the control plane for a deterministic business. If the agent doesn’t know exactly what it is authorized to touch and which systems it is allowed to affect, it will eventually hallucinate a path that leads to a security breach or a data integrity issue. I often point to a global bank we work with where the CIO correctly identified that technology alone isn’t the solution; the LLM must operate inside a rigid architecture of data and control. The goal is to keep the “smart” part of the system contained within a “safe” part of the architecture that handles the actual decision logic.

How does the process of verifying a system change when we no longer have a human reviewing every single action an agent takes?

Verification is shifting from a pre-deployment checklist to a continuous, real-time process that asks if the entire decision-and-action chain is trustworthy. In the past, we focused on model performance metrics like bias and drift, but with autonomous agents, we have to verify the authority and appropriateness of the action itself. We need to know: Was the agent authorized for this? Did it stay within its defined boundaries? The data shows a stark divide here, as 66% of trustworthy AI organizations have formal validation processes, while only 15% of less mature companies do. It’s no longer just about whether the prediction was accurate; it’s about whether the subsequent action was the right thing to do in a specific business context.

Can you elaborate on how fragmented data and poor system quality act as a literal blindfold for these autonomous agents?

An agent is only as effective as the world it can perceive; if a customer’s information is scattered across six different legacy systems with conflicting definitions, the agent is essentially flying blind. We see this quite clearly in the life sciences sector, where researchers at one major biopharmaceutical company were spending 80% of their time just cleaning and reconciling data before any analysis could even begin. You cannot automate your way out of a broken data foundation, no matter how sophisticated your AI agents are. If the data lineage is unclear and the logic is fragmented, the agent will make decisions based on a distorted reality, leading to outcomes that are at best useless and at worst dangerous.

You’ve made the provocative claim that governance is actually “infrastructure for speed.” How does having strict controls actually help an organization move faster?

We have historically treated governance as a “no” department—a gatekeeper that slows down innovation with red tape and endless meetings. But if you think about it, the brakes on a car aren’t there just to make it go slow; they are there so you can drive fast safely. When policies are translated into executable, programmatic controls, engineering teams don’t have to debate every single deployment with the legal and security departments. Organizations with high trustworthiness scores are actually 15 times more likely to report a high ROI on their AI investments because they have a reusable operating system for innovation. By establishing clear thresholds for escalation and logging, these companies can delegate tasks to agents with total confidence, knowing the system will automatically flag anything that falls outside of its authority.

Where should engineering teams draw the line on machine autonomy, and what does a healthy “human-on-the-loop” architecture actually look like?

There isn’t a single universal line; the boundary has to move based on how much risk is involved and whether the action can be easily reversed. A low-risk task can be fully autonomous, but a high-consequence decision regarding credit, employment, or healthcare must have a human as the final arbiter. The mistake many make is putting a human in every loop, which turns that person into a bottleneck who eventually just “rubber stamps” everything without looking. A better architecture is “human-on-the-loop,” where humans set the authority, the boundaries, and the escalation criteria, while the machine handles the execution within those lines. We can and should delegate the heavy lifting of execution to AI, but we must never, under any circumstances, delegate the ultimate accountability.

What is your forecast for the evolution of AI governance over the next two years?

From 2026 to 2028, I expect we will see the total collapse of “documentation-only” governance as companies realize that a PDF policy cannot stop a malfunctioning agent. We will see a massive shift toward “Governance as Code,” where permissions and ethical boundaries are baked directly into the API calls and orchestration layers of the agentic ecosystem. The competitive landscape will split into two camps: those who are paralyzed by the risks of autonomy because they lack the controls to manage it, and the leaders who have built a robust “trust architecture” that allows them to scale agents across every department. Ultimately, the winners won’t be the ones who had the most powerful models, but the ones who had the most reliable systems to keep those models in check.

Explore more

Will Lower Payroll Taxes Help Solve Youth Unemployment?

The transition from a classroom desk to a professional office chair has become an increasingly treacherous journey for hundreds of thousands of young adults across the United Kingdom. With youth unemployment figures hovering around 750,000 individuals, the disconnect between academic achievement and labor market stability has reached a critical juncture. This roundup examines the proposed fiscal strategies aimed at reversing

Is Your Job Description Outdated in the Age of AI?

Ling-Yi Tsai is a titan in the HR technology landscape, possessing a wealth of experience in steering global organizations through the turbulent waters of digital transformation. As an expert in HR analytics and technology integration, she has spent years dismantling outdated legacy systems to make room for intelligent, data-driven talent management. Her perspective is particularly vital today as companies struggle

The Evolution of Revenue Operations via the CRO AI Framework

The effectiveness of an AI agent is fundamentally limited by the quality and connectivity of the data sources it is permitted to access. In the current enterprise sales environment of 2026, Chief Revenue Officers are no longer content with speculative pilot programs or isolated productivity tools that offer only marginal gains. Instead, there is a decisive move toward a comprehensive,

Cisco Transforms Webex with New Collaborative AI Agents

Digital collaboration has evolved from a utility for remote communication into a sophisticated ecosystem where artificial intelligence acts as a catalyst for deep organizational change. Cisco has recently pivoted the Webex platform toward “agentic” AI, signaling a fundamental move from simple, reactive chatbots to fully integrated digital colleagues. This transition is not merely a cosmetic update; it represents a strategic

Why Should HR Lead Your Agentic AI Talent Strategy?

When AI agents begin handling prospecting and data research, human representatives must be strategically redeployed into areas that prioritize relationship-building and strategic growth. This fundamental shift marks a departure from traditional automation, where software merely served as a static tool for human operators. In the current landscape of 2026, agentic Artificial Intelligence has evolved into a category of autonomous entities