Build Trusted Data for AI and Analytics in Eight Steps

In the rapidly evolving landscape of artificial intelligence and machine learning, few voices carry as much weight as Dominic Jainy. With a career spanning the intersection of blockchain, data science, and high-level IT infrastructure, Jainy has witnessed firsthand how the “garbage in, garbage out” mantra has transformed from a simple warning into a systemic risk for global enterprises. Today, the stakes for data integrity have shifted from mere reporting errors to the potential for autonomous AI agents to make catastrophic business decisions based on biased or incomplete information. We sit down with him to explore the shift from reactive data cleansing to the proactive construction of a trusted data foundation—a necessity for any organization aiming to leverage complex analytics without falling prey to operational failure.

The following discussion explores the strategic framework required to build trust in data, emphasizing the transition from periodic maintenance to a continuous, governance-led lifecycle. We delve into the nuances of departmental quality thresholds, the critical role of data stewards as guardians of information, and the way AI can be harnessed to monitor and safeguard the very systems it feeds upon. Jainy outlines how organizations must move beyond addressing symptoms to tackle the root causes of data decay, while fostering a culture where every employee feels the weight of responsibility for the information they create and consume.

Data quality dimensions like accuracy and timeliness often vary by department. How do you reconcile these conflicting thresholds?

Reconciling these differences is less about finding a single middle ground and more about establishing a documented, shared understanding of what “trusted data” means for a specific use case. In a process coordinated by data leaders, business stakeholders must sit down and define the required quality characteristics—such as accuracy, completeness, consistency, timeliness, validity, and uniqueness—based on the actual outcome they are trying to achieve. For instance, a financial reporting team might demand 100% accuracy for regulatory compliance, feeling the heavy pressure of legal scrutiny, while a customer analytics team might prioritize timeliness to catch a fleeting market trend. We reconcile these by documenting specific business rules and measurable quality standards for individual data assets, ensuring that a dataset used for AI model training meets a different, perhaps more rigorous, threshold than one used for a generic weekly internal update. By creating these clear, measurable benchmarks, we provide a foundation for continuous monitoring that respects the unique needs of different departments while maintaining enterprise-level integrity.

How is the rise of AI changing the stakes for data quality compared to the traditional era of business reporting?

In the past, poor-quality data certainly caused operational headaches and some head-scratching during quarterly reviews, but AI has amplified these issues into something far more volatile. AI systems take existing data quality problems and magnify them, producing unreliable recommendations, biased outcomes, and inaccurate predictions that can lead to disastrous choices by business executives. We are now seeing the rise of autonomous AI agents that act on data without human intervention, which means a single duplicate entry or a standardized value error can trigger a chain reaction of bad decisions. Because organizations are investing so heavily in AI applications, improving data quality is no longer a “nice to have” or a periodic chore; it is a fundamental requirement for the AI to function safely. If the foundation is cracked, the AI will not just lean—it will collapse, taking the organization’s competitive edge and financial performance down with it.

You have mentioned that the most cost-effective strategy is prevention at the source. How can organizations practically stop errors from entering their systems?

The most effective way to handle a data error is to ensure it never exists in the first place, which requires implementing validation controls right at the point of data capture. We should be using automated systems to ensure that all required fields are completed and that data formats are standardized before they ever hit the database. It is incredibly satisfying to see a system block a duplicate entry or validate incoming information against master data in real-time, preventing the “data smog” that usually clogs downstream analytics. Interestingly, we can now use AI to strengthen these preventive controls by having it detect anomalies or suggest missing information as a user types. While AI agents can even be configured to act autonomously in correcting these entries, we always maintain human oversight through business data stewards to ensure these automated “fixes” align with our broader business rules.

Why is it essential to focus on root cause analysis rather than just fixing individual data errors as they appear?

Correcting an individual record might make a specific report look better today, but without root cause analysis, you are essentially just treading water while the tide continues to rise. We must investigate why recurring issues happen—whether they stem from weaknesses in business processes, confusing data definitions, or perhaps a lack of user training on a specific system. Data stewards are perfectly positioned to act as detectives here, looking across organizational boundaries to see if a system integration is mangling a specific field every Tuesday. By addressing these systemic weaknesses, we reduce the ongoing costs of data management and slowly build a culture of trust where users don’t have to second-guess every number they see. AI can assist in this detective work by uncovering recurring patterns in quality incidents, but the final corrective action always requires a human touch to ensure we aren’t accidentally breaking a regulated or business-critical process.

Metadata is often viewed as a dry, technical requirement, but you argue it is vital for trust. How does metadata provide the context needed for AI?

Trusted data for AI is entirely dependent on consistency, and metadata is the “map” that ensures everyone—and every machine—is reading the data the same way. Without standardized metadata, including common definitions and naming conventions, organizations risk inconsistent reporting and flawed analytics where two different AI models might interpret the same field in conflicting ways. High-quality metadata provides the essential context of data lineage, ownership, and sensitivity classifications, which tells an AI user or a developer if a specific dataset is even appropriate for their application. It acts as a safety manual, documenting approved uses and quality rules so that when an AI model processes a piece of information, it understands the history and the “rules of engagement” for that data. Investing in data quality while ignoring metadata is like buying a high-performance engine but refusing to look at the dashboard; you might be moving fast, but you have no idea if you’re about to overheat.

What specific role do data stewards play in becoming the operational “guardians” of an organization’s data?

Data stewards are the bridge between the technical storage of data and the actual business utility, serving as the frontline protectors of data integrity. While technical stewards handle how data is accessed and delivered, business data stewards are responsible for translating high-level quality expectations into operational policies that people actually follow. In the age of AI, these stewards are now collaborating closely with governance teams to monitor how data changes might affect the performance and reliability of AI models. They are the ones who identify deteriorating quality before it disrupts a major project, coordinating remediation efforts and engaging with data owners to ensure the data remains fit for its intended use. An organization that fully empowers its stewards is one that can sustain reliable AI over time, because it has humans in the loop who understand both the data’s origin and its ultimate business impact.

How can a company transition from seeing data quality as a “technical problem” to building a true culture of shared responsibility?

Building a culture of responsibility starts with the realization that everyone who creates, manages, or uses data is a stakeholder in its quality. We need to integrate data quality into the very fabric of the organization through clear policies, performance metrics, and, most importantly, comprehensive training. When an employee understands how a small error in a source system can lead to a massive failure in an AI-driven forecast, the weight of that responsibility becomes tangible. Leadership must reinforce that data integrity isn’t just a task for the IT department; it is a shared enterprise value that directly impacts financial performance and competitive standing. By fostering this sense of accountability, organizations move away from a “clean-up” mentality and toward a state where high-quality data is produced as a natural byproduct of every business process.

What is your forecast for the future of AI-driven data governance?

I expect to see a significant shift toward “agentic” data governance, where AI agents are not just monitoring data quality but proactively managing the entire lifecycle with minimal human intervention. We will likely move toward a federated governance model where automated quality rules are embedded into every layer of the modern data stack, catching errors with a speed and precision that human stewards could never match. However, this will also create a “transparency gap” that will force organizations to invest even more heavily in data lineage and metadata documentation to understand how these AI agents are making decisions. Ultimately, the successful companies of the next decade will be those that treat data not as a static resource, but as a living, breathing asset that requires constant, automated, and human-led stewardship to remain a reliable foundation for intelligence.

Explore more

Modern ERP Systems Evolve from Documentation to Execution

When a sudden regional power outage strikes a major semiconductor manufacturing hub, the modern enterprise no longer waits for a manual report to trickle through various management tiers; instead, an intelligent system detects the disruption in real-time and immediately begins reallocating existing stock to high-priority orders. This transition marks a departure from the traditional role of Enterprise Resource Planning software,

Will AI Turn Windows Into a Monthly Subscription?

The traditional concept of a computer operating system as a one-time purchase is rapidly dissolving as Microsoft steers Windows 11 toward a persistent service-based architecture. For decades, the standard experience for any personal computer user involved purchasing a hardware device pre-installed with a permanent software license that remained functional for the entire lifespan of the machine. However, the rise of

How Do You Shift From Prompts to Agent Architecture?

The unprecedented acceleration of artificial intelligence integration within enterprise software development has now reached a pivotal inflection point where ad-hoc prompting is no longer sufficient for complex production systems. While the early months of adoption were characterized by individual developers finding clever ways to optimize small tasks, the current requirements of the industry demand a much more structured and predictable

Is Minnesota’s Water Safe After a Coordinated Cyberattack?

The recent identification of a sophisticated cyber intrusion targeting municipal water treatment facilities across Minnesota has raised significant alarms regarding the resilience of critical infrastructure in the Midwest. As state officials and cybersecurity experts analyze the aftermath of this coordinated event, the primary concern remains whether the safety and integrity of the public water supply were ever truly compromised. Initial

How Is AI Redefining B2B Go-To-Market Intelligence?

Modern revenue teams are no longer satisfied with the static spreadsheets and fragmented data silos that defined the previous decade of business-to-business sales operations. The emergence of sophisticated artificial intelligence has fundamentally shifted the baseline for market intelligence from reactive reporting to proactive, real-time strategic execution. Companies that once struggled to identify their most profitable segments are now leveraging neural