Data‑First Security Strategies for Enterprise AI
Emerj Artificial Intelligence Research has published a sponsored analysis titled "Data-First Security Strategies for Enterprise AI," aggregating interviews with practitioners from Google Cloud, Citi…
Shane Barrett·updated August 02, 2026

Emerj Artificial Intelligence Research has published a sponsored analysis titled "Data-First Security Strategies for Enterprise AI," aggregating interviews with practitioners from Google Cloud, Citi, Veeam, and Securiti into a four-part framework on data governance for enterprise AI deployments. The piece cites figures from Stanford HAI, the U.S. GAO, and the Cloud Security Alliance indicating that organizational visibility into unstructured data remains structurally incomplete, with direct consequences for any ML pipeline that ingests, transforms, or vectorizes that content downstream. For systems and ML engineering teams, the framing repositions a governance question as a data-pipeline architecture problem: control, it argues, can only be exerted before ingestion.
Empirical baseline
The cited figures establish the empirical baseline. Stanford HAI is reported as finding that 88% of organizations now use AI in at least one business function. Documented AI incidents are reported to have reached 362 in 2025, up from 233 the prior year. The GAO is cited as concluding that AI use in financial services introduces data quality, privacy, and cybersecurity risks that regulators are actively examining. The Cloud Security Alliance is reported as finding that only 35% of organizations maintain full visibility into where unstructured data resides, that 9% operate real-time scanning capabilities, and that 23% cannot scan unstructured data for risks at all.
The architectural hypothesis
Chris Joynt, Director of Product Marketing at Securiti, advances a specific architectural claim: once unstructured data is ingested, transformed, or vectorized, organizations lose meaningful visibility into how it is being used. Original forms become obscured, derivative copies proliferate, and lineage back to source is reported to break. Joynt cites customer environments operating across more than 200,000 data systems, generating billions of files and producing a petabyte of logs per day. At that scale, the claim is that any governance layer placed downstream of the model is structurally too late; pre-ingestion visibility becomes the only intervention point where sensitivity classification, access control, and lineage can be applied without re-engineering the model itself.
Limitations of the evidence base
The analysis is sponsored content produced in alignment with Emerj's sponsored guidelines and features vendors whose products address the visibility gap described. Survey methodologies underlying the Stanford HAI and Cloud Security Alliance figures are not reproduced in the piece, nor are confidence intervals, sample composition, or the operational definition of "unstructured data" disclosed. The 88% adoption figure and the 362-incident count should therefore be treated as inputs to a hypothesis rather than ground truth. The central architectural claim — that pre-ingestion is the only viable control point — remains untested against published benchmarks comparing pre-ingestion governance with post-hoc retrieval-side controls, retrieval-augmented generation audit trails, or model-level output filtering. No ablation is provided.