AI adoption is pushing organizations to retain more data for longer and bring older information back into active use, according to an IDC survey sponsored by storage company WD.
The white paper, Built for Scale: The Enduring Role of HDDs in the AI Era, is based on a survey of 763 IT and business decision-makers across seven countries who had direct responsibility for AI infrastructure or data-storage decisions. IDC also conducted three qualitative interviews with senior storage-industry participants.
Among surveyed organizations, 94.7 percent said they were storing more data because of AI and generative AI adoption over the previous 12 months. Sixty-one percent reported data growth of at least 25 percent over that period, while 74 percent expected their data volumes to grow by at least 25 percent over the next three years.
AI data persists beyond individual compute cycles
The study’s central argument is that AI infrastructure planning cannot focus only on compute. Training, inference and other workloads create datasets, checkpoints, model logs, synthetic data and outputs that can remain useful after the original processing job is complete.
IDC found that 85.4 percent of surveyed organizations had seen growth in their data-lake volumes over the previous year. About 59.4 percent identified AI-generated information, including synthetic data, inference outputs and model logs, as the leading driver of that growth.
Nearly 95 percent said the value of their organization’s data had increased as a result of AI and generative AI adoption. That finding reflects respondent perceptions rather than a direct valuation of the underlying data, but it helps explain why companies may be more reluctant to discard information that could later be used for model development, retrieval or analysis.
Archived data is moving back into active use
The survey also found that 74.3 percent of respondents said AI had caused their organizations to retain data for longer. Another 75.9 percent reported bringing more archived or cold-tier data back online to support AI workloads.
IDC said 96 percent expected to need faster archive retrieval for AI inference and retrieval-augmented generation applications. That can change how companies think about the boundary between active and archival storage because information that was previously kept mainly for compliance or long-term retention may become an input to new AI systems.
The study found that 74.6 percent of enterprise data among surveyed organizations resided in warm, cool and cold storage tiers. More than 60 percent of data-lake volume was described as cold or infrequently accessed.
Those findings do not imply that every AI workload should use the same type of storage. Different parts of the data lifecycle have different performance, capacity and cost requirements, and the study’s sponsor has a commercial interest in emphasizing the continuing role of hard-disk drives.
Storage economics becomes part of AI infrastructure planning
About 98.2 percent of respondents considered total cost of ownership per terabyte important or very important when making storage decisions. As AI datasets grow, the cost of retaining information that may be accessed only occasionally can become a larger part of infrastructure planning.
WD CEO Irving Tan said the AI infrastructure conversation has focused heavily on compute even as organizations generate more data, keep it longer and find new uses for historical information.
That shift is relevant in Asia as companies and governments commit more capital to local AI infrastructure. TNGlobal has covered plans for large-scale AI factory capacity in Indonesia through the Zankore platform, one example of the broader buildout of computing resources across the region.
Compute remains the most visible part of many AI infrastructure projects because GPUs and accelerators account for a large share of upfront investment. Storage becomes a different kind of constraint as data accumulates over time and needs to remain accessible across training, fine-tuning, inference and retrieval workflows.
The study highlights a practical issue that can be overlooked when AI infrastructure is discussed mainly in terms of compute. Organizations that generate more information, retain it for longer and repeatedly reactivate historical datasets will need to plan for the cost and operational complexity of the full data lifecycle, not only the processing capacity required for the next model run.

