Artificial intelligence (AI) and machine learning (ML) promise a step change in the automation traditional to IT, with applications ranging from simple chatbots to nearly unimaginable levels of complexity, affirming technology and retaining control.
Storage forms a key part of AI, to create data for training and store the potentially large volumes of data generated, or throughout inference when the implications of AI are applied to real-world workloads.
Listed here, we examine the essential characteristics of AI workloads, their storage input/output (I/O) profile, the types of storage suited to AI, the suitability of cloud and object storage for AI, and storage vendor approach and products for AI.
Technology tamfitronics What are the essential capabilities of AI workloads?
AI and ML are based mostly on training an algorithm to detect patterns in data, gain insight into data and often to trigger responses based on these findings. These can be very simple recommendations based on sales data, such as the “people who bought this also bought” type of recommendation. Or they can be the kind of advanced AI we see from natural language models (LLMs) in generative AI (GenAI) trained on large and multiple datasets to enable it to generate convincing text, images and video.
There are three key phases and deployment types for AI workloads:
Training, where patterns are learned into the algorithm from the AI model dataset, with varying degrees of human supervision;
Inference, in which the patterns identified in the training phase are applied, either in standalone AI deployments and/or;
Deployment of AI to an application or set of purposes.
The place and how AI and ML workloads are processed and run can vary significantly. On the one hand, they may resemble batch or one-off training and inference runs that resemble high-performance computing (HPC) processing on explicit datasets in science and research environments. On the other hand, AI, once trained, can also be applied to continuous application workloads, such as the types of sales and marketing operations described above.
The types of data in training and operational datasets may range from a large number of small data items, such as sensor readings in Internet of Things (IoT) workloads, to very large objects such as image and video data or discrete batches of scientific data. File size upon ingestion also depends on the AI frameworks in use (see below).
Datasets can also constitute a fraction of essential or secondary data storage, such as sales data or data held in backups, which is increasingly seen as a valuable source of corporate data.
Technology tamfitronics What are the I/O characteristics of AI workloads?
Processing performance needs to be high to handle AI training and inference in a reasonable timeframe and with as many iterations as possible to maximize quality.
Infrastructure also potentially needs a way to scale significantly to handle very large training datasets and outputs from training and inference. It also requires high I/O speed between storage and processing, and potentially also a way to maintain data portability between locations to enable the most efficient processing.
Data is at risk of being unstructured and in large volumes, rather than structured and in databases.
Technology tamfitronics What form of storage do AI workloads need?
As we’ve seen, massive parallel processing using GPUs is the core of AI infrastructure. So, in short, the role of storage is to feed these GPUs as quickly as possible to ensure these very expensive hardware components are used optimally.
Most of the time, that means flash storage for low latency in I/O. Capacity required will vary depending on the size of workloads and the potential scale of the implications of AI processing, but a few terabytes, even petabytes, is possible.
Ample throughput can also be a factor as different AI frameworks store data differently, such as between PyTorch (a series of smaller data files) and TensorFlow (the reverse). So, it’s not simply a matter of getting data to GPUs quickly, but also ensuring the right volume and appropriate I/O capabilities.
Not long ago, storage providers have promoted flash-based storage — typically using high-density QLC flash — as a general-purpose storage solution, including for datasets previously considered “secondary,” such as backup data, because users may now need to access it at higher speeds using AI.
Storage for AI initiatives will range from solutions offering very high performance during training and inference to various types of long-term retention, since it is not always clear at the outset of an AI project which data will prove valuable.
Technology tamfitronics Is cloud storage correct for AI workloads?
Cloud storage in general is a viable consideration for AI workload data. Essentially the most moving thing about holding data in the cloud brings a part of portability, with data ready to be “moved” nearer to its processing space.
Many AI initiatives start in the cloud attributable to it is possible you’ll doubtless exercise the GPUs for the time you will want them. The cloud just will not be cheap, however to deploy hardware on-premise, you like to occupy committed to a producing project sooner than it is far justified.
The total key cloud providers provide AI services that range from pre-professional units, application programming interfaces (APIs) into units, AI/ML compute with scalable GPU deployment (Nvidia and their hold) and storage infrastructure scalable to extra than one petabytes.
Technology tamfitronics Is object storage correct for AI workloads?
Object storage is correct for unstructured data, ready to scale hugely, usually located in the cloud, and can handle nearly any data type as an object. That makes it well-suited for the messy, unstructured data workloads possible in AI and ML applications.
The presence of rich metadata is one more plus to object storage. It can also be searched and browsed to support securing and organizing the appropriate data for AI training units. Data can even be held nearly anywhere, including in the cloud with communication via the S3 protocol.
However metadata, for all its advantages, could also overwhelm storage controllers and affect performance. And, if cloud is a place for cloud storage, cloud costs must still be taken into account as data is accessed and moved.
Technology What storage suppliers provide for AI
Nvidia provides reference architectures and hardware stacks that comprise servers, GPUs and networking. These are the DGX BasePOD reference structure and DGX SuperPOD turnkey infrastructure stack, that is at possibility of be specified for industry verticals.
Storage suppliers have also focused on the I/O bottleneck so data can be delivered efficiently to large numbers of (very costly) GPUs.
Those efforts have ranged from integrations with Nvidia infrastructure – the indispensable participant in GPU and AI server technology – via microservices such as NeMo for training and NIM for inference to storage product validation with AI infrastructure, and to total storage infrastructure stacks aimed at AI.
Supplier initiatives have also focused on the improvement of retrieval-augmented generation (RAG) pipelines and hardware architectures to enhance it. RAG validates the outputs of AI models by referencing external, trusted data, in part to address so-called hallucinations.
Technology tamfitronics Which storage suppliers provide products validated for Nvidia DGX?
Several storage suppliers offer products validated with DGX systems, including the following.
DataDirect Networks (DDN) provides its A³I AI400X2 all-NVMe storage appliances with SuperPOD. Each appliance delivers up to 90 GBps throughput and three million IOPS.
Dell’s AI Factory is an integrated hardware stack spanning desktop, laptop and server PowerEdge XE9680 compute, PowerScale F710 storage, systems and services, and validated with Nvidia’s AI infrastructure. It is available via Dell’s Apex as-a-service offering.
IBM has Spectrum Storage for AI with Nvidia DGX. It is a converged, yet independently scalable compute, storage and networking solution validated for Nvidia BasePOD and SuperPod.
Backup provider Cohesity announced at Nvidia’s GTC 2024 event that it would integrate Nvidia NIM microservices and Nvidia AI Enterprise into its Gaia multicloud data platform, which enables use of backup and archive data to create a source of training data.
Hammerspace has GPUDirect certification with Nvidia. Hammerspace markets its Hyperscale NAS as a global file system designed for AI/ML workloads and GPU-accelerated processing.
Hitachi Vantara has its Hitachi iQ, which presents industry-specific AI methods that use Nvidia DGX and HGX GPUs with the company’s storage.
HPE has GenAI supercomputing and enterprise methods with Nvidia parts, a RAG reference architecture, and plans to create in NIM microservices. In March 2024, HPE upgraded its Alletra MP storage arrays to connect twice the number of servers and 4 times the capacity in the same rackspace with 100Gbps connectivity between nodes in a cluster.
NetApp has product integrations with BasePOD and SuperPOD. At GTC 2024 NetApp announced integration of Nvidia’s NeMo Retriever microservice, a RAG system providing, with OnTap buyer hybrid cloud storage.
Pure Storage has AIRI, a flash-based AI infrastructure licensed with DGX and Nvidia OVX servers and the use of Pure’s FlashBlade//S storage. At GTC 2024, Pure announced it had created a RAG pipeline that uses Nvidia NeMo-based microservices with Nvidia GPUs and its storage, plus RAGs for explicit industry verticals.
Nice Records launched its Nice Records Platform in 2023, which combines its QLC flash-and-rapid-cache storage subsystems with database-like capabilities at the native storage I/O level, and achieved DGX certification.
In March 2024, hybrid cloud NAS maker Weka announced a hardware appliance licensed to work with Nvidia’s DGX SuperPod AI datacenter infrastructure.
Recent advances in machine perfusion technology are not only preserving donated livers but also making them biologically younger at a molecular level. This breakthrough…
What happens when human imagination, artificial intelligence, art and technology come together? From September 1 to 6, Samsung Electronics presented a special exhibition themed…