skip to main content

Lenovo EveryScale HPC & AI Software Stack

Product Guide

Home
Top
Updated
10 Aug 2026
Form Number
LP1651
PDF size
25 pages, 940 KB

Abstract

The Lenovo EveryScale HPC & AI Software Stack combines open-source with proprietary best-of-breed supercomputing software to provide the most consumable open-source HPC software stack embraced by all Lenovo HPC customers.

This product guide provides essential pre-sales information to understand the key features and components of the EveryScale HPC & AI Software Stack. The product guide is intended for technical specialists, sales specialists, sales engineers, IT architects, and other IT professionals who want to learn more about the Lenovo EveryScale HPC & AI Software Stack.

Change History

Changes in the August 10, 2026 update:

  • Updates under - Orchestration and management section
    • Replaced SchedMD parts with new NVIDIA parts
    • Added Slinky parts
    • Updated Slurm & Slinky product descriptions
  • Updates under - AI application workload and advanced gpu orchestration section
    • Updated the descriptions for NVIDIA AI Enterprise, Mission Control and Run:AI
    • Added NVIDIA AI Enterprise parts for RTX Pro 4500 and RTX Pro 6000
    • Added NVIDIA Run:AI parts for RTX Pro 4500 and RTX Pro 6000
  • Updates under - Energy monitoring and optimization section
    • Removed deprecated Energy Aware Runtime parts
    • Added new EAR product offering and updated its description

Introduction

The Lenovo EveryScale HPC & AI Software Stack combines open-source with proprietary best-of-breed Supercomputing software to provide the most consumable open-source HPC software stack embraced by all Lenovo HPC customers.

It provides a fully tested and supported, complete but customizable HPC software stack to enable the administrators and users in optimally and environmentally sustainable utilizing their Lenovo Supercomputers.

The software stack is built on the most widely adopted and maintained HPC community software for orchestration and management. It integrates third party components especially around programming environments and performance optimization to complement and enhance the capabilities, creating the organic umbrella in software and service to add value for our customers.

The software stack offers key software and support components for orchestration and management, programming environments and services and support, as shown in the following figure.

Lenovo HPC & AI Software Stack
Figure 1. Lenovo EveryScale HPC & AI Software Stack

Did you know?

Lenovo EveryScale HPC & AI Software Stack is a modular software stack tailored to our customer's needs. Thoroughly tested, supported and periodically updated, it combines the latest open-source HPC software releases to enable organizations with an agile and scalable IT infrastructure.

Benefits

The Lenovo EveryScale HPC & AI Software Stack provides the following benefits to customers.

Overcoming the Complexity of HPC Software

An HPC system software stack consists of dozens of components, that administrators must integrate and validate before an organization’s HPC applications can run on top of the stack. Ensuring stable, reliable versions of all stack components is an enormous task due to the numerous interdependencies. This task is very time consuming because of the constant release cycles and updates of individual components.

The Lenovo EveryScale HPC & AI Software Stack is fully tested, supported and periodically updated to combine the latest open-source HPC software releases, enabling organizations with an agile and scalable IT infrastructure.

Benefits of the Open-source Model

Going forward, in IDC's opinion, the development model exemplified by Linux is more workable. In this model, stack development is driven primarily by the open-source community and vendors offer supported distributions with additional capabilities for customers that require and are willing to pay for them. As the Linux initiative demonstrates, a community-based model like this has major advantages for enabling software to keep pace with requirements for HPC computing and storage hardware systems.

This model delivers new capabilities faster to users and makes HPC systems more productive and higher returning investments.

A fair number of foundational open source HPC software components already exist (e.g., Open MPI, Rocky Linux, Slurm, OpenStack, and others). Many HPC community members are already taking advantage of these.

Customers will benefit from the HPC community, as the community works to integrate a multitude of components that are commonly used in HPC systems and are freely available for open source distribution.

The key open-source components of the software stack are:

EAR is a powerful European energy management suite supporting anything from monitoring over power capping to live-optimization during the application runtime. Lenovo is collaborating with EAS on the continuous development and support and offers three products with differentiating capabilities.

  • Confluent Management

    Confluent is Lenovo-developed open-source software designed to discover, provision, and manage HPC clusters and the nodes that comprise them. Confluent provides powerful tooling to deploy and update software and firmware to multiple nodes simultaneously, with simple and readable modern software syntax.

  • Slurm Orchestration

    Slurm is integrated as an open source, flexible, and modern choice to manage complex workloads for faster processing and optimal utilization of the large-scale and specialized high-performance and AI resource capabilities needed per workload provided by Lenovo systems. Lenovo provides support in partnership with NVIDIA.

  • Slinky
    Slinky Bridge is a software solution that enables HPC and AI convergence by allowing Slurm and Kubernetes workloads to share the same compute infrastructure. Acting as the source of truth for scheduling decisions, Slurm extends its scheduling, accounting, and resource management capabilities to Kubernetes workloads while maintaining support for traditional HPC jobs. This unified approach reduces GPU fragmentation, improves resource utilization, and allows organizations to support both HPC and cloud-native AI workflows on a common platform.

  • Energy Aware Runtime

    EAR is a powerful European open-source energy management suite supporting anything from monitoring over power capping to live-optimization during the application runtime. Lenovo collaborates with Barcelona Supercomputing Centre (BSC) and Energy Aware Solutions (EAS) on the continuous development and support and offers three versions with differentiating capabilities.

  • Open OnDemand

    Open OnDemand (OOD) is an open source, web-based access portal that provides a unified Graphical User Interface (GUI) for seamless access to HPC and AI clusters. It enables users to securely access compute resources, submit and monitor jobs, and launch rich interactive applications, such as Jupyter Notebooks, VS Code, RStudio, and TensorBoard, directly from a standard web browser. Through these interactive environments, OOD enables end-users to run complete AI and data-science workflows based on popular frameworks like TensorFlow, PyTorch, RAPIDS, and more, without needing to manage the underlying cluster complexity. Supporting multiple resource managers, including Slurm, PBS Pro, LSF, and Kubernetes, Open OnDemand makes it easy to consolidate HPC and AI workloads on a shared infrastructure, significantly improving usability, accessibility, and overall productivity while eliminating the need for command-line expertise.

Orchestration and management

The following orchestration and management software is available with Lenovo EveryScale HPC & AI Software Stack:

  • Confluent (Best Recipe interoperability)
    Confluent is Lenovo-developed open source software designed to discover, provision, and manage HPC clusters and the nodes that comprise them. Confluent provides powerful tooling to deploy and update software and firmware to multiple nodes simultaneously, with simple and readable modern software syntax. Additionally, Confluent’s performance scales seamlessly from small workstation clusters to thousand-plus node supercomputers. For more information, see the Confluent Online documentation and the Lenovo Confluent Management Software paper.

  • Slurm
    Slurm is a modern, open-source scheduler designed specifically to satisfy the demanding needs of high-performance computing (HPC), high throughput computing (HTC), and AI. Slurm is developed and maintained by NVIDIA. Slurm maximizes workload throughput, scale, and reliability while optimizing resource utilization and enforcing organizational policies. It automates job scheduling across on-premises, hybrid, and cloud environments, helping administrators and users manage increasingly complex compute infrastructures. Slurm’s modern, plug-in-based architecture leverages REST APIs and supports deployments ranging from small clusters to some of the world's largest HPC and AI systems.

  • Slinky
    Slinky Bridge is a Slurm-centric HPC/AI convergence solution that enables organizations to run both traditional Slurm jobs and Kubernetes workloads on a shared pool of compute resources. By making Slurm the single scheduling authority, Slinky applies advanced scheduling capabilities such as fair-share, QoS, priorities, accounting, and topology awareness to both HPC and cloud-native AI workloads. This approach helps eliminate GPU silos, improve infrastructure utilization, and simplify operations by allowing customers to support HPC simulations, AI training, and AI inference workloads on a unified platform without statically partitioning cluster resources.

  • NVIDIA Slurm & Slinky Support
    Service capabilities for Lenovo HPC and AI systems include:

    • Level 3 Support: High-performance systems must operate at high levels of utilization and reliability to maximize return on investment. Customers covered by a support contract gain direct access to NVIDIA engineering experts for assistance with complex workload management issues, configuration challenges, troubleshooting, and operational best practices, helping accelerate problem resolution and minimize downtime. Customers with more demanding operational requirements may also purchase a Premium Support add-on, which provides enhanced Severity-2 and Severity-3 SLAs.
      For more information, see the NVIDIA Support Terms for Slurm and Slinky.

    • Proof of Concept (PoC): NVIDIA Proof of Concept services help organizations validate Slurm and/or Slinky deployments before production rollout. Working alongside NVIDIA experts, customers can evaluate architecture options, test workload scheduling and resource management capabilities, validate integration with existing infrastructure, and assess operational readiness. These engagements reduce deployment risk, accelerate adoption, and help ensure the solution meets business and technical requirements.

    • Slurm Training: NVIDIA provides a comprehensive 3-day instructor-led training program, delivered on-site at the customer location and supported by NVIDIA’s dedicated lab environment for hands-on exercises. The course combines technical sessions and practical labs for users and administrators, covering job submission and management, scheduling, GPU resource allocation, accounting and reporting, topology awareness, Quality of Service (QoS), Fair Share scheduling, reservations, containerized workloads, monitoring, troubleshooting, and advanced Slurm administration. The training is designed to accelerate user adoption, improve operational efficiency, and help organizations maximize the value of their Slurm environments.

    • Slinky Training: NVIDIA also offers Slinky training services to help users and administrators develop the skills required to deploy, operate, and manage a hybrid HPC/AI platform managed by Slinky.

  • NVIDIA Unified Fabric Manager (UFM) (ISV supported)
    NVIDIA Unified Fabric Manager (UFM) is InfiniBand networking management software that combines enhanced, real-time network telemetry with fabric visibility and control to support scale-out InfiniBand data centers. For more information, see the NVIDIA UFM product page.

    The two UFM offerings available from Lenovo are as follows:

    • UFM Telemetry for Real-Time Monitoring

      The UFM Telemetry platform provides network validation tools to monitor network performance and conditions, capturing and streaming rich real-time network telemetry information, application workload usage, and system configuration to an on-premises or cloud-based database for further analysis.

    • UFM Enterprise for Fabric Visibility and Control

      The UFM Enterprise platform combines the benefits of UFM Telemetry with enhanced network monitoring and management. It performs automated network discovery and provisioning, traffic monitoring, and congestion discovery. It also enables job schedule provisioning and integrates with industry-leading job schedulers and cloud and cluster managers, including Slurm and Platform Load Sharing Facility (LSF).

The following table lists all Orchestration and Management software available with Lenovo EveryScale HPC & AI Software Stack.

Table 1. Confluent
Part number Feature code Description
Lenovo Confluent support
7S090039WW S9VH Lenovo Confluent 1 Year Support per managed node
7S09003AWW S9VJ Lenovo Confluent 3 Year Support per managed node
7S09003BWW S9VK Lenovo Confluent 5 Year Support per managed node
7S09003CWW S9VL Lenovo Confluent 1 Extension Year Support per managed node
Table 2. NVIDIA Slurm and Slinky
Part number Description
NVIDIA Slurm and Slinky PoC and Training
7S02006QWW NVIDIA POC for Slurm, 1Year (SKU also valid for Slinky)
7S02006WWW NVIDIA Training for Slurm, 1Year
VLS NVIDIA Training for Slinky, 1Year
NVIDIA Slurm PoC and Training (EDU)
7S02006JWW NVIDIA POC for Slurm, EDU, 1Year (SKU also valid for Slinky)
7S02006SWW NVIDIA Training for Slurm, EDU, 1Year
VLS NVIDIA Training for Slinky, EDU, 1Year
NVIDIA Slurm Standard Support
7S02006PWW NVIDIA Standard Support for Slurm per Node, 1Year (Minimum Order Quantity: 200 units)
7S02006NWW NVIDIA Standard Support for Slurm per Node, 2Years (Minimum Order Quantity: 200 units)
7S02006RWW NVIDIA Standard Support for Slurm per Node, 3Years (Minimum Order Quantity: 200 units)
7S02006BWW NVIDIA Standard Support for Slurm per Node, 4Years (Minimum Order Quantity: 200 units)
7S020069WW NVIDIA Standard Support for Slurm per Node, 5Years (Minimum Order Quantity: 200 units)
7S020074WW NVIDIA Standard Support for Slurm per GPU, 1Year
7S020076WW NVIDIA Standard Support for Slurm per GPU, 2Years
7S020078WW NVIDIA Standard Support for Slurm per GPU, 3Years
7S02007AWW NVIDIA Standard Support for Slurm per GPU, 4Years
7S02007CWW NVIDIA Standard Support for Slurm per GPU, 5Years
NVIDIA Slurm Standard Support (EDU)
7S02006DWW NVIDIA Standard Support for Slurm per Node, EDU, 1Year (Minimum Order Quantity: 200 units)
7S02006EWW NVIDIA Standard Support for Slurm per Node, EDU, 2Years (Minimum Order Quantity: 200 units)
7S02006CWW NVIDIA Standard Support for Slurm per Node, EDU, 3Years (Minimum Order Quantity: 200 units)
7S02006LWW NVIDIA Standard Support for Slurm per Node, EDU, 4Years (Minimum Order Quantity: 200 units)
7S02006MWW NVIDIA Standard Support for Slurm per Node, EDU, 5Years (Minimum Order Quantity: 200 units)
7S020075WW NVIDIA Standard Support for Slurm per GPU, EDU, 1Year
7S020077WW NVIDIA Standard Support for Slurm per GPU, EDU, 2Years
7S020079WW NVIDIA Standard Support for Slurm per GPU, EDU, 3Years
7S02007BWW NVIDIA Standard Support for Slurm per GPU, EDU, 4Years
7S02007DWW NVIDIA Standard Support for Slurm per GPU, EDU, 5Years
NVIDIA Slurm Premium Support Add-On
7S02007QWW NVIDIA Premium Support for Slurm, 1Year
7S02007SWW NVIDIA Premium Support for Slurm, 2Years
7S02007UWW NVIDIA Premium Support for Slurm, 3Years
7S02007WWW NVIDIA Premium Support for Slurm, 4Years
7S02007YWW NVIDIA Premium Support for Slurm, 5Years
NVIDIA Slurm Premium Support Add-On (EDU)
7S02007RWW NVIDIA Premium Support for Slurm, EDU, 1Year
7S02007TWW NVIDIA Premium Support for Slurm, EDU, 2Years
7S02007VWW NVIDIA Premium Support for Slurm, EDU, 3Years
7S02007XWW NVIDIA Premium Support for Slurm, EDU, 4Years
7S02007ZWW NVIDIA Premium Support for Slurm, EDU, 5Years
NVIDIA Slinky Standard Support
7S02006AWW NVIDIA Standard Support for Slinky per Node, 1Year
7S02006VWW NVIDIA Standard Support for Slinky per Node, 2Years
7S02001VWW NVIDIA Standard Support for Slinky per Node, 3Years
7S02006HWW NVIDIA Standard Support for Slinky per Node, 4Years
7S02006KWW NVIDIA Standard Support for Slinky per Node, 5Years
7S02007EWW NVIDIA Standard Support for Slinky per GPU, 1Year
7S02007GWW NVIDIA Standard Support for Slinky per GPU, 2Years
7S02007JWW NVIDIA Standard Support for Slinky per GPU, 3Years
7S02007LWW NVIDIA Standard Support for Slinky per GPU, 4Years
7S02007NWW NVIDIA Standard Support for Slinky per GPU, 5Years
NVIDIA Slinky Standard Support (EDU)
7S02006UWW NVIDIA Standard Support for Slinky per Node, EDU, 1Year
7S02006XWW NVIDIA Standard Support for Slinky per Node, EDU, 2Years
7S02006TWW NVIDIA Standard Support for Slinky per Node, EDU, 3Years
7S02006GWW NVIDIA Standard Support for Slinky per Node, EDU, 4Years
7S02006FWW NVIDIA Standard Support for Slinky per Node, EDU, 5Years
7S02007FWW NVIDIA Standard Support for Slinky per GPU, EDU, 1Year
7S02007HWW NVIDIA Standard Support for Slinky per GPU, EDU, 2Years
7S02007KWW NVIDIA Standard Support for Slinky per GPU, EDU, 3Years
7S02007MWW NVIDIA Standard Support for Slinky per GPU, EDU, 4Years
7S02007PWW NVIDIA Standard Support for Slinky per GPU, EDU, 5Years


Notes on Slurm and Slinky PoC.

  • The NVIDIA POC for Slurm, 1 Year SKU is also used for Slinky PoC engagements.
  • The “1 Year” designation indicates the entitlement period during which the PoC may be scheduled and consumed. Customers have up to one year from the date of purchase to utilize the service.

Notes on Slurm and Slinky Training.

  • Training is optional but strongly recommended for new Slurm and/or Slinky users and administrators.
  • Training is delivered on-site by NVIDIA instructors and includes access to NVIDIA’s dedicated lab environment for hands-on exercises and practical learning.
  • The “1 Year” designation indicates the entitlement period during which the training may be scheduled and consumed. Customers have up to one year from the date of purchase to utilize the service.

Notes on Slurm Standard Support.

Slurm Standard Support subscriptions are based on the compute infrastructure being managed:

  • One per-node subscription is required for each compute node that executes workloads, excluding Slurm controller nodes.
  • GPU-accelerated nodes require both a per-node subscription and an additional per-GPU subscription for each installed GPU.
  • A minimum quantity of 200 units applies to the NVIDIA Standard Support for Slurm per-Node

Notes on Slinky Standard Support.

  • Slinky Standard Support follows the same licensing model as Slurm Standard Support.
  • One per-node subscription is required for each Slinky-managed compute node, with additional per-GPU subscriptions required for GPU-accelerated nodes.
  • Active Slurm Standard Support subscriptions are a prerequisite for Slinky support.
  • Customers requiring Slinky support must maintain active support subscriptions for both Slurm and Slinky.

Notes on Slurm Premium Support.

Slurm Premium Support is an optional add-on that requires an active Slurm Standard Support subscription.

  • Premium Support is intended for customers who require enhanced Severity-2 and Severity-3 SLAs.
  • A single Slurm Premium Support SKU applies to the entire cluster, regardless of the number of compute nodes being managed.
Table 3. NVIDIA Unified Fabric Manager (UFM)
Part number Feature code Description
UFM Telemetry
7S090011WW S921 NVIDIA UFM Telemetry 1-year License and 24/7 Support for Lenovo clusters
7S090012WW S922 NVIDIA UFM Telemetry 3-year License and 24/7 Support for Lenovo clusters
7S090013WW S923 NVIDIA UFM Telemetry 5-year License and 24/7 Support for Lenovo clusters
UFM Enterprise
7S09000XWW S91Y NVIDIA UFM Enterprise 1-year License and 24/7 Support for Lenovo clusters
7S09000YWW S91Z NVIDIA UFM Enterprise 3-year License and 24/7 Support for Lenovo clusters
7S09000ZWW S920 NVIDIA UFM Enterprise 5-year License and 24/7 Support for Lenovo clusters

Programming environment

The following programming software is available with Lenovo EveryScale HPC & AI Software Stack.

  • NVIDIA CUDA

    NVIDIA CUDA is a parallel computing platform and programming model for general computing on graphical processing units (GPUs). With CUDA, developers are able to dramatically speed up computing applications by harnessing the power of GPUs. When using CUDA, developers program in popular languages such as C, C++, Fortran, Python and MATLAB and express parallelism through extensions in the form of a few basic keywords. For more information, see the NVIDIA CUDA Zone.

  • NVIDIA HPC Software Development Kit

    The NVIDIA HPC SDK C, C++, and Fortran compilers support GPU acceleration of HPC modeling and simulation applications with standard C++ and Fortran, OpenACC directives, and CUDA. GPU-accelerated math libraries maximize performance on common HPC algorithms, and optimized communications libraries enable standards-based multi-GPU and scalable systems programming. Performance profiling and debugging tools simplify porting and optimization of HPC applications, and containerization tools enable easy deployment on-premises or in the cloud. For more information, see the NVIDIA HPC SDK.

  • Intel oneAPI
    The Intel oneAPI Base & HPC Toolkit is a comprehensive software development suite designed to empower developers in creating HPC & AI solutions that exploit the full potential of modern hardware architectures. This toolkit encompasses an array of advanced tools, libraries, and compilers, enabling programmers to efficiently design, optimize, and deploy parallel applications across diverse computing platforms, including CPUs, GPUs, and FPGAs. With a focus on fostering code portability and performance scalability, the Intel oneAPI Base & HPC Toolkit equips developers with the means to enhance productivity, streamline software development, and achieve exceptional performance outcomes in the realm of high-performance computing.
    • For more information, see Intel® oneAPI Base and HPC Toolkit
    • System-based pricing (introduced in mid-2023) can range from a small system (64-256 nodes), medium system (257-512 nodes), or large system (512+ nodes)
    • Developer-based pricing is for systems with fewer than 64 nodes and offers support for a limited number of users.
    • Part numbers are available for Commercial or Academic customers.
    • Support Renewals are available.
    • Commercial parts have different part numbers if they are quoted with or without Intel hardware.

The following tables list the relevant ordering part numbers.

Table 4. NVIDIA CUDA part numbers
Part number Description
NVIDIA CUDA
7S09001EWW NVIDIA CUDA Support and Maintenance (up to 200 GPUs), 1 Year
7S09001FWW NVIDIA CUDA Support and Maintenance (up to 500 GPUs), 1 Year
7S09002EWW NVIDIA CUDA Support and Maintenance (up to 1000 GPUs), 1 Year
7S09002FWW NVIDIA CUDA Support and Maintenance (up to 5000 GPUs), 1 Year
Table 5. NVIDIA HPC SDK part numbers
Part number Description
NVIDIA HPC SDK
7S090014WW NVIDIA HPC Compiler Support Services, 1 Year
7S090015WW NVIDIA HPC Compiler Support Services, 3 Years
7S09002GWW NVIDIA HPC Compiler Support Services, 5 Years
7S090016WW NVIDIA HPC Compiler Support Services, EDU, 1 Year
7S090017WW NVIDIA HPC Compiler Support Services, EDU, 3 Years
7S09002HWW NVIDIA HPC Compiler Support Services, EDU, 5 Years
7S09001CWW NVIDIA HPC Compiler Support Services - Additional Contact, 1 Year
7S09002JWW NVIDIA HPC Compiler Support Services - Additional Contact, 3 Years
7S09002KWW NVIDIA HPC Compiler Support Services - Additional Contact, 5 Years
7S09001DWW NVIDIA HPC Compiler Support Services - Additional Contact, EDU, 1 Year
7S09002LWW NVIDIA HPC Compiler Support Services - Additional Contact, EDU, 3 Years
7S09002MWW NVIDIA HPC Compiler Support Services - Additional Contact, EDU, 5 Years
7S09001AWW NVIDIA HPC Compiler Premier Support Services, 1 Year
7S09002NWW NVIDIA HPC Compiler Premier Support Services, 3 Years
7S09002PWW NVIDIA HPC Compiler Premier Support Services, 5 Years
7S09001BWW NVIDIA HPC Compiler Premier Support Services, EDU, 1 Year
7S09002QWW NVIDIA HPC Compiler Premier Support Services, EDU, 3 Years
7S09002RWW NVIDIA HPC Compiler Premier Support Services, EDU, 5 Years
7S090018WW NVIDIA HPC Compiler Premier Support Services - Additional Contact, 1 Year
7S09002SWW NVIDIA HPC Compiler Premier Support Services - Additional Contact, 3 Years
7S09002TWW NVIDIA HPC Compiler Premier Support Services - Additional Contact, 5 Years
7S090019WW NVIDIA HPC Compiler Premier Support Services - Additional Contact, EDU, 1 Year
7S09002UWW NVIDIA HPC Compiler Premier Support Services - Additional Contact, EDU, 3 Years
7S09002VWW NVIDIA HPC Compiler Premier Support Services - Additional Contact, EDU, 5 Years
Table 6. Intel oneAPI part numbers
Part number Description
Commercial - Small System (64 - 256 nodes)
7S09003DWW Intel oneAPI Base & HPC Toolkit (Multi-Node) Commercial Small System support for oneAPI Cluster runtimes 1YSupp with Intel HW
7S09003EWW Intel oneAPI Base & HPC Toolkit (Multi-Node) Commercial Small System support for oneAPI Cluster runtimes 3YSupp with Intel HW
7S09003FWW Intel oneAPI Base & HPC Toolkit (Multi-Node) Commercial Small System support for oneAPI Cluster runtimes 4YSupp with Intel HW
7S09003GWW Intel oneAPI Base & HPC Toolkit (Multi-Node) Commercial Small System support for oneAPI Cluster runtimes 5YSupp with Intel HW
Commercial - Medium System (257 - 512 nodes)
7S09003HWW Intel oneAPI Base & HPC Toolkit (Multi-Node) Commercial Medium System support for oneAPI Cluster runtimes 1YSupp with Intel HW
7S09003JWW Intel oneAPI Base & HPC Toolkit (Multi-Node) Commercial Medium System support for oneAPI Cluster runtimes 3YSupp with Intel HW
7S09003KWW Intel oneAPI Base & HPC Toolkit (Multi-Node) Commercial Medium System support for oneAPI Cluster runtimes 4YSupp with Intel HW
7S09003LWW Intel oneAPI Base & HPC Toolkit (Multi-Node) Commercial Medium System support for oneAPI Cluster runtimes 5YSupp with Intel HW
Commercial - Large System (512+ nodes)
7S09003QWW Intel oneAPI Base & HPC Toolkit (Multi-Node) Commercial Large System support for oneAPI Cluster runtimes 1YSupp with Intel HW
7S09003PWW Intel oneAPI Base & HPC Toolkit (Multi-Node) Commercial Large System support for oneAPI Cluster runtimes 3YSupp with Intel HW
7S09003NWW Intel oneAPI Base & HPC Toolkit (Multi-Node) Commercial Large System support for oneAPI Cluster runtimes 4YSupp with Intel HW
7S09003MWW Intel oneAPI Base & HPC Toolkit (Multi-Node) Commercial Large System support for oneAPI Cluster runtimes 5YSupp with Intel HW
Commercial – Developer Based (for systems < 64 nodes)
7S09003RWW Intel oneAPI Base & HPC Toolkit (Multi-Node) - 2 Concurrent User Commercial 1YSupp with Intel HW
7S09003SWW Intel oneAPI Base & HPC Toolkit (Multi-Node) - 2 Concurrent User Commercial 3YSupp with Intel HW
7S09003TWW Intel oneAPI Base & HPC Toolkit (Multi-Node) - 5 Concurrent User Commercial 1YSupp with Intel HW
7S09003UWW Intel oneAPI Base & HPC Toolkit (Multi-Node) - 5 Concurrent User Commercial 3YSupp with Intel HW
7S09003VWW Intel oneAPI Base & HPC Toolkit (Multi-Node) - 10 Concurrent User Commercial 1YSupp with Intel HW
7S09003WWW Intel oneAPI Base & HPC Toolkit (Multi-Node) - 10 Concurrent User Commercial 3YSupp with Intel HW
7S09003XWW Intel oneAPI Base & HPC Toolkit (Multi-Node) - Enterprise-Above 10 Concurrent Users
Commercial – Small System (64-256 nodes) with non-Intel HW
7S09004JWW Intel oneAPI Base & HPC Toolkit (Multi-Node) Commercial Small System support for oneAPI Cluster runtimes 1YSupp with non-Intel HW
7S09004KWW Intel oneAPI Base & HPC Toolkit (Multi-Node) Commercial Small System support for oneAPI Cluster runtimes 3YSupp with non-Intel HW
7S09004LWW Intel oneAPI Base & HPC Toolkit (Multi-Node) Commercial Small System support for oneAPI Cluster runtimes 4YSupp with non-Intel HW
7S09004MWW Intel oneAPI Base & HPC Toolkit (Multi-Node) Commercial Small System support for oneAPI Cluster runtimes 5YSupp with non-Intel HW
Commercial – Medium System (257 - 512 nodes) with non-Intel HW
7S09004NWW Intel oneAPI Base & HPC Toolkit (Multi-Node) Commercial Medium System support for oneAPI Cluster runtimes 1YSupp with non-Intel HW
7S09004PWW Intel oneAPI Base & HPC Toolkit (Multi-Node) Commercial Medium System support for oneAPI Cluster runtimes 3YSupp with non-Intel HW
7S09004QWW Intel oneAPI Base & HPC Toolkit (Multi-Node) Commercial Medium System support for oneAPI Cluster runtimes 4YSupp with non-Intel HW
7S09004RWW Intel oneAPI Base & HPC Toolkit (Multi-Node) Commercial Medium System support for oneAPI Cluster runtimes 5YSupp with non-Intel HW
Commercial – Large System (512+ nodes) with non-Intel HW
7S09004VWW Intel oneAPI Base & HPC Toolkit (Multi-Node) Commercial Large System support for oneAPI Cluster runtimes 1YSupp with non-Intel HW
7S09004UWW Intel oneAPI Base & HPC Toolkit (Multi-Node) Commercial Large System support for oneAPI Cluster runtimes 3YSupp with non-Intel HW
7S09004TWW Intel oneAPI Base & HPC Toolkit (Multi-Node) Commercial Large System support for oneAPI Cluster runtimes 4YSupp with non-Intel HW
7S09004SWW Intel oneAPI Base & HPC Toolkit (Multi-Node) Commercial Large System support for oneAPI Cluster runtimes 5YSupp with non-Intel HW
Commercial – Developer Based (for systems < 64 nodes) with non-Intel HW
7S09004WWW Intel oneAPI Base & HPC Toolkit (Multi-Node) - 2 Concurrent User Commercial 1YSupp with non-Intel HW
7S09004XWW Intel oneAPI Base & HPC Toolkit (Multi-Node) - 2 Concurrent User Commercial 3YSupp with non-Intel HW
7S09004YWW Intel oneAPI Base & HPC Toolkit (Multi-Node) - 5 Concurrent User Commercial 1YSupp with non-Intel HW
7S09004ZWW Intel oneAPI Base & HPC Toolkit (Multi-Node) - 5 Concurrent Users Commercial 3YSupp with non-Intel HW
7S090050WW Intel oneAPI Base & HPC Toolkit (Multi-Node) - 10 Concurrent User Commercial 1YSupp with non-Intel HW
7S090051WW Intel oneAPI Base & HPC Toolkit (Multi-Node) - 10 Concurrent Users Commercial 3YSupp with non-Intel HW
Academic - Small System (64 - 256 nodes)
7S09003YWW Intel oneAPI Base & HPC Toolkit (Multi-Node) Academic Small System support for oneAPI Cluster runtimes with 1YSupp
7S09003ZWW Intel oneAPI Base & HPC Toolkit (Multi-Node) Academic Small System support for oneAPI Cluster runtimes with 3YSupp
7S090040WW Intel oneAPI Base & HPC Toolkit (Multi-Node) Academic Small System support for oneAPI Cluster runtimes with 4YSupp
7S090041WW Intel oneAPI Base & HPC Toolkit (Multi-Node) Academic Small System support for oneAPI Cluster runtimes with 5YSupp
Academic - Medium System (257 - 512 nodes)
7S090042WW Intel oneAPI Base & HPC Toolkit (Multi-Node) Academic Medium System support for oneAPI Cluster runtimes with 1YSupp
7S090043WW Intel oneAPI Base & HPC Toolkit (Multi-Node) Academic Medium System support for oneAPI Cluster runtimes with 3YSupp
7S090044WW Intel oneAPI Base & HPC Toolkit (Multi-Node) Academic Medium System support for oneAPI Cluster runtimes with 4YSupp
7S090045WW Intel oneAPI Base & HPC Toolkit (Multi-Node) Academic Medium System support for oneAPI Cluster runtimes with 5YSupp
Academic - Large System (512+ nodes)
7S090046WW Intel oneAPI Base & HPC Toolkit (Multi-Node) Academic Large System support for oneAPI Cluster runtimes with 1YSupp
7S090047WW Intel oneAPI Base & HPC Toolkit (Multi-Node) Academic Large System support for oneAPI Cluster runtimes with 3YSupp
7S090048WW Intel oneAPI Base & HPC Toolkit (Multi-Node) Academic Large System support for oneAPI Cluster runtimes with 4YSupp
7S090049WW Intel oneAPI Base & HPC Toolkit (Multi-Node) Academic Large System support for oneAPI Cluster runtimes with 5YSupp
Academic - Developer Based (for systems < 64 nodes)
7S09004AWW Intel oneAPI Base & HPC Toolkit (Multi-Node) - 2 Concurrent User Academic 1YSupp with Intel HW
7S09004BWW Intel oneAPI Base & HPC Toolkit (Multi-Node) - 2 Concurrent User Academic 3YSupp with Intel HW
7S09004CWW Intel oneAPI Base & HPC Toolkit (Multi-Node) - 5 Concurrent User Academic 1YSupp with Intel HW
7S09004DWW Intel oneAPI Base & HPC Toolkit (Multi-Node) - 5 Concurrent User Academic 3YSupp with Intel HW
7S09004EWW Intel oneAPI Base & HPC Toolkit (Multi-Node) - 10 Concurrent User Academic 1YSupp with Intel HW
7S09004FWW Intel oneAPI Base & HPC Toolkit (Multi-Node) - 10 Concurrent User Academic 3YSupp with Intel HW

AI: application workload and advanced GPU orchestration

The following software suite is available with Lenovo EveryScale HPC & AI Software Stack.

  • NVIDIA AI Enterprise
    NVAIE is an end-to-end, cloud-native suite of AI and data analytics software, optimized, certified, and supported by NVIDIA to run on VMware vSphere and bare-metal with NVIDIA-Certified Systems™.  It includes key enabling technologies from NVIDIA for rapid deployment, management, and scaling of AI workloads in the modern hybrid cloud. NVAIE is licensed on a per-GPU basis and can be purchased as either a perpetual license with support services, or as an annual or multi-year subscription.
    • The perpetual license provides the right to use the NVIDIA AI Enterprise software indefinitely, with no expiration. NVIDIA AI Enterprise with perpetual licenses must be purchased in conjunction with one-year, three-year, or five-year support services. A one-year support service is also available for renewals.
    • The subscription offerings are an affordable option to allow IT departments to better manage the flexibility of license volumes. NVIDIA AI Enterprise software products with subscription includes support services for the duration of the software’s subscription license.

The key capabilities of NVAIE offerings include the following:

  • Full-Stack AI Platform – Includes frameworks like TensorFlow, PyTorch, RAPIDS, XGBoost, Triton Inference Server, and more.
  • Generative AI Enablement – Provides pre-built pipelines for Large Language Models (LLMs), Retrieval-Augmented Generation (RAG), and fine-tuning.
  • Enterprise-Grade Support – Backed by NVIDIA with 24/7 support, performance optimization, and validated infrastructure.
  • Security & Governance – Supports secure, multi-tenant environments with integrated access control and policy enforcement.
  • Hybrid-Cloud Ready – Deployable across VMware, Red Hat OpenShift, Kubernetes, or bare metal, with full container support.

Discounted software promotion: Order NVIDIA Enterprise with either NVIDIA RTX PRO 4500 Blackwell Server Edition or NVIDIA RTX PRO 6000 Blackwell Server Edition and get the software at a discount. For more information, see the Discounted software with RTX PRO Server GPU bundle.

License included: A 5-year NVIDIA Enterprise subscription is included with H100 and H200 PCIe double-wide GPUs on Lenovo ThinkSystem servers. Customer can redeem the license through this link: https://www.nvidia.com/en-us/data-center/activate-license/.

  • NVIDIA Run:AI
    As AI adoption accelerates, many organizations encounter a new bottleneck: underutilized or fragmented GPU resources across teams and workloads. NVIDIA Run:AI resolves this by introducing an intelligent orchestration and scheduling platform purpose-built for AI infrastructure. It enables multi-tenant GPU sharing, workload prioritization, and centralized resource management, dramatically improving ROI on AI infrastructure.

    When deployed on Lenovo ThinkSystem servers, Run:AI transforms traditional GPU clusters into fully virtualized AI infrastructure, capable of allocating GPUs dynamically based on policies, project needs, or user demand: on-prem, in hybrid environments, or across Kubernetes clusters.

    The key capabilities of Run:AI include the following:

    Discounted software promotion: Order Run:ai with either NVIDIA RTX PRO 4500 Blackwell Server Edition or NVIDIA RTX PRO 6000 Blackwell Server Edition and get the software at a discount. For more information, see the Discounted software with RTX PRO Server GPU bundle.

    • Dynamic GPU Scheduling – Allocates GPU resources to users and projects in real time—full, fractional, or virtualized.
    • Fair Share Quotas – Ensures teams get guaranteed GPU access without overprovisioning or idle capacity.
    • Multi-Tenancy – Supports multiple departments or teams with isolated workloads, policies, and quotas.
    • Kubernetes-Native Integration – Runs seamlessly on K8s, OpenShift, or any CNCF-compliant container environment.
    • Visibility & Dashboards – Offers real-time metrics, cost tracking, and usage visualization for IT and leadership.
    • AI Workload Prioritization – Automatically prioritizes high-value, production workloads over experimentation.

      For more information, see the NVIDIA Run:AI product page:
      https://www.nvidia.com/en-us/software/run-ai/
  • NVIDIA Mission Control
    NVIDIA Mission Control™ is an enterprise-grade operations software platform designed to deploy, manage, and operate large-scale AI infrastructure with consistency, resilience, and efficiency. Built for modern data centers and hybrid environments, Mission Control enables organizations to operationalize AI faster, reduce complexity, and maintain full visibility across compute, networking, and accelerated workloads.

    Mission Control abstracts infrastructure complexity and delivers a centralized control plane for AI clusters, enabling IT and platform teams to focus on delivering AI outcomes, not managing infrastructure silos.

    Key capabilities of NVIDIA Mission Control include:
    • Unified AI Infrastructure Management
      • Centralized orchestration of NVIDIA-accelerated AI clusters
      • Lifecycle management for compute, networking, and system software
      • Policy-driven operations across multi-node, multi-rack deployments
    • Accelerated Deployment & Provisioning
      • Rapid, repeatable deployment of AI-ready infrastructure
      • Pre-validated configurations aligned with NVIDIA reference architectures
      • Automated setup of system software, drivers, and dependencies
    • Operational Visibility & Control
      • Real-time monitoring of cluster health, utilization, and performance
      • End-to-end observability across GPUs, CPUs, networking, and storage
      • Integrated alerting and diagnostics for proactive operations
    • Resilience & Reliability
      • Built-in support for fault detection and remediation
      • Consistent operations across environments to reduce configuration drift
      • Designed for mission-critical AI workloads
    • Enterprise-Grade Security & Governance
      • Role-based access control (RBAC)
      • Secure, policy-driven operations aligned with enterprise IT standards
      • Designed to support regulated and sovereign AI environments.

For more information, see the NVIDIA Mission Control product page:
https://www.nvidia.com/en-us/data-center/mission-control/

  • ClearML: Innovating AI at Scale with Lenovo

As a Lenovo innovator in the AI software stack, ClearML delivers an enterprise-grade AI and GenAI platform that accelerates the entire AI lifecycle. Designed for flexibility and performance, ClearML provides a cloud-like experience with the efficiency and security of Lenovo’s infrastructure solutions.

ClearML’s architecture is built on three layers – Infrastructure Control Plane, AI Development Center, and GenAI App Engine – each enabling Lenovo customers to unlock the full potential of AI:

  • Infrastructure Control Plane – ClearML enables GPU optimization and management across heterogeneous environments with dynamic resource allocation for NVIDIA and AMD GPUs, granular quota management with per-tenant billing, and isolated multi-tenancy. The platform supports silicon-agnostic workload scheduling across any environment, enables hybrid multi-cluster orchestration with cloud bursting, and deploys flexibly across on-premise or air-gapped infrastructure.
    • Dynamic Fractional GPUs – ClearML enables highly efficient use of GPU resources by allowing granular allocation and sharing of GPU fractions. This dynamic approach optimizes performance for a wide range of AI workloads, supporting both NVIDIA MIG-capable and non-MIG devices and AMD MI300X and newer GPUs. It allows organizations to maximize their hardware investment and adapt resource allocation in real time as workload demands change.
    • Multi-tenant, Multi-cluster – The platform supports secure, dynamic resource sharing across multiple users and clusters. Its secure multi-tenancy model includes robust network isolation, ensuring privacy and compliance for different teams or projects. Resources are allocated flexibly based on real-time demand, and new tenants can be onboarded quickly without the need for dedicated clusters, maximizing overall resource utilization.
  • AI Development Center – The platform offers a comprehensive workbench for AI/ML development, integrating tools for data management, experiment tracking, and pipeline orchestration. Users can launch familiar development environments like Jupyter and VSCode with a single-click, manage datasets, and automate complex AI workflows—all from a unified interface. This abstraction allows data scientists and engineers to focus on their code and experiments, minimizing the need to interact directly with Kubernetes or underlying infrastructure.
  • GenAI App Engine – ClearML simplifies the testing, deployment, and management of large language models (LLMs). With one-click access, users can deploy custom or off-the-shelf (Hugging Face) models, NVIDIA NIM containers, or custom applications with pre-configured networking and RBAC authentication. An endpoint dashboard provides teams with full visibility on usage and performance, and static routing provides horizontal scalability for peak periods. The GenAI App Engine supports secure deployment of both open-source and custom models.
  • NVIDIA AI Enterprise integration – ClearML integrates with NVIDIA AI Enterprise (NVAIE), including dynamic license allocation that allows organizations to optimize costs and scale efficiently. The platform orchestrates NVIDIA-optimized containers (such as NIM and Triton), accelerating AI development and providing observability for inference workloads.

Lenovo and ClearML together deliver a secure, scalable, and innovative AI platform – empowering organizations to accelerate innovation and turn vision into production-grade AI solutions.

For more information: https://clear.ml/

The following tables list the ordering part numbers.

Table 7. NVIDIA AI Enterprise Software (NVAIE)
Part number Description
NVIDIA Enterprise for use bundled with the NVIDIA RTX PRO 4500 Blackwell Server Edition GPU
7S02006YWW NVIDIA RTX PRO 4500 Blackwell Server Edition Software Kit, per GPU, 3Years
7S02006ZWW NVIDIA RTX PRO 4500 Blackwell Server Edition Software Kit, per GPU, 5Years
NVIDIA Enterprise for use bundled with the NVIDIA RTX PRO 6000 Blackwell Server Edition GPU
7S020063WW NVIDIA RTX PRO 6000 Blackwell Server Edition Software Kit, per GPU, 3 Years
7S020064WW NVIDIA RTX PRO 6000 Blackwell Server Edition Software Kit, per GPU, 5 Years
AI Enterprise Perpetual License
7S02001BWW NVIDIA Enterprise (NVIDIA AI Enterprise and NVIDIA Omniverse Enterprise) Perpetual License & Support per GPU, 5 Years
7S02001EWW NVIDIA Enterprise (NVIDIA AI Enterprise and NVIDIA Omniverse Enterprise) Perpetual License & Support per GPU, EDU, 5 Years
AI Enterprise Subscription License
7S02001FWW NVIDIA Enterprise (NVIDIA AI Enterprise and NVIDIA Omniverse Enterprise) Subscription per GPU, 1 Year
7S02005XWW NVIDIA Enterprise (NVIDIA AI Enterprise and NVIDIA Omniverse Enterprise) Subscription per GPU, 2 Years
7S02001GWW NVIDIA Enterprise (NVIDIA AI Enterprise and NVIDIA Omniverse Enterprise) Subscription per GPU, 3 Year
7S02005YWW NVIDIA Enterprise (NVIDIA AI Enterprise and NVIDIA Omniverse Enterprise) Subscription per GPU, 4 Years
7S02001HWW NVIDIA Enterprise (NVIDIA AI Enterprise and NVIDIA Omniverse Enterprise) Subscription per GPU, 5 Year
7S02001JWW NVIDIA Enterprise (NVIDIA AI Enterprise and NVIDIA Omniverse Enterprise) Subscription per GPU, EDU, 1 Year
7S02005ZWW NVIDIA Enterprise (NVIDIA AI Enterprise and NVIDIA Omniverse Enterprise) Subscription per GPU, EDU, 2 Years
7S02001KWW NVIDIA Enterprise (NVIDIA AI Enterprise and NVIDIA Omniverse Enterprise) Subscription per GPU, EDU, 3 Years
7S020060WW NVIDIA Enterprise (NVIDIA AI Enterprise and NVIDIA Omniverse Enterprise) Subscription per GPU, EDU, 4 Years
7S02001LWW NVIDIA Enterprise (NVIDIA AI Enterprise and NVIDIA Omniverse Enterprise) Subscription per GPU, EDU, 5 Years
Business Critical Support Services for NVIDIA Enterprise
7S02001MWW Business Critical Support Services for NVIDIA Enterprise (NVIDIA AI Enterprise and Omniverse Enterprise) per GPU, 1 Year
7S02001NWW Business Critical Support Services for NVIDIA Enterprise (NVIDIA AI Enterprise and Omniverse Enterprise) per GPU, 3 Years
7S020061WW Business Critical Support Services for NVIDIA Enterprise (NVIDIA AI Enterprise and Omniverse Enterprise) per GPU, 4 Years
7S02001PWW Business Critical Support Services for NVIDIA Enterprise (NVIDIA AI Enterprise and Omniverse Enterprise) per GPU, 5 Years
7S02001QWW Business Critical Support Services for NVIDIA Enterprise (NVIDIA AI Enterprise and Omniverse Enterprise) per GPU, EDU, 1 Year
7S02001RWW Business Critical Support Services for NVIDIA Enterprise (NVIDIA AI Enterprise and Omniverse Enterprise) per GPU, EDU, 3 Years
7S020062WW Business Critical Support Services for NVIDIA Enterprise (NVIDIA AI Enterprise and Omniverse Enterprise) per GPU, EDU, 4 Years
7S02001SWW Business Critical Support Services for NVIDIA Enterprise (NVIDIA AI Enterprise and Omniverse Enterprise) per GPU, EDU, 5 Years
Table 8. NVIDIA Run: AI
Part number Description
NVIDIA Run:ai for use bundled with RTX PRO 4500 Blackwell Server Edition GPU
7S020072WW NVIDIA Run:ai for RTX PRO 4500 Blackwell Server Edition, per GPU, Self-Hosted, 3Years
7S020073WW NVIDIA Run:ai for RTX PRO 4500 Blackwell Server Edition, per GPU, Self-Hosted, 5Years
7S020070WW NVIDIA Run:ai for RTX PRO 4500 Blackwell Server Edition, per GPU, SaaS, 3Years
7S020071WW NVIDIA Run:ai for RTX PRO 4500 Blackwell Server Edition, per GPU, SaaS, 5Years
NVIDIA Run:ai for use bundled with NVIDIA Run:ai for RTX PRO 6000 Blackwell Server Edition GPU
7S020065WW NVIDIA Run:ai for RTX PRO 6000 Blackwell Server Edition, per GPU, Self-Hosted, 3 Years
7S020066WW NVIDIA Run:ai for RTX PRO 6000 Blackwell Server Edition, per GPU, Self-Hosted, 5 Years
7S020067WW NVIDIA Run:ai for RTX PRO 6000 Blackwell Server Edition, per GPU, SaaS, 3 Years
7S020068WW NVIDIA Run:ai for RTX PRO 6000 Blackwell Server Edition, per GPU, SaaS, 5 Years
Software subscription
7S02004UWW NVIDIA Run:ai Subscription per GPU 1 Year
7S02004XWW NVIDIA Run:ai Subscription per GPU 3 Years
7S020050WW NVIDIA Run:ai Subscription per GPU 5 Years
7S02004VWW NVIDIA Run:ai Subscription per GPU EDU 1 Year
7S02004YWW NVIDIA Run:ai Subscription per GPU EDU 3 Years
7S020051WW NVIDIA Run:ai Subscription per GPU EDU 5 Years
7S02004WWW NVIDIA Run:ai Subscription per GPU INC 1 Year
7S02004ZWW NVIDIA Run:ai Subscription per GPU INC 3 Years
7S020052WW NVIDIA Run:ai Subscription per GPU INC 5 Years
Support Services subscription
7S020053WW 24x7 Support Services for NVIDIA Run:ai Subscription per GPU 1 Year
7S020056WW 24x7 Support Services for NVIDIA Run:ai Subscription per GPU 3 Years
7S020059WW 24x7 Support Services for NVIDIA Run:ai Subscription per GPU 5 Years
7S020054WW 24x7 Support Services for NVIDIA Run:ai Subscription per GPU EDU 1 Year
7S02005AWW 24x7 Support Services for NVIDIA Run:ai Subscription per GPU EDU 3 Years
7S020057WW 24x7 Support Services for NVIDIA Run:ai Subscription per GPU EDU 5 Years
7S020055WW 24x7 Support Services for NVIDIA Run:ai Subscription per GPU INC 1 Year
7S020058WW 24x7 Support Services for NVIDIA Run:ai Subscription per GPU INC 3 Years
7S02005BWW 24x7 Support Services for NVIDIA Run:ai Subscription per GPU INC 5 Years

Note: SKUs marked ‘INC’ are part of NVIDIA’s Inception Program, a global initiative supporting AI startups.
Learn more at: https://www.nvidia.com/en-us/startups/

Table 9. NVIDIA Mission Control
Part number Description
NVIDIA Mission Control with Business Critical Support
7S02004KWW NVIDIA Mission Control SW Subscription per GPU per Year With Business Critical Support 3 Years
7S02004LWW NVIDIA Mission Control SW Subscription per GPU per Year With Business Critical Support 4 Years
7S02004MWW NVIDIA Mission Control SW Subscription per GPU per Year With Business Critical Support 5 Years
7S02004NWW NVIDIA Mission Control SW Subscription per GPU per Year With Business Critical Support EDU 3 Years
7S02004PWW NVIDIA Mission Control SW Subscription per GPU per Year With Business Critical Support EDU 4 Years
7S02004QWW NVIDIA Mission Control SW Subscription per GPU per Year With Business Critical Support EDU 5 Years
7S02004RWW NVIDIA Mission Control SW Subscription per GPU per Year With Business Critical Support INC 3 Years
7S02004SWW NVIDIA Mission Control SW Subscription per GPU per Year With Business Critical Support INC 4 Years
7S02004TWW NVIDIA Mission Control SW Subscription per GPU per Year With Business Critical Support INC 5 Years
NVIDIA Mission Control with Standard Support
7S02004AWW NVIDIA Mission Control SW Subscription per GPU per Year With Std Support 3 Years
7S02004BWW NVIDIA Mission Control SW Subscription per GPU per Year With Std Support 4 Years
7S02004CWW NVIDIA Mission Control SW Subscription per GPU per Year With Std Support 5 Years
7S02004DWW NVIDIA Mission Control SW Subscription per GPU per Year With Std Support EDU 3 Years
7S02004EWW NVIDIA Mission Control SW Subscription per GPU per Year With Std Support EDU 4 Years
7S02004FWW NVIDIA Mission Control SW Subscription per GPU per Year With Std Support EDU 5 Years
7S02004GWW NVIDIA Mission Control SW Subscription per GPU per Year With Std Support INC 3 Years
7S02004HWW NVIDIA Mission Control SW Subscription per GPU per Year With Std Support INC 4 Years
7S02004JWW NVIDIA Mission Control SW Subscription per GPU per Year With Std Support INC 5 Years

Note: SKUs marked ‘INC’ are part of NVIDIA’s Inception Program, a global initiative supporting AI startups.
Learn more at: https://www.nvidia.com/en-us/startups/

Energy monitoring and optimization

The following Energy monitoring and optimization software is available with Lenovo EveryScale HPC & AI Software Stack.

  • Energy Aware Runtime (EAR)
    Energy Aware Runtime (EAR) is a system-level energy management software solution originally developed by the Barcelona Supercomputing Center (BSC) and the Universitat Politècnica de Catalunya (UPC). The software solution is now developed, commercialized, delivered, and supported by the EAR team, operating under the corporate name Energy Aware Solutions (EAS).

    EAR helps organizations monitor, analyze, and optimize energy consumption across HPC and AI environments. The solution supports both CPU- and GPU-based systems and integrates with popular workload managers such as Slurm, PBS Pro, and Kubernetes.

    EAR combines open-source technologies with advanced commercial capabilities to provide a comprehensive platform for energy monitoring, optimization, accounting, and power management. The solution is available through Lenovo and is delivered with installation services, training, and support from the EAR team.
  • EAR Product Options
    EAR offers three products that can be combined to meet customer requirements.

    • EAR Analytics (Required)
      EAR Analytics is the base module required for all EAR deployments. It provides the foundation of the EAR solution and includes:

      • Cluster-wide energy and power monitoring

      • CPU and GPU energy accounting

      • Job-level energy consumption reporting

      • Data analysis and visualization through Grafana dashboards

      • Job- and system-wide energy analytics and reporting

      • EAR support services

    • EAR Optimizer (Optional)
      EAR Optimizer extends EAR Analytics by adding intelligent GPU energy optimization policies that dynamically improve the performance-per-watt ratio of GPU workloads, helping organizations reduce energy consumption and operating costs while maximizing system efficiency. This module is particularly valuable for GPU-accelerated HPC and AI environments where energy consumption represents a significant operational cost.

      Note: EAR Optimizer is focused on GPU optimization capabilities. Additional CPU optimization services may be available upon request and are provided separately from the standard EAR Optimizer offering.

    • EAR Smart Power Cap (Optional)
      EAR Smart Power Cap extends EAR Analytics with advanced power management capabilities that maximize application performance within a specified power budget. This module enables organizations to:
      • Operate systems within variable facility power constraints
      • Improve overall system utilization
      • Reduce power peaks across the cluster
      • Optimize performance per watt
  • Deployment Configuration
    Customers can select one of the following configurations:
    • EAR Analytics
    • EAR Analytics + EAR Optimizer
    • EAR Analytics + EAR Smart Power Cap
    • EAR Analytics + EAR Optimizer + EAR Smart Power Cap.

Trial Program
Customers may evaluate EAR through a Trial Program that provides temporary access to selected EAR capabilities for validation purposes. The trial enables organizations to assess energy monitoring, reporting, and optimization features in their own HPC or AI environment before making a purchasing decision. Trial engagements are typically limited to four weeks.

Trial deployments are supported by the EAR team. EAR provides the system software either as an RPM package or as a containerized deployment, together with the required documentation, allowing customers to perform the installation independently. Following deployment, the EAR team conducts weekly meetings with the customer to answer questions, provide guidance, and discuss observed results.

Ordering Information
EAR products can be purchased through Lenovo and are delivered and supported by the EAR team.
To prepare a quotation, the following information is required:

  • The desired EAR configuration.
  • The total number of compute nodes eligible to execute workloads (for example, all compute nodes defined in the Slurm configuration).
  • The number of CPU-only nodes and the associated number of CPU sockets.
  • The number of GPU nodes and the associated number of GPU sockets (CPU sockets in GPU nodes are not counted).

There are no predefined limits on the number of supported nodes or GPU sockets.
A quotation from EAR is required for all EAR deployments.

Resources

For more information, see these resources:

Related product families

Product families related to this document are the following:

Trademarks

Lenovo and the Lenovo logo are trademarks or registered trademarks of Lenovo in the United States, other countries, or both. A current list of Lenovo trademarks is available on the Web at https://www.lenovo.com/us/en/legal/copytrade/.

The following terms are trademarks of Lenovo in the United States, other countries, or both:
Lenovo®
ThinkSystem®

The following terms are trademarks of other companies:

AMD is a trademark of Advanced Micro Devices, Inc.

Intel®, the Intel logo is a trademark of Intel Corporation or its subsidiaries.

Linux® is the trademark of Linus Torvalds in the U.S. and other countries.

VS Code® is a trademark of Microsoft Corporation in the United States, other countries, or both.

NVIDIA®, CUDA®, NVIDIA Omniverse®, NVIDIA RTX®, and OpenACC® are trademarks of NVIDIA Corporation.

Other company, product, or service names may be trademarks or service marks of others.