skip to main content

Lenovo LLM Sizing Tool

Tool

Home
Top
Updated
7 Oct 2026
Form Number
LP2487

Introduction

This sizing tool helps you determine the hardware and infrastructure required to deploy large language models for a wide range of AI workloads. Whether you're building a chatbot, Retrieval-Augmented Generation (RAG) application, AI agent, code assistant, or another generative AI solution, this tool helps you estimate the hardware resources needed to meet your performance and scalability requirements.

The tool includes two sizing options:

  • Calculator: Sizes GPU-based deployments for your model and workload.
  • CPU Sizing: Estimates capacity and cost for running inference on CPU-only Lenovo servers.

Version 2.0. Click the Full Change History link to see what's new.

Getting started

Scroll down to use the tool. To get started with the GPU calculator:

  1. Select the Use Case
  2. Select the desired GPU Type
  3. Click the Calculate GPU Requirements button

The tool automatically populates recommended preset values for the model and workload parameters, including concurrency, input and output token lengths, precision, and other deployment settings, which you can adjust as needed. Based on these inputs, it analyzes your workload and recommends an optimal infrastructure configuration. Learn more.

Use the Subscribe to Updates link to get notified when the tool gets updated, and use the Feedback link to send us comments and suggestions.

Sizing Tool

Using the sizing tool

The tool includes two sizing options:

  • Calculator: Sizes GPU-based deployments for your model and workload.
  • CPU Sizing: Estimates capacity and cost for running inference on CPU-only Lenovo servers.

Using the Calculator tool

Get started with these steps:

  1. Select the Use Case
  2. Select the desired GPU model
  3. Click the Calculate GPU Requirements button

The tool automatically populates recommended preset values for the model and workload parameters, including concurrency, input and output token lengths, precision, and other deployment settings, which you can adjust as needed. Based on these inputs, it analyzes your workload and recommends an optimal infrastructure configuration.

Custom use cases and models:

  • Custom use case: Select Custom as the use case and set your own values under Advanced Parameters.
  • Custom model: Select Custom as the model and enter its HuggingFace model ID in the Custom Model Name field (for example, mistralai/Mixtral-8x7B-v0.1), then press Enter. The model details are filled in automatically and can be reviewed under Advanced Parameters.

The Calculator results include:

  • Recommended GPU count and a detailed VRAM breakdown
  • Server platform and a recommended Lenovo Hybrid AI platform
  • CPU and system memory requirements
  • Storage, network adapters, and switching infrastructure
  • Estimated power usage
  • Inference performance metrics, including Time to First Token (TTFT), Inter-Token Latency (ITL), end-to-end request latency, throughput, and maximum supported concurrent users

Results can be downloaded as a PDF report, enabling you to evaluate different deployment scenarios before provisioning hardware.

Using the CPU Sizing tool

For CPU-based inference, open the CPU Sizing tab:

  1. Select a business use case
  2. Enter the number of concurrent users, active hours per day, server lifecycle, and electricity cost
  3. Review the recommended Lenovo servers and CPUs

For each option, the tool shows the supported capacity and the cost per million tokens compared with running the same workload in the cloud.

Use this tool to compare infrastructure options, optimize resource utilization, and make informed decisions when designing scalable and cost-effective AI deployments on Lenovo infrastructure. All estimates are based on analytical models and benchmark data and should be used as guidance for capacity planning rather than guaranteed production performance.

Use the Subscribe to Updates link to get notified when the tool gets updated, and use the Feedback link to send us comments and suggestions.

Related product families

Product families related to this document are the following: