Virtualization and cloud platforms / Compute infrastructure

GPU and
High-Performance Computing

Start with a model, simulation or dataset, coordinating compute, memory, storage and interconnects so valuable resources run workloads rather than wait for data or compete for capacity.

Workload assessment · CPU / GPU · Cluster scheduling · High-speed interconnects

Measure a compute platform by how well it completes work, not merely by GPU count.

What high-performance computing means

Coordinate complex computation
across the right resources

HPC uses powerful nodes or coordinated multi-node systems for tasks beyond ordinary office hardware. GPUs suit adapted parallel algorithms, but do not automatically accelerate every application. CPUs, memory and data paths also determine suitability.

TRAIN / INFER

AI training and inference

Training emphasizes model size, GPU memory and node communication. Online inference also requires concurrency, response time and service reliability; the same selection metrics do not apply to both.

SIMULATE

Engineering simulation and scientific computing

Choose CPU, GPU or mixed configurations by software support and solver behavior. Parallel licensing, memory requirements and numerical accuracy affect usable configurations.

PROCESS / RENDER

Data processing and rendering

Workload partitioning, file sharing and I/O intensity determine whether to add compute nodes or improve storage and queue management first.

Workload execution architecture

Scheduling allocates resources,
compute nodes access data directly

Users submit jobs and resource requirements. Scheduling assigns nodes under queue policies, while compute nodes access storage through the data network. The scheduler manages jobs; it is not the transit point for all workload data.

Job control and data paths / Two coordinated relationships
Submit

Users and job submission

Scripts / Parameters / Data paths
Request CPU, GPU, memory and runtime

Allocate

Queues and scheduling

Permissions / Quotas / Priorities
Select Slurm or other schedulers as needed

Execute

CPU / GPU nodes

Matched drivers, libraries and applications
Run jobs and report status

Data path

Shared / Parallel storage

Input data, model checkpoints and results

Compute-node reads and writes

Multi-node jobs also need suitable node interconnects

Complete job → Check results and logs → Save outputs → Release resources

Resource records support capacity and cost analysis. Retry and checkpoint recovery depend on application capabilities and scheduling policies.

The diagram illustrates queued jobs. Online inference, interactive development and container services may use different entry points and schedulers. Animation shows relationships, not measured bandwidth or speedup.

GPU and HPC solution overview: software platforms, CPU/GPU resources, high-performance storage and networking
Conceptual overview. GPU models, software, schedulers, network specifications and ecosystem brands are examples. Final configurations depend on workloads, compatibility, licensing and facility conditions, not a fixed supply list or performance commitment. Open the full image for details.

Balanced configuration

Do not let one bottleneck
hold back the entire platform

First confirm that the workload runs, then compare efficiency. Insufficient GPU memory, inter-node communication waits or inadequate data delivery can prevent high-specification hardware from performing effectively.

Compute compatibility
Check supported architectures, numerical precision, drivers and libraries instead of comparing peak compute alone.
Memory capacity
Calculate GPU and system memory for models, batches, caches and workspace, not just file size.
Communication throughput
Assess single-node or multi-node parallelism and communication volume before choosing interconnects. Network specifications are not workload speedup figures.
Data delivery
Plan storage around read patterns, file counts and checkpoint writes to avoid prolonged compute waits.
Exhibit photo of an NVIDIA HGX B200 multi-GPU board and cooling structure
Multi-GPU hardware reference, not a project configuration or Yuqi Intelligent case study. Photo: Pokiiri / Wikimedia Commons,CC BY-SA 4.0; existing resized version retained without further modification.

Choose the operating model

Compute resources need not
all be used in the same way

Workload typeOperating modelPrimary considerations
Periodic simulation / Batch trainingQueue jobs; allocate and release resources per taskCompletion time, fair scheduling, failure diagnosis and reproducibility
Interactive research / DevelopmentControlled development sessions, shared environments or dedicated nodesEnvironment consistency, user isolation, resource use and data permissions
Online inference / API servicesContinuously running services, optionally using container orchestrationConcurrency, tail latency, releases and availability

Validate the platform with real workloads

Before handover,
complete a real computation

Agree reproducible samples and acceptance conditions before installation, tuning and expansion. Record software versions, input size, node count and resource configuration with performance results to avoid comparing unlike measurements.

  1. 01

    Single-job baseline

    Confirm correctness, runtime and resource use, identifying environment or application compatibility issues.

  2. 02

    Concurrency and scaling

    Test multi-user sharing, multi-GPU nodes and multi-node behavior, identifying communication and storage waits.

  3. 03

    Sustained operation

    Observe temperatures, power, alerts and failure handling; validate checkpoint recovery where supported by the application.

Frequently asked questions

Confirm these conditions before selection

Does HPC always require GPUs?

No. Some simulation, data processing and scientific workloads depend on CPUs and large memory. Only GPU-adapted algorithms can use GPU parallelism. Check software support, licenses and workloads before selecting CPU, GPU or mixed systems.

Does adding GPUs reduce runtime proportionally?

Not necessarily. Speedup depends on parallelism, GPU memory, node communication, data reads and software implementation. Compare real models or simulations on single GPUs, single nodes and multiple nodes using effective runtime and utilization.

How should Slurm and Kubernetes be chosen?

Assess Slurm or similar schedulers for queued jobs, reservations and batch work. Assess Kubernetes and required extensions for container services and application lifecycle. They are not automatically chained together; workload and operations determine the choice.

What information is needed before deployment?

Provide software and versions, sample jobs, data volume, concurrency, runtime or latency objectives, and facility power, cooling and networking conditions. Sanitized samples are sufficient; handle sensitive data under the agreed process.

Compute-platform implementation and coordination

Yuqi Intelligent selects and supplies CPU/GPU servers, storage and interconnects, deploys clusters, configures drivers/software, integrates scheduling and tests workloads. Application licensing, algorithm changes and specialist solver results remain subject to the agreed responsible parties.

View implementation scope
  • Assess software, samples, data scale and concurrency; plan compute nodes, storage, networking and facility support.
  • Deploy equipment, configure drivers/dependencies, queues and permissions, and integrate monitoring and resource records.
  • Validate correctness and operating performance with agreed samples; hand over environments, procedures and maintenance methods.
View project records and handover
  • Hardware/software configuration, compatible versions and topology
  • Users, queues, quotas and data-access rules
  • Submission examples, environment documentation and validation records
  • Alerts, maintenance, training and expansion guidance

Next step / Technical discussion

Start configuring compute around one workload.

Bring software versions, data scale, concurrency and target runtime improvements to identify the actual bottleneck.

Talk to a technical adviser