Integrated Optimization from Communication to Fault-tolerant Recovery

Designed for large-scale model training scenarios at the 1,000 to 10,000-GPU scale

Deep optimization for distributed computing

Deep optimization for distributed computing

Asynchronous Checkpoint saving

Asynchronous Checkpoint saving

Intelligent topology-aware scheduling

Intelligent topology-aware scheduling

Visualized monitoring and automated alerting

Visualized monitoring and automated alerting

30%

End-to-end training efficiency improvement

Minute-level

Fault Recovery Time

Efficient & Stable

Training at scale

Capabilities

Distributed Training Optimization and Acceleration
Full Fault Tolerance
Visual monitor for training
Native Support for Cutting-Edge Algorithms
Distributed Training Optimization and Acceleration
Distributed Training Optimization and Acceleration

Multi-node collaborative acceleration: Supports multi-dimensional optimizations including collective communication, communication topology, and pipeline parallelism to improve cross-node and cross-GPU efficiency.

Data loading optimization: Eliminates data I/O bottlenecks through data pipelines, sampling strategies, distributed data loading, and other methods.

Startup and scheduling optimization: Supports preloading and caching of images and models, priority-based queues, and elastic training to reduce startup time and scaling delays for large jobs.

Full Fault Tolerance
Full Fault Tolerance

Asynchronous checkpointing: Enables faster iteration and rapid rollback through asynchronous checkpointing, combined with distributed sharding and multi-version management.

Multi-level anomaly detection and recovery: Supports detection of a wide range of hardware and job anomalies, including overheating, GPU failures, slow GPUs, and hanging jobs, with automatic restart and fault recovery.

Automatic task migration: Supports sub-second failover during hot-standby node failures, cross-node migration, and elastic fault tolerance for partial failures, minimizing training losses and reducing manual intervention.

Visual monitor for training
Visual monitor for training

Real-time metric tracking: Provides clear visibility into more than 50 key metrics, including core training metrics, performance metrics, and resource utilization.

Transparent training process: Provides multidimensional training dashboards with logs and key-event timelines.

Automated anomaly alerts: Supports intelligent alert configuration for anomalies in loss, performance, and resource issues, with severity levels and real-time notifications across multiple channels.

Native Support for Cutting-Edge Algorithms
Native Support for Cutting-Edge Algorithms

Reinforcement learning scenario optimization: Natively supports one-click setup and launch for leading RL frameworks, plus multi-role parallel training and multiple update strategies.

Integrated simulation and training: Comes with built-in support for multiple leading simulators, enabling multi-dimensional Sim-to-Real optimization and multimodal alignment.

Native support for multiple parallelization strategies: Supports data, model, pipeline, and tensor parallelism, with automatic parallel strategy search. Provides expert parallelism optimization and accelerated expert routing communication for MoE models.

Why Us

Use Case

A Large Research Laboratory

The customer needed to carry out pre-training, post-training, and reinforcement learning for a large language model with hundreds of billions of parameters on a cluster of several thousand GPUs.

A Large Research Laboratory
Before Optimization

Startup timeouts: Training jobs frequently failed to complete initialization within the NCCL communication timeout threshold of 60 minutes, triggering repeated fault-tolerant restarts and making it difficult to start training successfully.

Execution freezes: At the 2,000-GPU scale, 60%–70% of jobs froze after running for around one hour, preventing further progress and severely reducing training efficiency and resource utilization.

Performance instability: At the 1,000-GPU scale, training throughput fluctuated significantly, with TFLOPS varying between 180 and 330. This made it difficult to accurately estimate training completion timelines.

Schedule risk: Model iteration cycles exceeded one month, resulting in low production efficiency.

After Optimization

Faster initialization: Data initialization time was reduced from 60 minutes to under 15 minutes, effectively eliminating NCCL timeout issues and increasing training startup success rates to nearly 100%.

Highly stable operation: At the 2,000-GPU scale, jobs run continuously and stably for over two weeks without interruption or freezes. Throughput remained steady, and resource utilization was maximized.

Stable training performance: At the 1,000-GPU scale, training throughput remained stable at 330 TFLOPS with no sudden slowdowns, enabling precise control over iteration cycles.

Shorter iteration cycles:With the same dataset and cluster size, iteration cycles were reduced to just over 20 days, allowing the model training objectives to be completed ahead of schedule.

An Embodied AI Unicorn Company

The customer needed to deliver usable policy versions on a fixed timeline. Its R&D workflow spanned multimodal large-scale model training, offline replay and evaluation, as well as reinforcement learning and simulation-based training within a realdevice data closed loop. Training workloads were large and highly concurrent, placing high demands on distributed efficiency, stability, and observability.

An Embodied AI Unicorn Company
Before Using the Platform

Unpredictable startup and scaling: ThousandGPUscale training jobs depended on image and model pulls, data initialization, and resource queuing. Startup and scale-out wait times were unpredictable, affecting iteration schedules.

High barriers to RL/simulation experiments: Environment setup and scaling costs were high. Configuration of environments, parallelization strategies, update mechanisms, and Sim-to-Real alignment parameters was complex. Experiments were difficult to reproduce and scale, resulting in low convergence efficiency.

Delayed anomaly detection: Reinforcement learning pipelines involved many engineering details but lacked unified monitoring, event timelines, and tiered alerting. Training interruptions or freezes often surfaced only after several days of execution, leading to long troubleshooting cycles.

High delivery risk: Long model iteration cycles and uncertain release timelines increased delivery risk and reduced investor confidence.

After Using the Platform

Fast startup and stable throughput: Startup and scale-out wait times for large jobs were significantly reduced. Cross-node efficiency improved, data I/O bottlenecks were alleviated, training throughput became stable, and delivery timelines became predictable.

Scalable RL simulation: Configuration and reproducibility costs were significantly reduced, allowing teams to focus on training strategies and algorithm optimization. RL and simulation experiments could run concurrently at scale, accelerating convergence.

Automatic fault recovery: Training interruptions and freezes were greatly reduced. Failures could be automatically recovered or migrated, minimizing manual intervention. Performance degradation and loss anomalies were detected early, avoiding ineffective long-running jobs.

End-to-end observability: Training and evaluation workflows moved from slow startup, frequent interruptions, and unstable throughput to fast startup, stable execution, full observability, and rollback support.

Infinite Computing, Accelerating the Future of AGI

Contact us for customized AI infrastructure solutions

Infinite Computing, Accelerating the Future of AGI