End-to-end Capabilities: Inference Job Startup, Inference Optimization, Elasticity & Cache Acceleration

Designed for two core enterprise inference scenarios:
real-time online LM inference and asynchronous AIGC generation

P/D-Separated Architecture

P/D-Separated Architecture

Multimodal Batch Optimization

Multimodal Batch Optimization

Minute-level Hundred-Instance Elastic Scaling

Minute-level Hundred-Instance Elastic Scaling

Mixed-Length Texts

Stable Operation & Tail Latency Control

Stable Support

High Concurrency & Bursty Production Traffic

Scalable

Delivery for generation tasks

Capabilities

P/D Separation Inference
Async Inference & Workflow Opt
Elastic Scaling & Smooth Upgrades
Production Observability & Data Infra
P/D Separation Inference
P/D Separation Inference

Distributed inference service: Native support for fast, one-click service deployment across multiple nodes, with flexible extension options.

P/D separation architecture: Decoupled deployment of prefill and decode instances with an automatic routing mechanism, enabling adaptive control of P/D resource allocation and cluster-level scaling based on workload characteristics.

MoE inference-specific optimizations: Optimizations for multi-node, multi-GPU parallelism, communication, and routing overhead, improving both throughput and system stability.

Async Inference & Workflow Opt
Async Inference & Workflow Opt

Hosted workflows such as ComfyUI: Automated task-based and queue-based execution with support for asynchronous generation.

Batch high-concurrency processing: Supports automatic batch processing, concurrency control, and priority scheduling.

Multi-layer caching acceleration: Supports caching acceleration for critical dependencies such as images and models, reducing latency and jitter caused by cold starts and repeated loading.

Elastic Scaling & Smooth Upgrades
Elastic Scaling & Smooth Upgrades

Multiple scaling strategies: Configurable scaling policies based on manual control, scheduled actions, real-time workload signals, and custom rules.

Smooth upgrade capabilities: Support for canary releases, rolling upgrades, and version management, ensuring service continuity and minimizing operational risk during upgrades.

Ultimate elasticity: Ability to scale hundreds of instances within minutes, delivering maximum elasticity and optimal cost efficiency across workload peaks and valleys.

Production Observability & Data Infra
Production Observability & Data Infra

Comprehensive log collection: End-to-end logging and traceability across request ingress, routing and scheduling, prefill and decode stages, as well as model-side and resource-side components.

Metrics monitoring and alerting: Full observability into latency, throughput, queue depth, error rates, GPU utilization, and GPU memory usage.

Data-driven auto-scaling: Closed-loop scaling decisions driven by multi-dimensional signals such as concurrency, queue length, latency, and resource utilization.

Multi-strategy request routing: Policy-driven request distribution supporting load balancing, parallelism balancing, and cache-aware routing to meet different business objectives.

Why Us

Use Case

An AIGC Startup

The company provides multiple AI-powered image generation services online. Platform traffic fluctuates sharply with trending topics, re

An AIGC Startup
Before Optimization

Slow response to traffic spikes : Sudden surges driven by trending topics frequently exceeded preprovisioned compute capacity. When QPS rose above 200, hundreds of thousands of new users were lost per minute due to degraded user experience.

Severe performance degradation: Under heavy request loads, inference performance dropped by 30%. Single-task latency exceeded 20 seconds, causing a sharp increase in paid user churn.

Inefficient scaling: Adding temporary compute capacity required reconfiguring environments and services, resulting in slow iteration cycles that could not keep pace with traffic volatility.

After Optimization:

Stable handling of peak traffic: During peak events, the platform handled up to 200× normal traffic with zero service interruptions, fully meeting image generation demand.

Doubled inference speed: Single-task inference was reduced from over 20 seconds to under 10 seconds, cutting user bounce rates by more than 70%.

5-minute rapid integration: New workflows were integrated and deployed within 5 minutes using existing environments and model assets on the platform.

Significant cost reduction: During peak periods, per-user operational costs dropped by more than 50%, driving sustained revenue growth.

A Large-Model Unicorn Company

The company deployed a large-scale MoE model with tens of billions of parameters to provide unified inference services across multiple business lines. With pronounced traffic peaks and valleys and a mix of request lengths, the platform needed to deliver high concurrency stability while maintaining cost efficiency.

A Large-Model Unicorn Company
Before Using the Platform:

Inefficient manual scaling: Resource scaling relied on 24/7 manual monitoring and adjustments, with complex multi-node configurations, creating a heavy operational burden.

Long-tail requests impacting performance: Mixed short and long requests caused certain long-context workloads to slow down overall response times, leading to frequent user complaints.

High cost pressure: High concurrency was supported through extensive redundant capacity, while long-context inference requests drove up costs. Scaling strategies remained coarse-grained and inefficient.

After Using the Platform:

Automatic elastic scaling:Business-driven auto-scaling strategies with one-click deployment and PD configuration reduced manual operations effort by 80%.

Precise request bucketing: Requests are bucketed based on input length and cache hit patterns, ensuring both high throughput and consistently low response latency.

Accelerated long-text inference: First-token latency was reduced by 1 second, overall throughput increased by more than 2×, and token-level cost-performance improved by 3×, resulting in a healthier and more sustainable production inference operation.

Infinite Computing, Accelerating the Future of AGI

Contact us for customized AI infrastructure solutions

Infinite Computing, Accelerating the Future of AGI