【WAIC】Jensen Huang “Sells GPUs,” Lisa Su “ Breaks Down Barriers ”: AMD Competes for China’s Developer Mindshare

About Infinigence-AI

【WAIC】On May 13, Jensen Huang made a last-minute trip to Beijing aboard Air Force One alongside Donald Trump. Six days later, Lisa Su took the stage in Shanghai and asked an audience of more than 2,000 developers a simple question: “Are you excited?”

Within one week, two of the world’s most influential semiconductor leaders visited China in succession—delivering sharply different narratives.Huang’s visit was framed around market access and the boundaries of what could be sold. Su’s message, by contrast, focused on what could be built next.From the stage, she noted: “China has the world’s most dynamic AI ecosystem.”

Beyond competition in hardware shipments, AMD’s deeper ambition lies elsewhere: securing long-term influence over the choices Chinese developers make in shaping the next phase of AI computing.


The Only Question in the Room: Can GPUs still keep up?

 

This was not just an AMD showcase—it felt more like a collective reckoning by 2,000 developers, aimed at one of the AI industry’s biggest assumptions of the past two years: that more compute automatically means better outcomes.

Kai-Fu Lee, CEO of 01.AI, was the first to challenge that assumption.“If your AI deployment doesn’t change the numbers on your quarterly financial reports, you’re not transforming your business—you’re just running an AI lab.”

 

He argued that many enterprises are still stuck in what he called “cosmetic AI”—deploying HR chatbots, internal search tools, and customer service automation that improve interfaces but not outcomes.His advice to developers was blunt: stop building superficial features, and start building systems that structurally drive revenue and productivity.But his most striking argument was about agents. A single agent, he said, is not enough. The future belongs to “agent committees”—systems where multiple agents specialize in planning, execution, risk control, and evaluation, continuously debating, coordinating, and iterating with one another.

 

If one person can eventually manage 5, 10, or even 100 agents capable of handling nearly all digital tasks, then the traditional logic of simply scaling compute begins to look incomplete.To this question, Jack Huynh, AMD’s Senior Vice President and General Manager of Computing and Graphics, offered a hardware-centric reframing that shifted the room’s attention:“Compute remains critical in the prefill stage, but memory is becoming the new bottleneck—and the new center of gravity.”

 

AMD’s Answer: Develop, Test, Deploy—Start Locally First

 

AMD’s approach breaks AI development into three stages, with a dedicated hardware strategy for each phase.“This isn’t about replacing the cloud; it’s about complementing it,” said Jack. In this model, development and testing happen locally, where token costs effectively approach zero. Only the final deployment stage moves to the cloud. Today, more than 35 Ryzen AI Max+ agent-optimized PCs have already been launched by HP, ASUS, Lenovo, Acer, and emerging local brands, spanning laptops, all-in-one systems, and mini workstations.


Kai-Fu Lee says “Don’t burn cash”; Jack’s response is more operational: “I’ll help you save it”The implication is straightforward: on-device inference efficiency is beginning to reshape the cost structure of AI development itself. Agents no longer need to start in the cloud to get real work done—they can begin locally.

 

Data Closed Loop: Running a 200B Model at 100 Tokens/s on a Laptop

 

Zhu Yibo, CTO of StepFun, demonstrated with Step 3.5 Flash: “this really works.”He showed that Step 3.5 Flash—a 200B-parameter model designed for agentic workloads—can achieve near 100 tokens/s decoding speed on a Ryzen AI Max+ 395 laptop with 128GB of unified memory, after 4-bit quantization. In certain scenarios, this performance even surpasses that of many cloud-based models.


His conclusion was straightforward:“Everyone will have multiple personal agents, and at least one of them will run locally on a PC.”This signals a clear shift in where agents are being deployed—from cloud-centric architectures toward local devices. As hardware capabilities and model efficiency converge, a new question emerges naturally:


What will the software ecosystem look like in this new on-device agent era?

 

Ecosystem Self-Evolution: AMD Is Using AI to Write Its Own Code

 

Nick Ni, Senior Director of AMD’s Artificial Intelligence Group, highlighted two notable developments on site.First, AMD is now using AI agents to optimize its own open-source ecosystem. Thousands of AI agents continuously monitor open-source projects, identify gaps in ROCm support, generate pull requests, and run validation tests. As a result, engineers have shifted from producing “a few pull requests per week” to “several per day.” In a global ROCm performance competition, the winning teams reportedly relied heavily on AI-assisted tooling, achieving performance improvements of more than 2×.


Second, AMD has launched its first free Radeon GPU public cloud service tailored for Chinese AI developers.Users can register and immediately access AMD GPUs within the ModelScope ecosystem, without the need to install drivers or configure local environments.Together, these initiatives reflect a dual-track strategy: accelerating ROCm’s maturity through AI-driven automation, while simultaneously lowering the barrier to entry for developers through free cloud access.That said, ROCm’s open ecosystem is still in the process of catching up with CUDA.


NVIDIA’s platform benefits from a 20-year head start and an ecosystem of more than four million developers. Industry reports suggest that NVIDIA hardware still accounts for roughly 90% of citations in AI research papers. More importantly, switching costs are driven less by hardware itself and more by software migration complexity—existing codebases are not easily portable in a “one-click” manner.

 

Still, AMD’s bet is clear: in the emerging era of on-device inference and edge agents, developers want optionality. In this context, “free access” and “openness” may become the most effective entry points into the ecosystem.

 

Endgame Scenario: AI Must Move from Screens to the Physical World

 

Wang Yu, Professor in the Department of Electronic Engineering, Tsinghua University and founder of Infinigence, closed the conference while also opening up its next frontier.He introduced a simple but powerful productivity identity for AI:

 

AI Productivity = Intelligence Scale× Token Efficiency × Value Conversion

 

The first two variables had already been demonstrated throughout the day. Scale of intelligence is being driven by multi-agent systems. Token efficiency is increasingly achieved through on-device inference and local execution.The remaining term—value conversion—defines the real boundary ahead: AI must generate measurable impact in the physical world, not just within digital interfaces.


For Wang, this shift implies that the next stage is not simply larger models, but world models. Instead of reacting to the physical world, AI systems should first simulate it internally—building a “mental model” of physical dynamics before acting in reality. In theory, this approach can accelerate learning efficiency by orders of magnitude.


To advance this direction, Wang Yu and AMD co-developed the open-source framework Olive, focused on physical AI workloads. Within four months of release, the project has accumulated over 3,300 GitHub stars, been adopted by more than 20 leading companies, and is fully integrated with AMD’s ROCm stack.


By the end of the conference, a clear narrative arc had emerged—from digital intelligence to physical embodiment: Kai-Fu Lee emphasized that agents must deliver real economic value. Jack Huynh argued that local devices are already capable of running them. Zhu Yibo demonstrated that large models can run efficiently on laptops. Nick Ni showed how ecosystems are beginning to self-optimize. And Wang Yu ultimately pointed to the next frontier: AI stepping into the physical world.


Holding Onto Today vs. Securing Tomorrow


So we return to the opening question: what are Jensen Huang and Lisa Su really competing over?
Jensen Huang is fighting for today’s revenue—whether H20 can still be shipped, and whether CUDA can continue to lock in customers. It is, fundamentally, a battle over installed base and existing market share.
Lisa Su, by contrast, is competing for tomorrow’s developers—pushing ROCm open source, offering free GPU cloud access, enabling direct integration through the ModelScope ecosystem, and building an end-to-end pipeline that spans cloud, edge, and on-device inference. This is a strategy oriented toward expanding the future ecosystem.

But in the era of AI agents, the real competition is no longer about who burns more compute—it is about who uses it more intelligently.

As Jack Huynh put it:
“The hardware is already in place, the software ecosystem is already in place, and China has one of the world’s strongest open-source AI communities.” What remains missing is the connective layer—the ability to integrate these pieces into a coherent system.
And that missing piece may already be in the room, among the 2,000 developers gathered in Shanghai.


Infinite Computing, Accelerating the Future of AGI

Contact us for customized AI infrastructure solutions

Infinite Computing, Accelerating the Future of AGI