China’s domestic artificial intelligence (AI) ecosystem is gaining momentum as major models demonstrate the ability to train and run on locally developed chips, while Nvidia supply constraints and Beijing’s restrictions on foreign processors accelerate the shift towards domestic hardware, Maybank Investment Bank said on Monday.

Meituan’s LongCat-2.0, released on June 30, became the first publicly confirmed trillion-parameter model to complete full-process training and inference on a 50,000-chip domestic computer cluster, according to the research report.

The milestone challenges the view that Chinese-made chips remain suitable mainly for inference rather than large-scale AI training, Maybank said in a note.

Baidu has also demonstrated the reliability of domestic hardware, with a version of its Ernie-5.1 model trained on a fully domestic Kunlunxin cluster achieving a 97 percent effective training rate.

Separately, DeepSeek-V4-Pro completed full-parameter post-training on a cluster of at least 1,000 Huawei Ascend 910C chips, although this involved post-training rather than pre-training.

The developments come as access to Nvidia’s AI chips becomes increasingly constrained. H100 one-year rental rates rose to $2.35 per GPU-hour in March 2026 from $1.70 in October 2025, a 40 percent increase, while China-specific H100 rates increased to RMB 80,000 – RMB 90,000 ($11,868-$13,352), including rack costs.

The research report said most GPU suppliers were reporting H-series shortages, with some H100 renewal contracts extending to four years. Nvidia’s Blackwell capacity before August-September 2026 was also fully booked, against a reported backlog of about US$1 trillion through 2027.

At the same time, Beijing has banned foreign AI chips in state-funded data centers and ordered projects that are less than 30% complete to remove installed foreign-made silicon.

Domestic vendors are forecast to account for about 56 percent of China’s AI server chip market in 2026, up from 46 percent, while the foreign share is expected to fall to about 21 percent from 34 percent.

The report said the near-term beneficiaries could be companies operating at the software-abstraction layer rather than chipmakers themselves.

HeteroFlow v2, launched on Aug. 8, unifies nine domestic GPU brands through a single application programming interface, or API. The software uses an open-source, multi-level intermediate representation-based compiler framework, allowing domestic chips to leverage existing community-developed tools rather than rebuilding Nvidia’s software ecosystem from scratch.

However, significant challenges remain. The research report cited Epoch’s assessment that China could still be at least a decade behind in hardware, with frontier AI training feasible but materially more expensive.

A key obstacle is the “migration tax” involved in moving AI workloads from Nvidia’s CUDA platform to domestic alternatives. Existing training pipelines require substantial rewriting of custom code, adding at least 50 percent to the time and cost for each team, the report said.

As a result, the transition is increasingly split by workload. Pre-training remains largely dependent on Nvidia hardware because of interconnect limitations and weaker compute utilization at scale, while large-scale inference, standard workloads and edge and automotive applications are gradually shifting towards domestic chips.

The report said automated AI coding and porting tools could help narrow the software gap, with portability rates reaching as high as 90 percent in some cases.

China’s domestic AI ecosystem is therefore increasingly using AI-generated code to accelerate the migration to local hardware, potentially creating a self-reinforcing cycle of software development and chip adoption, the report said.

Alipay launches agentic commerce platform in China to bring AI tools to merchants