01 Data Factor Layer
1. Industry Chain Positioning: AI's "Crude Oil" and "Engine"
The Data Factor Layer is the "crude oil extraction and refining" stage of the AI industry chain — it determines what AI models "eat" and whether they eat well. In our framework, it is the first tier of the "physical foundation" and the bedrock of the entire AI edifice.
In the logic of industry-chain transmission: Data feeding → Power driving → Chip computing → Model thinking → Application landing, data occupies the uppermost position. If data supply is insufficient or of poor quality, the entire chain will "run out of fuel."
Data plays an extremely critical dual role in the AI industry chain:
- AI's "food/oil": Training and inference of large models rely on massive amounts of high-quality data. Without data, compute and models are out of the question.
- AI's "value engine": AI is not only a consumer of data, but also an amplifier of data value. High-quality industry data combined with AI forms a positive loop of "data → AI → business → data," transforming data from a cost center into a profit center.
2. Key Turning Point: From Training to Inference
2025 has become a watershed for AI moving from "training" to "large-scale application." For the first time, the volume of AI inference data in China exceeded training data.
Meanwhile, the National Data Administration has officially named the Token in the AI field "Ciyuan", treating it as the smallest standardized unit of value after semantic decomposition of data by large models — marking a new unit of measure for AI development.
3. Core Trend: High-Quality Data Replaces Model Selection
With the surge of open-source / low-cost models such as DeepSeek triggering the wave of large-model democratization, the functionality of foundation models is becoming increasingly homogeneous. Against this backdrop, high-quality data is replacing model selection as the absolute decisive factor in AI core competitiveness.
4. Benchmark Companies in This Layer
Databricks
Unlisted | United StatesCore Business: Provides a unified data analytics and AI platform, pioneering the Lakehouse architecture to unify data storage and AI governance.
⚔️ Moat: Unifies fragmented data silos (data lake and data warehouse), providing enterprises with full-stack AI services from data preparation to model training to deployment.
Key Competitors:
Scale AI
Unlisted | United StatesCore Business: Provides high-quality human annotation (RLHF) and synthetic data generation services for AI large model training, and is a benchmark enterprise in AI data infrastructure.
⚔️ Moat: Possesses a vast and professional global network of annotation experts, with its data platform deeply integrated with leading large-model companies such as OpenAI and Meta, as well as the U.S. Department of Defense, making it the highest-valued unicorn in the data element layer.
Key Competitors:
IntSig Information
688615:SH | ChinaCore Business: Provides AI-based intelligent character recognition (OCR) and commercial big data mining services, with well-known apps such as Qixinbao and CamScanner.
⚔️ Moat: Has deep expertise in document digitization and data mining for complex scenarios, pioneered the exploration of including data assets in balance sheets, and its large-model accelerator products provide intelligent text processing infrastructure for AI.
Key Competitors:
Scale AI
Unlisted | United StatesCore Business: Provide high-quality data annotation and synthetic data generation services for AI models
⚔️ Moat: Through a human-in-the-loop annotation platform and synthetic data technology, it has built a data flywheel for generating massive high-quality data, deeply binding large model customers and defense departments
Key Competitors: Appen, Labelbox, Sama
NavInfo
002405:SZ | ChinaCore Business: Provides navigation maps, high-definition maps, autonomous driving data services, and connected vehicle solutions; it is the largest vehicle-location data service provider in China.
⚔️ Moat: Holds the country's most scarce national-scale high-definition location data assets, and is an indispensable data provider in the high-definition map and autonomous driving data sector in China.
Key Competitors:
TRS
300229:SZ | ChinaCore Business: Deeply engaged in natural language processing (NLP) and big data technologies, providing semantic intelligence and knowledge mining services for vertical industries such as government and media.
⚔️ Moat: Holds an extremely high market share in domestic government and central media data processing, skilled at extracting knowledge from unstructured text and building 'data brains' for specialized fields.
Key Competitors:
Haitian Ruisheng
688787:SH | ChinaCore Business: Provides high-quality datasets and customized annotation services for AI model training in speech, vision, natural language, and other modalities.
⚔️ Moat: A rare domestic enterprise with high-quality multilingual and multimodal data accumulation, deeply involved in the construction standards of domestic large-model corpora, and the first to launch an embodied intelligence data engineering service platform.
Key Competitors: