Skip to content

01 Data Factor Layer

1. Industry Chain Positioning: AI's "Crude Oil" and "Engine"

The Data Factor Layer is the "crude oil extraction and refining" stage of the AI industry chain — it determines what AI models "eat" and whether they eat well. In our framework, it is the first tier of the "physical foundation" and the bedrock of the entire AI edifice.

In the logic of industry-chain transmission: Data feeding → Power driving → Chip computing → Model thinking → Application landing, data occupies the uppermost position. If data supply is insufficient or of poor quality, the entire chain will "run out of fuel."

Data plays an extremely critical dual role in the AI industry chain:

  • AI's "food/oil": Training and inference of large models rely on massive amounts of high-quality data. Without data, compute and models are out of the question.
  • AI's "value engine": AI is not only a consumer of data, but also an amplifier of data value. High-quality industry data combined with AI forms a positive loop of "data → AI → business → data," transforming data from a cost center into a profit center.

2. Key Turning Point: From Training to Inference

2025 has become a watershed for AI moving from "training" to "large-scale application." For the first time, the volume of AI inference data in China exceeded training data.
Meanwhile, the National Data Administration has officially named the Token in the AI field "Ciyuan", treating it as the smallest standardized unit of value after semantic decomposition of data by large models — marking a new unit of measure for AI development.

3. Core Trend: High-Quality Data Replaces Model Selection

With the surge of open-source / low-cost models such as DeepSeek triggering the wave of large-model democratization, the functionality of foundation models is becoming increasingly homogeneous. Against this backdrop, high-quality data is replacing model selection as the absolute decisive factor in AI core competitiveness.


4. Benchmark Companies in This Layer

Databricks

Unlisted | United States
Data Lakehouse and AI Governance

Core Business: Provides a unified data analytics and AI platform, pioneering the Lakehouse architecture to unify data storage and AI governance.

⚔️ Moat: Unifies fragmented data silos (data lake and data warehouse), providing enterprises with full-stack AI services from data preparation to model training to deployment.

Key Competitors:

Scale AI

Unlisted | United States
Data Annotation & Synthetic Data

Core Business: Provides high-quality human annotation (RLHF) and synthetic data generation services for AI large model training, and is a benchmark enterprise in AI data infrastructure.

⚔️ Moat: Possesses a vast and professional global network of annotation experts, with its data platform deeply integrated with leading large-model companies such as OpenAI and Meta, as well as the U.S. Department of Defense, making it the highest-valued unicorn in the data element layer.

Key Competitors:

IntSig Information

688615:SH | China
Commercial Big Data and Document Intelligence

Core Business: Provides AI-based intelligent character recognition (OCR) and commercial big data mining services, with well-known apps such as Qixinbao and CamScanner.

⚔️ Moat: Has deep expertise in document digitization and data mining for complex scenarios, pioneered the exploration of including data assets in balance sheets, and its large-model accelerator products provide intelligent text processing infrastructure for AI.

Key Competitors:

Scale AI

Unlisted | United States
Synthetic Data

Core Business: Provide high-quality data annotation and synthetic data generation services for AI models

⚔️ Moat: Through a human-in-the-loop annotation platform and synthetic data technology, it has built a data flywheel for generating massive high-quality data, deeply binding large model customers and defense departments

Key Competitors: Appen, Labelbox, Sama

NavInfo

002405:SZ | China
High-Definition Maps and Autonomous Driving Data

Core Business: Provides navigation maps, high-definition maps, autonomous driving data services, and connected vehicle solutions; it is the largest vehicle-location data service provider in China.

⚔️ Moat: Holds the country's most scarce national-scale high-definition location data assets, and is an indispensable data provider in the high-definition map and autonomous driving data sector in China.

Key Competitors:

TRS

300229:SZ | China
Semantic Intelligence and Government Data

Core Business: Deeply engaged in natural language processing (NLP) and big data technologies, providing semantic intelligence and knowledge mining services for vertical industries such as government and media.

⚔️ Moat: Holds an extremely high market share in domestic government and central media data processing, skilled at extracting knowledge from unstructured text and building 'data brains' for specialized fields.

Key Competitors:

Haitian Ruisheng

688787:SH | China
AI Training Data Services

Core Business: Provides high-quality datasets and customized annotation services for AI model training in speech, vision, natural language, and other modalities.

⚔️ Moat: A rare domestic enterprise with high-quality multilingual and multimodal data accumulation, deeply involved in the construction standards of domestic large-model corpora, and the first to launch an embodied intelligence data engineering service platform.

Key Competitors: