Haitian Ruisheng
Provides high-quality datasets and customized annotation services for AI model training in speech, vision, natural language, and other modalities.
A rare domestic enterprise with high-quality multilingual and multimodal data accumulation, deeply involved in the construction standards of domestic large-model corpora, and the first to launch an embodied intelligence data engineering service platform.
Company Overview
Founded in 2005 and headquartered in Beijing, Haitian Ruisheng is one of the earliest companies in China to provide AI training data services. In 2021, it was listed on the STAR Market, becoming the first "AI data stock" in China.
The core business of Haitian Ruisheng is to provide "food" for AI models — high-quality annotated data. From early speech recognition corpora (dialects, multilingual) to today's LLM instruction fine-tuning (SFT) data and RLHF preference data, Haitian Ruisheng's business covers all categories of AI training data. Following the rise of embodied intelligence (robots), the company took the lead in releasing an embodied data platform for robot training.
Core Business
Intelligent Speech Data (Traditional Advantage)
Speech recognition corpora covering over 200 languages and dialects, serving as the underlying data foundation for speech products from companies such as iFlytek, Baidu, and Alibaba.
LLM Training Data
Provides LLM instruction fine-tuning (SFT) datasets, RLHF preference ranking datasets, and domain-specific data for vertical industries (legal, medical).
Embodied Intelligence Data
A newly launched data engineering service platform offering multimodal data collection and annotation services for robotics, including visual, tactile, and motion capture data.
Market Landscape
| Dimension | Data |
|---|---|
| Domestic AI Data Services | Leading enterprise |
| Main Competitors | Scale AI, IntSig Information |
| Core Customers | Baidu, Alibaba, iFlytek, domestic large model companies |
Risks and Summary
Haitian Ruisheng is the Chinese equivalent of Scale AI. Amid the domestic large model R&D boom, demand for high-quality Chinese corpus annotation and construction remains strong. Haitian Ruisheng's data accumulation in the speech and NLP domains constitutes its core moat.