DeepSeek releases V3 large model: 671B MoE base, performance benchmarking GPT-4o and Claude 3.5
On December 26, 2024, DeepSeek released the 671B parameter MoE base model DeepSeek-V3, with a single token activation of about 37B. It added multi-token prediction and adopted FP8 mixed precision training. Its performance benchmarked GPT-4o and Claude 3.5 Sonnet, and attracted industry attention with extremely low training cost.
On December 26, 2024, DeepSeek released
Picture: DeepSeek official website dialogue interface. V3 is pre-trained on 14.8T tokens multi-lingual corpus, using FP8 mixed precision and trained on the H800 cluster. The official disclosure of the total training cost is approximately US$5.576 million (approximately 2.788 million GPU hours).
Version overview
| Project | Content |
|---|
Reviews