DeepSeek releases V3 large model: 671B MoE base, performance benchmarking GPT-4o and Claude 3.5

On December 26, 2024, DeepSeek released the 671B parameter MoE base model DeepSeek-V3, with a single token activation of about 37B. It added multi-token prediction and adopted FP8 mixed precision training. Its performance benchmarked GPT-4o and Claude 3.5 Sonnet, and attracted industry attention with extremely low training cost.

On December 26, 2024, DeepSeek released , a 671B parameter MoE base model, which laid the foundation for the capabilities of the subsequent series.

DeepSeek-V3 671B MoE open source large model release overview

Picture: DeepSeek official website dialogue interface. V3 is pre-trained on 14.8T tokens multi-lingual corpus, using FP8 mixed precision and trained on the H800 cluster. The official disclosure of the total training cost is approximately US$5.576 million (approximately 2.788 million GPU hours).

Version overview

Project Content
Copyright: Content sourced from DeepSeek official . This platform has compiled and organized this content for informational purposes and learning exchange only. If there are any copyright concerns, please contact us for resolution.

Reviews

  • Loading reviews...