DeepSeek V4 深度评测:38项任务横评 Claude、GPT-5.4

FUNDA·2026年4月24日

Opus 4.6/4.7 tie #1 at 8.72; DeepSeek V4 Pro close at 8.27 with lower cost; GPT-5.4 fastest.

重要说明:本报告不是研究报告,而是由 FundaAI 工程团队完成的模型评估报告,并非由 FundaAI 分析师团队撰写,也不代表 FundaAI 分析师团队观点。

所有测试用例均基于 FundaAI Platform 的真实工作环境。

截稿时,GPT-5.5 尚未正式开放 API。仅通过 Codex 5.5 进行测试,可能无法完整反映其 API 版本的真实表现。我们目前只对 DeepSeek V4 进行了紧急测试;GPT-5.5 的 API 正式开放后,我们会尽快补充其测试结果。


核心结论

  • Claude Opus 4.6 (Thinking) 与 Claude Opus 4.7 并列综合第一​:二者加权平均分均为 8.72。Opus 4.6 Thinking 更强在编程和 hard reasoning,Opus 4.7 更强在写作和完整覆盖的多步任务。

Continue reading with FUNDA

This report is available to subscribers. Sign in or subscribe to read the full analysis.