DeepSeek V4 vs Claude vs GPT-5.4: A 38-Task Benchmark Across Coding, Reasoning, and Financial Research

FUNDA·April 24, 2026

Opus 4.6/4.7 tie #1 at 8.72; DeepSeek V4 Pro close at 8.27 with lower cost; GPT-5.4 fastest.

Important note: This report is not a research report. It is an evaluation report completed by the FundaAI Engineering Team, not written by the FundaAI Analyst Team. It does not represent the views of the FundaAI Analyst Team.

All test cases are based on the actual working environment of the FundaAI Platform.

As of time of publication, GPT-5.5 has not yet officially released its API. Testing solely through Codex 5.5 may not fully reflect the complete performance of the API. We have currently only conducted urgent testing on DeepSeek V4, and will include GPT-5.5 test results as soon as its API becomes officially available.


Key Takeaways

  • Claude Opus 4.6 (Thinking) and Claude Opus 4.7 tie for #1 overall (both 8.72 weighted avg). They lead for different reasons: Opus 4.6 Thinking is strongest on coding and hard reasoning, while Opus 4.7 leads writing and full-coverage multi-step work.

Continue reading with FUNDA

This report is available to subscribers. Sign in or subscribe to read the full analysis.