DeepSeek-V4 Flash 0731 vs GPT-5.6 Luna on DeepSWE: Cost and Coding
Together AI 3 weeks ago 36 ● 6 sources
DeepSeek-V4 Flash 0731 achieved 53.3% pass@1 on the DeepSWE coding benchmark versus GPT-5.6 Luna's 67.2%, but costs $0.10 per task compared to Luna's $0.61—a 6x price difference. When used in cascade (DeepSeek first, escalating to Luna on failure), the pairing solves 78.9% of tasks at $0.385 each, outperforming Luna alone on accuracy while costing 37% less. This strategy leverages DeepSeek's low cost to handle routine problems and reserves Luna's stronger reasoning for harder cases, fundamentally changing the economics of software engineering task automation.