Cross-trial summary comparison
| Metric | DeepSeek-V4-Flash-Vision-Exp (Winner) | Qwen3.8-Flash-Next-NVFP4 | Δ |
|---|---|---|---|
| Mean Score | 88.6 | 85.4 | +3.2 |
| Std Dev | ±1.6 | ±2.0 | -0.4 |
| Mean Points | 148.9 | 143.2 | +5.7 |
| Deployability (α=0.7) | 79 / 100 | 69 / 100 | +10 |
| Quality | 90 / 100 | 83 / 100 | +7 |
| Responsiveness | 52 / 100 | 36 / 100 | +16 |
| Safety Warnings (max) | 1 | 2 max | Winner |
| Reliability (Pass⁸) | 75.0% | 67.9% | +7.1 |
| Category | DeepSeek-V4-Flash-Vision-Exp | Qwen3.8-Flash-Next-NVFP4 | Diff |
|---|---|---|---|
| Tool Selection | 100% | 100% | — |
| Parameter Precision | 88% | 96% | -8 |
| Multi-Step Chains | 100% | 91% | +9 |
| Restraint & Refusal | 100% | 100% | — |
| Error Recovery | 100% | 100% | — |
| Localization | 100% | 100% | — |
| Structured Reasoning | 100% | 100% | — |
| Instruction Following | 90% | 80% | +10 |
| Context & State | 78% | 74% | +4 |
| Code Patterns | 100% | 96% | +4 |
| Safety & Boundaries | 83% | 78% | +5 |
| Toolset Scale | 97% | 97% | — |
| Autonomous Planning | 96% | 85% | +11 |
| Creative Composition | 83% | 85% | -2 |
| Structured Output | 94% | 87% | +7 |
| Hard Mode | 80% | 77% | +3 |