SciForma: Structure-Faithful Generation of Scientific Diagrams
SciFormaBench-2K Main Results:
| Method | Average Score | Component | Arrow | Text |
|---|---|---|---|---|
| GPT-Image-2 (Proprietary) | 85.62 | 83.34 | 89.61 | 83.53 |
| Nano Banana Pro (Proprietary) | 81.34 | 81.10 | 83.60 | 78.70 |
| SciForma-9B + Edit | 72.40 | 76.70 | 69.91 | 70.14 |
| SciForma-9B (M-DPO) | 69.51 | 74.49 | 66.46 | 67.00 |
| GPT-Image-1.5 (Proprietary) | 68.96 | 75.70 | 62.50 | 68.20 |
| SciForma-Base (SFT) | 67.59 | 73.52 | 64.64 | 63.84 |
| FLUX.2-klein-base-9B (Base Model) | 33.87 | 51.50 | 25.20 | 23.60 |
- M-DPO improves the SFT model's average score from 67.59 to 69.51 (+1.92), with the largest gains on the weakest axes: Arrow (+1.82) and Text (+3.16). Iterative refinement adds another +2.89 points.
AIBench Main Results:
| Method | Overall Score | Component | Topology | Phase | Semantics |
|---|---|---|---|---|---|
| GPT-Image-2 | 80.27 | 90.69 | 81.77 | 84.18 | 90.24 |
| SciForma-9B | 70.29 | 77.53 | 64.17 | 74.71 | 80.79 |
| Original Image (Human) | 70.09 | 82.65 | 57.98 | 79.57 | 79.13 |
| GPT-Image-1.5 | 61.62 | 66.23 | 50.87 | 55.95 | 77.55 |
- On AIBench, which uses VQA to test logical understanding, SciForma-9B (70.29) slightly outperforms human-drawn originals (70.09) and surpasses GPT-Image-1.5 by 8.67 points. The largest margin over originals is in Topology (+6.19).
Ablation Studies:
- M-DPO achieves a +1.92 score improvement in 4K training steps, while continuing SFT for 30K steps yields only a +0.37 gain.
- Scalar reward-based alternatives fail to improve over the SFT baseline: GDRO achieves a score of 67.49 (-0.10) and GRPO scores 66.30 (-1.29).
| Objective | Average Score | Δ vs SFT | Component | Arrow | Text |
|---|---|---|---|---|---|
| SciForma (SFT Baseline) | 67.59 | - | 73.52 | 64.64 | 63.84 |
| + DPO (Scalar) | 68.42 | +0.83 | 74.14 | 65.32 | 65.08 |
| + M-DPO (mean) | 69.31 | +1.72 | 74.54 | 66.28 | 66.47 |
| + M-DPO (Conjunctive) | 69.51 | +1.92 | 74.49 | 66.46 | 67.00 |
- Peer multi-dimensional DPO methods like CaPO and MCDPO cause the Text axis to regress (-1.34 and -1.85 respectively), while M-DPO's conjunctive objective improves all axes.
Additional Evaluations:
- On the PaperBanana agentic benchmark, SciForma achieves a 30.7% pairwise win rate against human originals, surpassing GPT-Image-1.5 (19.0%).
- A user study with 36 graduate students found that SciForma-9B outperforms Wan2.7-Image and is competitive with Nano Banana Pro. User preferences showed a Pearson correlation of r=0.76 with SciFormaBench-2K scores.
The authors conclude that SciForma enables the structure-faithful generation of scientific diagrams. By decomposing diagram quality into verifiable axes and enforcing their joint correctness with the M-DPO objective, the SciForma-9B model achieves structural fidelity and aesthetic quality approaching that of proprietary systems. The work demonstrates the effectiveness of incorporating explicit, multi-dimensional structural constraints into the training process for complex, structured image generation tasks.
Don't read this site daily. Get it in your inbox.
The daily brief and Sunday deep dive — distilled, scored, and opinionated. For builders only.