Cura 1T: Specialized Model for Agentic Healthcare
Healthcare Benchmark Performance vs. Base Model:
| Benchmark | Base (Kimi-K2.6) | Cura 1T | Δ |
|---|---|---|---|
| MedAgentBench | 0.847 | 0.940 | +0.093 |
| HealthBench Professional | 0.503 | 0.662 | +0.159 |
| HealthBench Hard | 0.222 | 0.368 | +0.146 |
| MedXpertQA | 0.569 | 0.655 | +0.086 |
| AgentClinic | 0.754 | 0.796 | +0.042 |
MedAgentBench (Task Success):
- Cura 1T achieves a task success rate of 0.940, outperforming Claude Opus 4.8 (0.937), Gemini 3.1 Pro (0.913), and GPT-5.5 (0.894).
- The self-evolution loop improved the score from a base of 0.883 to 0.973 in a benchmark-specific round by targeting brittle execution of EHR writes with synthetic tool-use trajectories.
HealthBench (Rubric Score):
- On HealthBench Professional, Cura 1T scores 0.662, the highest among evaluated models.
- On HealthBench Hard, Cura 1T scores 0.368.
- An early training round that improved aggregate scores was reverted because it caused significant degradation on a subset of tasks (e.g., a 0.508 drop on Hard for that subset), highlighting the need for targeted, clean data mixtures.
MedXpertQA (pass@1):
| Model | Text | Multimodal | Overall |
|---|---|---|---|
| Cura 1T | 0.600 | 0.722 | 0.655 |
| GPT-5.5 | 0.596 | 0.771 | 0.675 |
| Claude Opus 4.8 | 0.562 | 0.710 | 0.628 |
| Kimi-K2.6 (Base) | 0.484 | 0.672 | 0.569 |
- Cura 1T is the second-strongest model overall. The improvement was driven by adding closed-book clinical knowledge and retention examples, as reasoning-only corrections were found to be unstable and were reverted.
AgentClinic (pass@1):
- Using a tool-native harness, Cura 1T achieves an overall pass@1 of 0.796, outperforming Claude Opus 4.8 (0.794) and GPT-5.5 (0.684).
- The key improvement came from training on interactive trajectories with retention anchors, which taught the model to complete a clinical workup before making a premature diagnosis.
Out-of-Domain Performance:
- On reasoning benchmarks (AIME 2025, AIME 2026, GPQA-Diamond), Cura 1T performs on par with frontier models.
- On agentic benchmarks (τ²-Bench), Cura 1T is on par with comparators on the airline domain and surpasses publicly reported scores on the retail and telecom domains, indicating that healthcare specialization did not erode general capabilities.
Cura 1T demonstrates that a specialized healthcare LLM can be effectively trained via a self-evolution loop that focuses on data mixture curation. By analyzing benchmark failures and synthesizing targeted data, this method improves performance across patient care, clinical reasoning, and agentic healthcare tasks. The model achieves state-of-the-art or near state-of-the-art results on a suite of healthcare benchmarks while preserving general reasoning and agentic capabilities on out-of-domain evaluations.
Don't read this site daily. Get it in your inbox.
The daily brief and Sunday deep dive — distilled, scored, and opinionated. For builders only.