EvidenceChain answer

How do the benchmark scores of Ornith-1.0-9B, such as its 43.1 on Terminal-Bench 2.1 and 69.4 on SWE-Bench Verified, dem

3

Ornith-1.0-9B reports a score of 43.1 on Terminal-Bench 2.1 and 69.4 on SWE-Bench Verified [2]. These figures put the 9‑billion‑parameter model in the same league as much larger systems. According to the available evidence, the model has outperformed Google’s Gemma 4 31B on real coding benchmarks and also challenges Qwen 3.6 35B [4]. One source specifically states that Ornith‑1.0‑9B outperforms Qwen 3.6 35B in various benchmarks [1]. The strong benchmark results chip away at the assumption that only giant, frontier‑scale models can handle agentic coding work [3]. In short, the scores themselves are high enough that the small model matches or exceeds the performance of models several times its size—directly showing that superior coding ability does not require tens of billions of parameters.

Discussion

Comments

0

No comments yet. Be the first to add a useful angle.