EvidenceChain answer
How do the benchmark scores of Ornith-1.0-9B, such as its 43.1 on Terminal-Bench 2.1 and 69.4 on SWE-Bench Verified, dem
Ornith-1.0-9B reports a score of 43.1 on Terminal-Bench 2.1 and 69.4 on SWE-Bench Verified [2]. These figures put the 9‑billion‑parameter model in the same league as much larger systems. According to the available evidence, the model has outperformed Google’s Gemma 4 31B on real coding benchmarks and also challenges Qwen 3.6 35B [4]. One source specifically states that Ornith‑1.0‑9B outperforms Qwen 3.6 35B in various benchmarks [1]. The strong benchmark results chip away at the assumption that only giant, frontier‑scale models can handle agentic coding work [3]. In short, the scores themselves are high enough that the small model matches or exceeds the performance of models several times its size—directly showing that superior coding ability does not require tens of billions of parameters.
Discussion
Comments
Sign in to join the discussion
Comments are open to registered users so replies and notifications stay tied to a real account.
No comments yet. Be the first to add a useful angle.