Reports

Coding

Latest reports tagged Coding.

Technology

How does Claude Sonnet 5's agentic performance on coding, tool use, and knowledge work benchmarks compare to its predecessor Sonnet 4.6 and the larger Opus 4.8 model?

What the evidence tells us Your question asks about Claude Sonnet 5’s agentic performance on coding, tool use, and knowledge work benchmarks, compared with Sonnet 4.6 and Opus 4.8. The available evidence gives us a partial picture — enough to see a couple of bright spots, but plenty of gaps...

1000
Technology

How do the benchmark scores of Ornith-1.0-9B, such as its 43.1 on Terminal-Bench 2.1 and 69.4 on SWE-Bench Verified, demonstrate that it matches or exceeds the performance of much larger models like Gemma 4-31B?

Ornith 1.0 9B reports a score of 43.1 on Terminal Bench 2.1 and 69.4 on SWE Bench Verified 2 . These figures put the 9‑billion‑parameter model in the same league as much larger systems. According to the available evidence, the model has outperformed Google’s Gemma 4 31B on real coding...

3000