Show HN: FrontierHarness Eval – 9 harness, same model, cost per pass varies 17x

72 points by shiqimei 12 hours ago on hackernews | 52 comments

Pass rate

Median cost per successful task

Median cost per task

Median cache hit rate per successful task

Median time per successful task

Beyond the numbers

  1. 01

    OpenCode: failures excluded.

    It only covers 15 passes. Count failed attempts and the number becomes $3.24 per task.

  2. 02

    Cache hit rate is not cost.

    A cached 300-turn failure can still burn more than a short cache miss.

  3. 03

    Quality and cost can diverge.

    Claude Code passes 19 tasks, but reaches $18.34 in cost per task.

Run your harness on Runta.

If you want to test your own harness on Runta, we’ll give you $100 in credits to get started.

Tested harnesses

Codex

v0.148.0

DeepSeek Harness

v0.1.0-rc.8

Claude Code

v2.1.237

Pi

v0.84.2

Oh My Pi

v17.4.0

Kimi Code

v0.37.2

Exo Harness

v0.1.0

OpenCode

v1.18.19

Hermes

v0.20.4
  • FrontierHarness v1.0 focuses on software engineering contexts and terminal-based tasks. It may not generalize to other areas of knowledge work.
  • Evaluated on Runta agent runtimes. All harnesses and the task environment are prepared once as a golden checkpoint. Every run is a fresh restore with identical vCPU, memory, disk size, disk contents, and memory state.