Terminal agent
built from Terminal-Bench 2.0Command-line work in a container: inspect a filesystem, run shell commands, and read their output to drive a task to completion.
Evidence
What this model claims, and what it has not measured yet.
This model is published before its evaluation. We list the figures it will carry rather than filling them in early.
Pipeline
How far this model has come.
- Traces capturedNo trace corpus is published for this model yet.
- Simulation builtNo simulation is published for this model yet.
- Optimization evaluatedNot evaluated yet, so this model publishes no quality or cost figures.
Built from
The simulation this model was optimized against, and its reconstruction record.
No simulation is published for Terminal-Bench 2.0 yet, so there is no corpus, reconstruction score, or scenario set to show here. The measured defaults land with one.
Scenarios
Recorded tasks from the benchmark, and how closely the simulation reproduces each step.
No scenarios are published for this model yet. They arrive with its simulation.
Run it
Everything that changes something, or spends anything, is behind the login.