world-model-optimizer

Customer service agent

built from τ²-bench

Multi-turn customer-service tool use: look up a user, book or modify an order, and hold to the domain's policy across the conversation.

Evidence

What this model claims, and what it has not measured yet.

task successnot yet measured
cost per runnot yet measured
savings vs baselinenot yet measured
benchmarkτ²-bench
tool-callscustomer-servicemulti-turn

This model is published before its evaluation. We list the figures it will carry rather than filling them in early.

Pipeline

How far this model has come.

  • Traces capturedReal agent sessions from the benchmark, recorded step by step.
  • Simulation builtA simulator of the environment, scored against held-out scenarios.
  • Optimization evaluatedNot evaluated yet, so this model publishes no quality or cost figures.

Built from

The simulation this model was optimized against, and its reconstruction record.

τ²-bench customer service - Customer-service tool environment (airline, retail, telecom) reconstructed from real τ²-bench agent sessions. Step it with the benchmark's tool calls (look up users, book flights, modify orders), and it responds as the live backend would.

Sign in to open this simulationStep it, inspect its full trace corpus, or import it into your workspace.
reconstruction fidelity91.5%How closely the model's observations match the real environment's, on a held-out test slice (±0.003).
corpus12 traces / 67 steps
serve LLMOpus 4.8
providerbedrock
tasktau-bench
tool-callscustomer-service

fidelity from trace-scaling (open-loop RAG)

Scenarios

Recorded tasks from the benchmark, and how closely the simulation reproduces each step.

Recorded traces, grouped by task. Open one to see how faithfully the simulation reconstructs each step against what really happened.

Run it

Everything that changes something, or spends anything, is behind the login.

Models
by Experiential Labs. Simulating reality for hypothesis testing.