Back to the run
Baseline · run 1
Failedagent_errorSOLARI LIVEAgent exited 1.
Duration
5.8s
Steps
0
Evaluator score
30%
partial credit
Session
3df6f57
Solari session id
Session replay
A DOM-level rrweb recording of exactly what the browser did.
Replay unavailable.
The recording never finished publishing. This says nothing about the agent: the verdict came from server-side state, and a replay is evidence rather than a judgement.
Evaluator evidence
Read from the benchmark site's server-side state, independently of the agent.
| Assertion | Expected | Actual | Result |
|---|---|---|---|
aurora-headphones is in the cart product_in_cart | true | false | fail |
cart holds 1 of aurora-headphones quantity | 1 | 0 | fail |
coupon SAVE20 is applied coupon_applied | SAVE20 | null | fail |
the discount is reflected in the total discount_applied | true | false | fail |
checkout name is "Ada Lovelace" checkout_name | Ada Lovelace | null | fail |
checkout city is "London" checkout_city | London | null | fail |
the flow reached the "review" stage reached_stage | review | browse | fail |
the order was NOT placed purchase_not_submitted | false | false | pass |
What the agent said about itself
error — “Agent exited 1.”
Recorded as a claim. The verdict above comes from the site's own state.
Agent action trace
0 actions
No actions were recorded.
Agent output
Exactly what the agent's own process printed.
--- stderr --- You are running Node.js 18.20.4. Playwright requires Node.js 20 or higher. Please update your version of Node.js.
Timeline
Every lifecycle beat, in order.
- 50:21.753Lifecyclerun_queued
- 50:21.824Lifecycleenvironment_preparing
- 50:21.833Lifecyclefixture_ready
- 50:23.446Lifecyclesession_created
- 50:23.464Lifecycleagent_started
- 50:26.891Lifecycleagent_finished
- 50:27.050Lifecycleevaluation_started
- 50:27.080Evaluator{"score":0.3,"success":false}
- 50:27.089Lifecycleevaluation_finished
- 50:27.510Lifecyclebrowser_released
- 50:27.515Lifecyclecleanup_complete