Eval-Actions

How to read this picture

Real-robot episodes with expert quality grades, used to train and test models that judge how well manipulation was executed.1

Sources
Last checked 10 Oct 2026Basic entry15 of 23 facts checked at the sourceNext check 8 Apr 2027
Runs in
Recorded data1
Checked against real robots
Not checked
Skill
Handling objects
Robot
One arm, Two arms1
Used by
52
citations
Licence
Apache-2.03

Comparisons with real robots

Real robots1

Details

Data are real-robot recordings scored offline, so sim-to-real does not apply. Validity evidence is judge-vs-expert agreement: AutoEval-S SRCC 0.81 (EG) and 0.84 (RG), success accuracy 90.6% / 91.0%; AutoEval-P SRCC 0.70 (CoT); inter-expert leave-one-out SRCC 0.91 +/- 0.06, ICC(2,1) 0.88. Policy-level use: 5 policies x 5 tasks x 20 rollouts; AutoEval-EG ranking differs from success-rate ranking (RDT vs ACT). All measured by the authors.

Known problems 1

  1. Limited scope of quality score

    Authors: scores mainly reflect spatial generalisation, not language/object/long-horizon generalisation; collision/safety assessment incomplete; judge may fail under occlusion.1

Details

About

What it is
Dataset Inferred4
More

Classified by the Atlas from how the authors describe and distribute it.

Built by
Unknown
More

No affiliations in arXiv v1 or v2 HTML; v2 and project page say 'Anonymous'. Repo owner is GitHub user LogSSim.

Released
2026-01 (arXiv v1 2026-01-26)4
More

v1 title: 'Trustworthy Evaluation of Robotic Manipulation: A New Benchmark and AutoEval Methods' (https://arxiv.org/abs/2601.18723v1).

Version
arXiv v2; dataset 'coming soon'; code repo initial commits only5
More

Project page: 'Code (Coming Soon)', 'Dataset (Coming Soon)'. Repo README: 'Dataset coming soon', download links 'coming soon'.

Last update
arXiv v2 2026-06-28 (new title; reframed as diagnostic methodology; abstract adds 13K+ episodes / 150+ tasks / 52 h)4
More

Repo last commit 2026-01-27 (README update).

Status
Active Inferred4
More

Paper revised 2026-06-28; data release still pending.

Setup

Runs in
Recorded data1
More

Judges score recorded episodes; no control loop at evaluation time.

Robot
One arm, Two arms1
Robot model
ARX R5 and UR5 manipulators (single-arm and bimanual)1
More

Joint trajectories 7/14-DoF.

Setting
Tabletop Inferred1
More

Task examples are tabletop; scene types not enumerated.

Size
13K+ episodes; 150+ tasks; about 52 h; 2.8K failures. EAS subset: 6K+ episodes, 50+ tasks, 12 h, 37.4% failure ratio, split 80/10/101
More

Table III; project page repeats it.

Scoring and access

Scored by
Human rating, Automatic judge1
More

Expert Grading: 10 experts score 1-10 on success, collisions, smoothness, efficiency. Rank-Guided labels: physical indicators weighted to match expert rankings (fit on train split). CoT text explanations. Evaluators are scored by SRCC to these labels and success accuracy.

Leaderboard
None5
More

No leaderboard on page, paper or repo.

Code licence
Apache-2.03
More

LICENSE file is Apache 2.0, but the README shows an MIT badge and a commented-out 'released under the MIT License' line: conflict, LICENSE file taken as authoritative.

Data licence
Unknown
More

Dataset not released; checked project page, repo README and paper.

Published at
Unknown
More

arXiv comments give only site and code links. Project page links a file named 'T_RO.pdf'; we did not treat that as evidence of a venue.

Sources 5

  1. 1Eval-Actions: Fine-Grained Execution Quality Evaluation for Robotic Manipulation (full text)Paper · Jan 2026 · checked 10 Oct 2026
  2. 2Semantic Scholar API recordIndex · checked 10 Oct 2026
  3. 3LogSSim/TERM-Bench on GitHub (blob)Repository · checked 10 Oct 2026
  4. 4Eval-Actions: Fine-Grained Execution Quality Evaluation for Robotic ManipulationPaper · Jan 2026 · checked 10 Oct 2026
  5. 5Eval-Actions: Fine-Grained Execution Quality Evaluation for Robotic ManipulationOfficial site · checked 10 Oct 2026

Change history

  1. Created as a basic entry: identity facts checked at primary sources (phase 1 re-verification).