Eval-Actions
Real-robot episodes with expert quality grades, used to train and test models that judge how well manipulation was executed.1
Comparisons with real robots
Real robots1
Details
Data are real-robot recordings scored offline, so sim-to-real does not apply. Validity evidence is judge-vs-expert agreement: AutoEval-S SRCC 0.81 (EG) and 0.84 (RG), success accuracy 90.6% / 91.0%; AutoEval-P SRCC 0.70 (CoT); inter-expert leave-one-out SRCC 0.91 +/- 0.06, ICC(2,1) 0.88. Policy-level use: 5 policies x 5 tasks x 20 rollouts; AutoEval-EG ranking differs from success-rate ranking (RDT vs ACT). All measured by the authors.
Known problems 1
Limited scope of quality score
Authors: scores mainly reflect spatial generalisation, not language/object/long-horizon generalisation; collision/safety assessment incomplete; judge may fail under occlusion.1
Details
About
- What it is
- Dataset Inferred4
More
Classified by the Atlas from how the authors describe and distribute it.
- Built by
- Unknown
More
No affiliations in arXiv v1 or v2 HTML; v2 and project page say 'Anonymous'. Repo owner is GitHub user LogSSim.
- Released
- 2026-01 (arXiv v1 2026-01-26)4
More
v1 title: 'Trustworthy Evaluation of Robotic Manipulation: A New Benchmark and AutoEval Methods' (https://arxiv.org/abs/2601.18723v1).
- Version
- arXiv v2; dataset 'coming soon'; code repo initial commits only5
More
Project page: 'Code (Coming Soon)', 'Dataset (Coming Soon)'. Repo README: 'Dataset coming soon', download links 'coming soon'.
- Last update
- arXiv v2 2026-06-28 (new title; reframed as diagnostic methodology; abstract adds 13K+ episodes / 150+ tasks / 52 h)4
More
Repo last commit 2026-01-27 (README update).
- Status
- Active Inferred4
More
Paper revised 2026-06-28; data release still pending.
Setup
- Runs in
- Recorded data1
More
Judges score recorded episodes; no control loop at evaluation time.
- Robot
- One arm, Two arms1
- Robot model
- ARX R5 and UR5 manipulators (single-arm and bimanual)1
More
Joint trajectories 7/14-DoF.
- Setting
- Tabletop Inferred1
More
Task examples are tabletop; scene types not enumerated.
- Size
- 13K+ episodes; 150+ tasks; about 52 h; 2.8K failures. EAS subset: 6K+ episodes, 50+ tasks, 12 h, 37.4% failure ratio, split 80/10/101
More
Table III; project page repeats it.
Scoring and access
- Scored by
- Human rating, Automatic judge1
More
Expert Grading: 10 experts score 1-10 on success, collisions, smoothness, efficiency. Rank-Guided labels: physical indicators weighted to match expert rankings (fit on train split). CoT text explanations. Evaluators are scored by SRCC to these labels and success accuracy.
- Leaderboard
- None5
More
No leaderboard on page, paper or repo.
- Code licence
- Apache-2.03
More
LICENSE file is Apache 2.0, but the README shows an MIT badge and a commented-out 'released under the MIT License' line: conflict, LICENSE file taken as authoritative.
- Data licence
- Unknown
More
Dataset not released; checked project page, repo README and paper.
- Published at
- Unknown
More
arXiv comments give only site and code links. Project page links a file named 'T_RO.pdf'; we did not treat that as evidence of a venue.
Sources 5
- 1Eval-Actions: Fine-Grained Execution Quality Evaluation for Robotic Manipulation (full text)Paper · Jan 2026 · checked 10 Oct 2026
- 2Semantic Scholar API recordIndex · checked 10 Oct 2026
- 3LogSSim/TERM-Bench on GitHub (blob)Repository · checked 10 Oct 2026
- 4Eval-Actions: Fine-Grained Execution Quality Evaluation for Robotic ManipulationPaper · Jan 2026 · checked 10 Oct 2026
- 5Eval-Actions: Fine-Grained Execution Quality Evaluation for Robotic ManipulationOfficial site · checked 10 Oct 2026
Change history
- Created as a basic entry: identity facts checked at primary sources (phase 1 re-verification).