Benchmarks / Remote Labor Index

Reported by Remote Labor Index

Remote Labor Index

Share of 240 real paid freelance projects (design, software, data, architecture and more) an agent delivers at a standard a client would accept.

Results dated
1 Oct 2026
Models
14
Unit
% of projects
Licence
Published with permission (Scale AI and the Center for AI Safety results via Epoch AI)

Projects done to client standard: gemini-3-pro-preview

14 results · % of projects, higher is better. Choose a model to highlight it.Clear highlight

  1. 1 GPT-6 AstraOpenAI 20.8%
  2. 2 Claude Fable 5.1Anthropic 17.9%
  3. 3 Claude Fable 5Anthropic 16.1%
  4. 4 Claude Opus 4.8Anthropic 8.3%
  5. 5 GPT-5.5OpenAI 6.2%
  6. 6 Gemini 3.7 FlashGoogle 5.0%
  7. 7 Claude Opus 4.6Anthropic 4.2%
  8. 8 Claude Opus 4.5Anthropic 3.8%
  9. 9 GPT-5.2 (medium reasoning)OpenAI 2.5%
  10. 10 Claude Sonnet 4.5Anthropic 2.1%
  11. 10 GPT-5.2OpenAI 2.1%
  12. 12 GPT-5OpenAI 1.7%
  13. 13 gemini-3-pro-previewGoogle 1.2%
  14. 14 Gemini 2.5 Pro Preview 06-05Google 0.8%

Full results

Remote Labor Index: projects done to client standard, % of projects, higher is better
#ModelProjects done to client standard
% of projects, higher is better
1 GPT-6 AstraOpenAI
20.8%
2 Claude Fable 5.1Anthropic
17.9%
3 Claude Fable 5Anthropic
16.1%
4 Claude Opus 4.8Anthropic
8.3%
5 GPT-5.5OpenAI
6.2%
6 Gemini 3.7 FlashGoogle
5.0%
7 Claude Opus 4.6Anthropic
4.2%
8 Claude Opus 4.5Anthropic
3.8%
9 GPT-5.2 (medium reasoning)OpenAI
2.5%
10 Claude Sonnet 4.5Anthropic
2.1%
10 GPT-5.2OpenAI
2.1%
12 GPT-5OpenAI
1.7%
13 gemini-3-pro-previewGoogle
1.2%
14 Gemini 2.5 Pro Preview 06-05Google
0.8%

Results as published by Remote Labor Index; we do not re-run them.

What it measures

Share of 240 real paid freelance projects (design, software, data, architecture and more) an agent delivers at a standard a client would accept.

What it does not measure

Not speed or cost; a small set of projects, so a few points either way are noise.

Failures

Failures count against a model: a product with no usable output is a failed listing. They are listed here so you can see why.

Not ranked

  • claude-fable-5_unknown: listed twice

Scale AI and the Center for AI Safety; collected by Epoch AI. Licence: Published with permission (Scale AI and the Center for AI Safety results via Epoch AI).