In late 2025, a research team led by Mantas Mazeika, affiliated with the Center for AI Safety and working with collaborators across industry and academia, introduced a new benchmark called the Remote Labor Index (RLI).
RLI asks a simple and practical question: can AI agents autonomously complete the same kinds of remote work tasks that humans are paid to do?
To answer this properly, RLI is anchored in real economic activity and outcomes. All projects are drawn from professional freelance platforms, primarily Upwork, and represent work that a human has already completed, delivered, and been paid for.
What RLI does is straightforward. For each project, it provides the original human brief, all input files, and the human gold-standard deliverable. An AI agent is then asked to attempt the same project without human intervention. Independent evaluators judge whether a reasonable client would accept the agent’s output.
The result is consistent across models: Even the strongest current AI agents complete only a very small fraction of these projects at professional quality. That fraction peaks at roughly 2.5% across the benchmark.
What RLI actually measures
RLI is built from 240 real freelance projects spanning 23 Upwork domains. These are projects completed by professionals, often involving complex deliverables and many interdependent files.
Each project includes:
-
a full written client brief
-
all necessary input files
-
The gold-standard deliverable created by a human professional
-
reported data on effort and economic value
The projects range across:
-
3D modeling, CAD, and product design
-
video editing and animation
-
architecture plans and rendering
-
browser games and interactive prototypes
-
data dashboards and analytic visualizations
-
long technical reports and formatted research documents
Across the dataset, the mean human completion time was approximately 28.9 hours, with a median of 11.5 hours. The average project value was around $632.
In total, this represents more than 6,000 hours of verified human work and roughly $144,000 in real economic value.
This matters because it reflects real remote labor, not an academic test or a collection of isolated tasks.
Which AI agents were evaluated
The Remote Labor Index has been used to compare a set of frontier autonomous AI systems available in late 2025 and early 2026. Publicly referenced models include:
-
GPT-5 agent setups
-
Claude Sonnet 4.5
-
Gemini 2.5 Pro
-
Grok-4
-
ChatGPT Agent frameworks
-
Manus, which is widely reported as the top performer on the leaderboard
All systems were evaluated on the same project set, under the same input conditions.
Performance is measured primarily through automation rate, defined as the percentage of projects where the AI deliverable meets or exceeds the quality of the human benchmark.
The highest-performing agent achieves an automation rate of roughly 2.5%. Other models cluster slightly below that.
This means that approximately 97.5% of real freelance work in the benchmark remains beyond current AI agents’ end-to-end capabilities when strict acceptance criteria are applied.
Where agents succeed and fail
RLI does not suggest a binary outcome. It reveals a pattern that aligns with what many practitioners have already observed: AI is not equally competent in all areas.
AI agents perform comparatively better on:
-
scripted audio tasks such as simple editing or vocal separation
-
basic image generation and simple graphics
-
simple data visualization prototypes
-
drafting long documents or templated reports
-
isolated code snippets under narrow constraints
These tasks tend to be narrow in scope, with limited dependencies and clear validation criteria.
The most consistent failure modes involve:
-
multi-asset deliverables with interconnected parts
-
file handling and export in client-required formats
-
consistency across views, versions, or outputs
-
logical completeness, including missing components of the brief
-
professional quality in design, layout, integration, or user experience
These are the issues that arise when an agent is tasked with running the entire project.
Why adoption has outpaced economic output
Large language models are being adopted quickly because they remove friction from everyday work. They help with writing, research, summarization, drafting, and exploration, all common steps across knowledge work, so usage spreads even while the tools remain uneven.
As of February 2026, end-to-end delivery is still unreliable.
Most economic value shows up when work becomes a finished, accepted outcome. A project that ships. An asset that can actually be used.
The Remote Labor Index reflects this clearly. Even when agents perform well on individual steps, projects tend to break somewhere in between. Files don’t line up, constraints are missed, outputs drift, and someone still has to notice and fix the result.
What scales first is speed rather than completion. Drafts multiply, iterations accelerate, and work moves faster inside workflows, while the final handoff remains human.
That gap between faster work and finished work explains why adoption has moved faster than economic growth.
Why this matters for practical work
In fields like marketing and the remote work economy, the distinction matters.
Partial automation, such as writing, analysis, and summarization, already delivers real value and explains the speed of adoption. These capabilities sit inside many workflows and remove friction where it is easiest to remove.
End-to-end project automation, where a brief is fulfilled to a professional standard without follow-up, remains a much harder frontier. That is where coordination, judgment, and acceptance still concentrate.
The Remote Labor Index provides a shared, empirical way to track progress across that boundary as models improve. It helps keep strategy grounded in measurable capability rather than hype.
Till next time 👋
Ilias