Sixty-four percent, then ninety-three percent. Those are the completion rates Claude Sonnet 4 achieved on real data science and analytics projects, first working alone and then after a human expert spent roughly twenty minutes reviewing its output. The gap between those two numbers is the clearest picture anyone has published of where freelance work is actually heading, and it points somewhere more specific than the usual advice about staying adaptable.
The study, and why its design matters
Upwork built a benchmark it calls the Human+Agent Productivity Index, testing Gemini 2.5 Pro, GPT-5 and Claude Sonnet 4 against more than 300 real client projects posted to its platform across writing, data science, web development, engineering, sales and translation. The full paper went through double-blind peer review and was accepted to NeurIPS.
Two design choices make the results more useful than the usual benchmark. First, these were jobs real clients paid for, not synthetic tasks. Second, and more revealing, Upwork deliberately picked easy ones. The projects were priced under $500 and represent less than 6 percent of Upwork’s total gross services volume. Andrew Rabinovich, Upwork’s CTO and head of AI and machine learning, told VentureBeat that the team chose simpler tasks specifically to give the agents traction, because further up the value chain “we don’t think they can solve them at all, even to scratch the surface.”
So the numbers below are the optimistic case. They describe agents working on the problems most favourable to them.
Where the agents held up and where they collapsed
Performance split along a clean line: work with objectively verifiable answers went well, work requiring judgment did not.
Claude Sonnet 4 completed 68 percent of web development jobs and 64 percent of data science projects unaided. Gemini 2.5 Pro reached 74 percent on certain technical tasks. Rabinovich’s explanation is blunt, that most coding tasks resemble each other, which is why coding agents have improved so fast.
Qualitative work told a different story. Gemini 2.5 Pro managed 17 percent on sales and marketing projects working alone. GPT-5 reached 30 percent on engineering and architecture tasks. Website layouts, marketing copy and translation requiring cultural nuance all faltered without expert direction.
That spread is the actionable part. If what you sell is template-driven and has a single correct output, an agent already completes roughly two thirds of it on the easy end of the market. If what you sell requires deciding what “good” means for a particular client, the agent is starting below one in five.
Human feedback moved the needle more than model choice
The headline finding is that project completion improved by up to 70 percent when agents worked with human experts rather than alone. The per-category detail is more instructive.
Gemini’s sales and marketing completion rose from 17 to 31 percent with human input. GPT-5’s engineering and architecture work climbed from 30 to 50 percent. Writing and translation, the categories where agents did worst alone, gained up to 17 percentage points per feedback cycle, and engineering and architecture projects improved by as much as 23 percentage points with human oversight. The improvement compounded across rounds rather than arriving all at once.
Each review cycle cost about twenty minutes of expert time. Rabinovich described the total time investment as “orders of magnitude different” from a human doing the work alone, with projects that might take days delivered in hours through alternating cycles of agent work and expert correction.
The uncomfortable read on this
There is a genuinely encouraging interpretation and a genuinely worrying one, and both are supported by the data.
The encouraging one is Upwork’s: expertise is what makes agents useful, so expertise gets more valuable, not less. AI-related work on the platform grew 53 percent year over year in the third quarter of 2025. Rabinovich argues simpler tasks get automated while jobs grow more complex in the number of tasks they contain, so freelancer earnings rise.
The worrying one is about pricing. If a client can hire you for twenty minutes of review instead of two days of production, your role shifts from producing work to validating it, and validation has historically commanded lower rates than creation. The study measured completion against rubrics, not what a client would actually pay for. The paper is careful about this, stating that rubric-based completion rates should not be read as a measure of whether an agent would be paid in a real marketplace, only of its ability to fulfil explicitly defined requests. An agent can technically satisfy every stated requirement and still produce work a client rejects.
Which of these plays out depends heavily on how you position what you sell, a tension we examined in our look at which AI-adjacent freelance work pays more and which is getting cheaper.
What to do with this if you freelance
The data supports a few concrete moves rather than a general posture.
Sell the judgment layer explicitly. The categories where agents fail hardest alone are exactly where human feedback produces the largest gains. Direction, editorial judgment and cultural context are the scarce inputs. Price them as the service rather than treating them as something bundled free with delivery.
Treat agent supervision as a named skill. Rabinovich points to emerging work in designing human-machine workflows, guiding agents to improve, and verifying that agentic output is actually correct. These barely existed as categories two years ago. The demand pattern is already visible on platforms, as we found when looking at what Claude Code specialist work actually involves.
Reconsider fully templated offerings. If your service produces one correct output from a predictable input and sits at the low end of the price range, the benchmark says an agent finishes most of it today on the favourable cases. That is the part of a service catalogue worth rebuilding first, and the practical route is usually toward selling AI automation as a service rather than competing with it.
Watch how clients are being routed to you. Upwork is building Uma, a meta-orchestration agent intended to sit between clients and talent, analysing requirements, deciding which tasks need humans, and coordinating the work. The company has said clients would increasingly interact with Uma rather than hiring freelancers directly. That is a change in how work reaches you, and it rewards freelancers who are legible to a routing system, a shift already showing up in the way clients now search for AI roles rather than AI tasks.
The number to keep
On deliberately simplified, sub-$500 projects representing under 6 percent of a major platform’s volume, the best agents finished somewhere between 17 and 74 percent of jobs alone depending on category. Every figure above comes from the easy end of the market, measured against explicit requirements rather than client satisfaction.
Read against the alarm about AI replacing freelancers wholesale, that is a considerably more modest picture. Read against anyone selling straightforward, template-driven work at the bottom of the price range, it is a clear instruction to move.




