New Benchmark Questions Workplace Readiness of AI Agents
Nearly two years ago, Microsoft CEO Satya Nadella forecasted that AI would reshape knowledge work — the domain of lawyers, investment bankers, librarians, accountants, IT professionals, and similar white-collar roles.
However, despite significant advancements in foundation models, a transformation in knowledge work has been slow to materialize. While models excel at deep research and agentic planning, most white-collar professions have seen relatively little disruption for reasons that remain unclear.
This stands as one of AI's great puzzles. New research from training-data leader Mercor is now providing crucial insights.
The study evaluates how top AI models handle genuine white-collar tasks from consulting, investment banking, and law. This led to the creation of the APEX-Agents benchmark — and currently, every AI lab is failing. When presented with queries from real professionals, even the best models correctly answered fewer than a quarter. Most often, they returned an incorrect response or none at all.
According to Mercor CEO Brendan Foody, who contributed to the research, the models' primary weakness was synthesizing information across multiple domains — a core component of human knowledge work.
"A key innovation in this benchmark is that we constructed a complete environment modeled on real professional services," Foody explained to TechCrunch. "Our work doesn't involve a single person handing us all the context in one place. In reality, you operate across Slack, Google Drive, and various other tools." For many agentic AI models, that type of cross-domain reasoning remains inconsistent.

Screenshot The test scenarios were sourced from actual professionals on Mercor's expert marketplace, who designed the queries and defined the criteria for a successful answer. Reviewing the publicly available questions on Hugging Face reveals the complexity of these tasks.
Techcrunch event Disrupt 2026 Tickets: One-time offer
Tickets are now available! Save up to $680 during this limited-time offer and be among the first 500 registrants to receive 50% off a +1 pass. TechCrunch Disrupt gathers top leaders from Google Cloud, Netflix, Microsoft, Box, a16z, Hugging Face, and more for 250+ sessions aimed at fueling growth and sharpening your competitive edge. Connect with hundreds of innovative startups and participate in curated networking that drives deals, insights, and inspiration.
Disrupt 2026 Tickets: One-time offer
Tickets are now available! Save up to $680 during this limited-time offer and be among the first 500 registrants to receive 50% off a +1 pass. TechCrunch Disrupt gathers top leaders from Google Cloud, Netflix, Microsoft, Box, a16z, Hugging Face, and more for 250+ sessions aimed at fueling growth and sharpening your competitive edge. Connect with hundreds of innovative startups and participate in curated networking that drives deals, insights, and inspiration.
San Francisco | October 13-15, 2026 REGISTER NOW One example from the "Law" section asks:
During the first 48 minutes of an EU production outage, Northstar's engineering team exported one or two bundled sets of EU production event logs containing personal data to a U.S. analytics vendor… According to Northstar's own policies, can it reasonably treat these one or two log exports as compliant with Article 49?
The correct answer is yes, but arriving at it requires a detailed analysis of both the company's internal policies and relevant EU privacy regulations.
Such a question could challenge even a knowledgeable human, but the researchers aimed to simulate real professional work. An LLM that can reliably answer these queries could potentially replace many practicing lawyers. "This is arguably the most significant economic topic today," Foody told TechCrunch. "The benchmark accurately reflects the actual work these professionals perform."
OpenAI previously attempted to gauge professional skills with its GDPval benchmark, but the APEX-Agents test differs meaningfully. While GDPval assesses broad general knowledge across many fields, APEX-Agents measures a system's capability to execute sustained tasks within a select few high-value professions. This makes it more challenging for models and more directly relevant to the potential for job automation.
Although no model is ready to step in as an investment banker, some clearly came closer than others. Gemini 3 Flash led the group with 24% one-shot accuracy, closely followed by GPT-5.2 at 23%. Opus 4.5, Gemini 3 Pro, and GPT-5 all scored around 18%.
While these initial results are underwhelming, the AI field has a track record of rapidly overcoming difficult benchmarks. With the APEX-Agents test now public, it poses an open challenge to AI labs confident they can improve — an outcome Foody fully anticipates in the coming months.
"The pace of improvement is remarkably fast," he told TechCrunch. "Currently, it's fair to compare the technology to an intern who gets it right a quarter of the time, but last year it was an intern who succeeded 5 or 10% of the time. That rate of year-over-year progress can create substantial impact very quickly."
Related article
Suno to Watermark Songs Amid Legal Battles
Suno, the platform enabling users to generate AI-created music, has unveiled new features to label platform-produced tracks, restrict downloads, and update community standards to curb unauthorized replicas. These updates arrive as Suno confronts mult
Musk Admits Grok Build Leaked User Code, Promises to Erase All Historical Data
Elon Musk directly addressed the privacy controversy surrounding Grok Build, beginning with a simple "True" to confirm the incident's validity. He pledged that all user data previously uploaded to SpaceXAI would be permanently erased, stating, "not a
U.S. Stocks Hit Historic Milestone as AI and Aerospace Giants Prepare for Trillion-Dollar Debut
Elon Musk, Sam Altman, and Dario Amodei, three titans of the technology sector, are advancing toward initial public offerings for their respective ventures. With SpaceX, OpenAI, and Anthropic—three industry behemoths nearing trillion-dollar valuation
Related Special Topic Recommendations
Comments (0)
0/500
Nearly two years ago, Microsoft CEO Satya Nadella forecasted that AI would reshape knowledge work — the domain of lawyers, investment bankers, librarians, accountants, IT professionals, and similar white-collar roles.
However, despite significant advancements in foundation models, a transformation in knowledge work has been slow to materialize. While models excel at deep research and agentic planning, most white-collar professions have seen relatively little disruption for reasons that remain unclear.
This stands as one of AI's great puzzles. New research from training-data leader Mercor is now providing crucial insights.
The study evaluates how top AI models handle genuine white-collar tasks from consulting, investment banking, and law. This led to the creation of the APEX-Agents benchmark — and currently, every AI lab is failing. When presented with queries from real professionals, even the best models correctly answered fewer than a quarter. Most often, they returned an incorrect response or none at all.
According to Mercor CEO Brendan Foody, who contributed to the research, the models' primary weakness was synthesizing information across multiple domains — a core component of human knowledge work.
"A key innovation in this benchmark is that we constructed a complete environment modeled on real professional services," Foody explained to TechCrunch. "Our work doesn't involve a single person handing us all the context in one place. In reality, you operate across Slack, Google Drive, and various other tools." For many agentic AI models, that type of cross-domain reasoning remains inconsistent.

The test scenarios were sourced from actual professionals on Mercor's expert marketplace, who designed the queries and defined the criteria for a successful answer. Reviewing the publicly available questions on Hugging Face reveals the complexity of these tasks.
Techcrunch eventDisrupt 2026 Tickets: One-time offer
Tickets are now available! Save up to $680 during this limited-time offer and be among the first 500 registrants to receive 50% off a +1 pass. TechCrunch Disrupt gathers top leaders from Google Cloud, Netflix, Microsoft, Box, a16z, Hugging Face, and more for 250+ sessions aimed at fueling growth and sharpening your competitive edge. Connect with hundreds of innovative startups and participate in curated networking that drives deals, insights, and inspiration.
Disrupt 2026 Tickets: One-time offer
Tickets are now available! Save up to $680 during this limited-time offer and be among the first 500 registrants to receive 50% off a +1 pass. TechCrunch Disrupt gathers top leaders from Google Cloud, Netflix, Microsoft, Box, a16z, Hugging Face, and more for 250+ sessions aimed at fueling growth and sharpening your competitive edge. Connect with hundreds of innovative startups and participate in curated networking that drives deals, insights, and inspiration.
San Francisco | October 13-15, 2026 REGISTER NOWOne example from the "Law" section asks:
During the first 48 minutes of an EU production outage, Northstar's engineering team exported one or two bundled sets of EU production event logs containing personal data to a U.S. analytics vendor… According to Northstar's own policies, can it reasonably treat these one or two log exports as compliant with Article 49?
The correct answer is yes, but arriving at it requires a detailed analysis of both the company's internal policies and relevant EU privacy regulations.
Such a question could challenge even a knowledgeable human, but the researchers aimed to simulate real professional work. An LLM that can reliably answer these queries could potentially replace many practicing lawyers. "This is arguably the most significant economic topic today," Foody told TechCrunch. "The benchmark accurately reflects the actual work these professionals perform."
OpenAI previously attempted to gauge professional skills with its GDPval benchmark, but the APEX-Agents test differs meaningfully. While GDPval assesses broad general knowledge across many fields, APEX-Agents measures a system's capability to execute sustained tasks within a select few high-value professions. This makes it more challenging for models and more directly relevant to the potential for job automation.
Although no model is ready to step in as an investment banker, some clearly came closer than others. Gemini 3 Flash led the group with 24% one-shot accuracy, closely followed by GPT-5.2 at 23%. Opus 4.5, Gemini 3 Pro, and GPT-5 all scored around 18%.
While these initial results are underwhelming, the AI field has a track record of rapidly overcoming difficult benchmarks. With the APEX-Agents test now public, it poses an open challenge to AI labs confident they can improve — an outcome Foody fully anticipates in the coming months.
"The pace of improvement is remarkably fast," he told TechCrunch. "Currently, it's fair to compare the technology to an intern who gets it right a quarter of the time, but last year it was an intern who succeeded 5 or 10% of the time. That rate of year-over-year progress can create substantial impact very quickly."
Suno to Watermark Songs Amid Legal Battles
Suno, the platform enabling users to generate AI-created music, has unveiled new features to label platform-produced tracks, restrict downloads, and update community standards to curb unauthorized replicas. These updates arrive as Suno confronts mult
Musk Admits Grok Build Leaked User Code, Promises to Erase All Historical Data
Elon Musk directly addressed the privacy controversy surrounding Grok Build, beginning with a simple "True" to confirm the incident's validity. He pledged that all user data previously uploaded to SpaceXAI would be permanently erased, stating, "not a
U.S. Stocks Hit Historic Milestone as AI and Aerospace Giants Prepare for Trillion-Dollar Debut
Elon Musk, Sam Altman, and Dario Amodei, three titans of the technology sector, are advancing toward initial public offerings for their respective ventures. With SpaceX, OpenAI, and Anthropic—three industry behemoths nearing trillion-dollar valuation





Home






