2026 suite: inbox, AI, accounting.Open app
← Workspace369 Research

Data, with the working shown.

AI Productivity and Time-Savings Benchmarks

Compare experimental AI results for writing, consulting and coding, with model-era limits, measured versus perceived outcomes and reproducible time-index calculations.

The short version

2026 edition

Key takeaways

Each takeaway links to its canonical data row. Reported figures belong to their cited producer; modeled results are labeled as calculations.

  1. Writing experiment: lower task time with AI assistance: 40% (2023 experiment; published July 2023).MIT / Noy and Zhang
  2. Writing experiment: higher evaluated output quality: 18% (2023 experiment; published July 2023).MIT / Noy and Zhang
  3. Consulting experiment: increase in subtasks completed: 12.2% (Study first released in 2023; published-study summary April 2026).Harvard Business School AI Institute / Dell’Acqua and colleagues
  4. Consulting experiment: approximately higher quality on the innovation exercise: 32% (Study first released in 2023; published-study summary April 2026).Harvard Business School AI Institute / Dell’Acqua and colleagues
  5. Strategy task: correct answers without AI: 84.5% (Study first released in 2023; published-study summary April 2026).Harvard Business School AI Institute / Dell’Acqua and colleagues
  6. Strategy task: correct answers with GPT-4 access: 70.6% (Study first released in 2023; published-study summary April 2026).Harvard Business School AI Institute / Dell’Acqua and colleagues
  7. Strategy task: correct answers with GPT-4 and prompt overview: 60% (Study first released in 2023; published-study summary April 2026).Harvard Business School AI Institute / Dell’Acqua and colleagues
  8. Copilot experiment: reduced completion time on one coding assignment: 55.8% (May 15–June 20, 2022).Peng, Kalliamvakou, Cihon and Demirer
  9. METR early-2025 experiment: longer completion time with AI allowed: 19% (February–June 2025).METR
  10. METR participants: perceived AI speedup after the experiment: 20% (February–June 2025).METR

Does AI help with short professional writing tasks?

The writing experiment found less time spent and higher evaluated quality with assistance. Its assignments did not require the full factual precision or organizational context of many real customer-facing documents.

AI productivity benchmarks — 2022–2025 experiments; institutional updates checked through September 2026. Study-specific recruitment; not a representative global worker sample.
Outcome versus controlReported change
Reduction in task time40%Randomized experiment
Increase in evaluated quality18%Randomized experiment

Quality-score improvement and time reduction are different outcomes and must not be added.

Primary source: MIT / Noy and Zhang — July 14, 2023

Can AI improve one consulting task and harm another?

Yes. The study separates tasks that the model could handle from a strategy problem on which it was unreliable. Productivity on the first set does not establish accuracy on the second.

AI productivity benchmarks — 2022–2025 experiments; institutional updates checked through September 2026. Study-specific recruitment; not a representative global worker sample.
Outcome and conditionReported percentage
More subtasks completed on the innovation exercise12.2%Randomized experiment
Higher quality on the innovation exercise, approximately32%Randomized experiment
Correct strategy answer: no AI84.5%Randomized experiment
Correct strategy answer: GPT-4 access70.6%Randomized experiment
Correct strategy answer: GPT-4 plus prompt overview60%Randomized experiment

The quality estimate is approximate. This is historical GPT-4 evidence, not a benchmark of today’s models.

Primary source: Harvard Business School AI Institute / Dell’Acqua and colleagues — April 9, 2026

Why do coding studies report different productivity effects?

A standalone coding assignment and maintenance inside a mature repository are different tasks. The studies also used different model generations and recruitment methods. Their results should remain separate rather than averaged.

AI productivity benchmarks — 2022–2025 experiments; institutional updates checked through September 2026. Study-specific recruitment; not a representative global worker sample.
Study and outcomeReported change
Copilot assignment: reduction in completion time55.8%Randomized experiment
METR: increase in actual completion time19%Randomized experiment
METR: anticipated speedup before the work24%Participant perception
METR: perceived speedup after the work20%Participant perception

METR’s later follow-up reported serious selection and time-measurement problems. Its authors do not treat the follow-up as a reliable estimate of the current effect; the historical slowdown is not a claim about all developers in 2026.

Primary sources: Peng, Kalliamvakou, Cihon and Demirer — February 13, 2023 · METR — July 10, 2025 · METR — February 24, 2026

What does each measured time change mean on the same baseline?

A baseline index makes the arithmetic easier to read without pooling the studies. Each control condition is assigned the same starting index; the resulting assisted index still refers only to its own experiment.

AI productivity benchmarks — 2022–2025 experiments; institutional updates checked through September 2026. Study-specific recruitment; not a representative global worker sample.
ExperimentAssisted time index; control = 100
Writing experiment60Normalized calculation
Copilot assignment44.2Normalized calculation
METR repository tasks119Normalized calculation

Workspace369 normalization, not a new observed result. Lower means less time. A reduction in time is not the same percentage increase in throughput, and no throughput estimate is claimed.

Primary sources: MIT / Noy and Zhang — July 14, 2023 · Peng, Kalliamvakou, Cihon and Demirer — February 13, 2023 · METR — July 10, 2025

How should a business test AI time savings?

Define a repeatable task and an acceptance standard before comparing assisted and unassisted work. Include prompt preparation, fact checking, review and rework. Keep customer-data access and human approval requirements unchanged across the comparison.

Track both time and quality. Do not apply a writing-task percentage to an entire payroll or assume the same result for an autonomous agent.

How this report was built

Methodology and limitations

  1. This is a selected evidence comparison, not a systematic review or meta-analysis. Studies were included for identifiable tasks, original-producer evidence and clearly stated outcome definitions.
  2. The writing and consulting figures use their research institutions’ public summaries of published studies. The coding paper and METR report provide original experimental detail. Working-paper and published versions are not mixed.
  3. Measured outcomes and perceived speedups carry separate evidence labels. The time-index calculation uses only measured completion-time changes; survey beliefs are excluded from it.
  4. Model versions and study dates remain attached to the findings. We checked the METR follow-up and retained its warning about the reliability of newer estimates.

What these numbers cannot tell you

  • Selected experiments cannot establish a universal AI return on investment. Different tasks, control conditions, participants and quality tests prevent a defensible pooled percentage.
  • The sample sizes describe study scope, not independent replications. Vendor involvement in the Copilot study is relevant to interpretation.
  • Historical results do not establish the capabilities of models available today. The normalization does not include implementation costs or demonstrate sustained annual savings.

Freshness and corrections

Maintained by the Workspace369 editorial team. Review quarterly and when a cited producer releases a replacement study. Next editorial review: December 2026. Retain historical model versions and observation dates. The edition date changes only when the evidence or content is substantively reviewed; it does not change the underlying observation period.

First edition: . Data extraction, source attribution and arithmetic checked for this edition. No independent peer review is claimed.

Found an error or a newer primary release? Send a correction with the source and affected statistic. Confirmed corrections should be recorded in the revision history before republishing.

Primary sources and provenance

Every reported numeric cell links directly to its producer. The downloads include exact table or workbook locators, observation periods, access dates, formulas and input references.

  1. MIT / Noy and ZhangMIT research announcement for the published Science writing experiment ↗Published July 14, 2023. Accessed September 12, 2026.Selected factual observations paraphrased with attribution. No source report, chart, participant data or proprietary database is redistributed; source rights remain with its producer.
  2. Harvard Business School AI Institute / Dell’Acqua and colleaguesBack to the Beginnings of AI at Work: institutional review of the published consulting experiment ↗Published April 9, 2026. Accessed September 12, 2026.Selected factual observations paraphrased with attribution. No source report, chart, participant data or proprietary database is redistributed; source rights remain with its producer.
  3. Peng, Kalliamvakou, Cihon and DemirerThe Impact of AI on Developer Productivity: Evidence from GitHub Copilot, arXiv v1 ↗Published February 13, 2023. Accessed September 12, 2026.Selected factual observations paraphrased with attribution. No source report, chart, participant data or proprietary database is redistributed; source rights remain with its producer.
  4. METRMeasuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity ↗Published July 10, 2025. Accessed September 12, 2026.Selected factual observations paraphrased with attribution. No source report, chart, participant data or proprietary database is redistributed; source rights remain with its producer.
  5. METRWe are Changing our Developer Productivity Experiment Design ↗Published February 24, 2026. Accessed September 12, 2026.Selected factual observations paraphrased with attribution. No source report, chart, participant data or proprietary database is redistributed; source rights remain with its producer.

Made to be checked, then cited

How to cite this report

For a source-reported statistic, credit the original publisher and link to the exact row here when using our compilation. For a modeled result, cite Workspace369 and include the assumptions. Linking to this page does not make us the original producer of third-party data.

Workspace369. (2026-09-12). AI Productivity and Time-Savings Benchmarks. https://workspace369.com/research/ai-productivity-time-savings-benchmarks/. Primary sources and observation periods as listed in the report.

The downloads are English-language reference datasets, including on translated pages. Source rights remain with their producers. Attribute Workspace369’s compilation and calculations; consult each source’s reuse terms.