The short version
2026 editionKey takeaways
Each takeaway links to its canonical data row. Reported figures belong to their cited producer; modeled results are labeled as calculations.
- Writing experiment: lower task time with AI assistance: 40% (2023 experiment; published July 2023).MIT / Noy and Zhang
- Writing experiment: higher evaluated output quality: 18% (2023 experiment; published July 2023).MIT / Noy and Zhang
- Consulting experiment: increase in subtasks completed: 12.2% (Study first released in 2023; published-study summary April 2026).Harvard Business School AI Institute / Dell’Acqua and colleagues
- Consulting experiment: approximately higher quality on the innovation exercise: 32% (Study first released in 2023; published-study summary April 2026).Harvard Business School AI Institute / Dell’Acqua and colleagues
- Strategy task: correct answers without AI: 84.5% (Study first released in 2023; published-study summary April 2026).Harvard Business School AI Institute / Dell’Acqua and colleagues
- Strategy task: correct answers with GPT-4 access: 70.6% (Study first released in 2023; published-study summary April 2026).Harvard Business School AI Institute / Dell’Acqua and colleagues
- Strategy task: correct answers with GPT-4 and prompt overview: 60% (Study first released in 2023; published-study summary April 2026).Harvard Business School AI Institute / Dell’Acqua and colleagues
- Copilot experiment: reduced completion time on one coding assignment: 55.8% (May 15–June 20, 2022).Peng, Kalliamvakou, Cihon and Demirer
- METR early-2025 experiment: longer completion time with AI allowed: 19% (February–June 2025).METR
- METR participants: perceived AI speedup after the experiment: 20% (February–June 2025).METR
Does AI help with short professional writing tasks?
The writing experiment found less time spent and higher evaluated quality with assistance. Its assignments did not require the full factual precision or organizational context of many real customer-facing documents.
| Outcome versus control | Reported change |
|---|---|
| Reduction in task time | 40%Randomized experiment |
| Increase in evaluated quality | 18%Randomized experiment |
Quality-score improvement and time reduction are different outcomes and must not be added.
Primary source: MIT / Noy and Zhang — July 14, 2023
Can AI improve one consulting task and harm another?
Yes. The study separates tasks that the model could handle from a strategy problem on which it was unreliable. Productivity on the first set does not establish accuracy on the second.
| Outcome and condition | Reported percentage |
|---|---|
| More subtasks completed on the innovation exercise | 12.2%Randomized experiment |
| Higher quality on the innovation exercise, approximately | 32%Randomized experiment |
| Correct strategy answer: no AI | 84.5%Randomized experiment |
| Correct strategy answer: GPT-4 access | 70.6%Randomized experiment |
| Correct strategy answer: GPT-4 plus prompt overview | 60%Randomized experiment |
The quality estimate is approximate. This is historical GPT-4 evidence, not a benchmark of today’s models.
Primary source: Harvard Business School AI Institute / Dell’Acqua and colleagues — April 9, 2026
Why do coding studies report different productivity effects?
A standalone coding assignment and maintenance inside a mature repository are different tasks. The studies also used different model generations and recruitment methods. Their results should remain separate rather than averaged.
| Study and outcome | Reported change |
|---|---|
| Copilot assignment: reduction in completion time | 55.8%Randomized experiment |
| METR: increase in actual completion time | 19%Randomized experiment |
| METR: anticipated speedup before the work | 24%Participant perception |
| METR: perceived speedup after the work | 20%Participant perception |
METR’s later follow-up reported serious selection and time-measurement problems. Its authors do not treat the follow-up as a reliable estimate of the current effect; the historical slowdown is not a claim about all developers in 2026.
Primary sources: Peng, Kalliamvakou, Cihon and Demirer — February 13, 2023 · METR — July 10, 2025 · METR — February 24, 2026
What does each measured time change mean on the same baseline?
A baseline index makes the arithmetic easier to read without pooling the studies. Each control condition is assigned the same starting index; the resulting assisted index still refers only to its own experiment.
| Experiment | Assisted time index; control = 100 |
|---|---|
| Writing experiment | 60Normalized calculation |
| Copilot assignment | 44.2Normalized calculation |
| METR repository tasks | 119Normalized calculation |
Workspace369 normalization, not a new observed result. Lower means less time. A reduction in time is not the same percentage increase in throughput, and no throughput estimate is claimed.
Primary sources: MIT / Noy and Zhang — July 14, 2023 · Peng, Kalliamvakou, Cihon and Demirer — February 13, 2023 · METR — July 10, 2025
How should a business test AI time savings?
Define a repeatable task and an acceptance standard before comparing assisted and unassisted work. Include prompt preparation, fact checking, review and rework. Keep customer-data access and human approval requirements unchanged across the comparison.
Track both time and quality. Do not apply a writing-task percentage to an entire payroll or assume the same result for an autonomous agent.
How this report was built
Methodology and limitations
- This is a selected evidence comparison, not a systematic review or meta-analysis. Studies were included for identifiable tasks, original-producer evidence and clearly stated outcome definitions.
- The writing and consulting figures use their research institutions’ public summaries of published studies. The coding paper and METR report provide original experimental detail. Working-paper and published versions are not mixed.
- Measured outcomes and perceived speedups carry separate evidence labels. The time-index calculation uses only measured completion-time changes; survey beliefs are excluded from it.
- Model versions and study dates remain attached to the findings. We checked the METR follow-up and retained its warning about the reliability of newer estimates.
What these numbers cannot tell you
- Selected experiments cannot establish a universal AI return on investment. Different tasks, control conditions, participants and quality tests prevent a defensible pooled percentage.
- The sample sizes describe study scope, not independent replications. Vendor involvement in the Copilot study is relevant to interpretation.
- Historical results do not establish the capabilities of models available today. The normalization does not include implementation costs or demonstrate sustained annual savings.
Freshness and corrections
Maintained by the Workspace369 editorial team. Review quarterly and when a cited producer releases a replacement study. Next editorial review: December 2026. Retain historical model versions and observation dates. The edition date changes only when the evidence or content is substantively reviewed; it does not change the underlying observation period.
First edition: . Data extraction, source attribution and arithmetic checked for this edition. No independent peer review is claimed.
Found an error or a newer primary release? Send a correction with the source and affected statistic. Confirmed corrections should be recorded in the revision history before republishing.
Primary sources and provenance
Every reported numeric cell links directly to its producer. The downloads include exact table or workbook locators, observation periods, access dates, formulas and input references.
- MIT / Noy and ZhangMIT research announcement for the published Science writing experiment ↗Published July 14, 2023. Accessed September 12, 2026.Selected factual observations paraphrased with attribution. No source report, chart, participant data or proprietary database is redistributed; source rights remain with its producer.
- Harvard Business School AI Institute / Dell’Acqua and colleaguesBack to the Beginnings of AI at Work: institutional review of the published consulting experiment ↗Published April 9, 2026. Accessed September 12, 2026.Selected factual observations paraphrased with attribution. No source report, chart, participant data or proprietary database is redistributed; source rights remain with its producer.
- Peng, Kalliamvakou, Cihon and DemirerThe Impact of AI on Developer Productivity: Evidence from GitHub Copilot, arXiv v1 ↗Published February 13, 2023. Accessed September 12, 2026.Selected factual observations paraphrased with attribution. No source report, chart, participant data or proprietary database is redistributed; source rights remain with its producer.
- METRMeasuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity ↗Published July 10, 2025. Accessed September 12, 2026.Selected factual observations paraphrased with attribution. No source report, chart, participant data or proprietary database is redistributed; source rights remain with its producer.
- METRWe are Changing our Developer Productivity Experiment Design ↗Published February 24, 2026. Accessed September 12, 2026.Selected factual observations paraphrased with attribution. No source report, chart, participant data or proprietary database is redistributed; source rights remain with its producer.
Made to be checked, then cited
How to cite this report
For a source-reported statistic, credit the original publisher and link to the exact row here when using our compilation. For a modeled result, cite Workspace369 and include the assumptions. Linking to this page does not make us the original producer of third-party data.
Workspace369. (2026-09-12). AI Productivity and Time-Savings Benchmarks. https://workspace369.com/research/ai-productivity-time-savings-benchmarks/. Primary sources and observation periods as listed in the report.
The downloads are English-language reference datasets, including on translated pages. Source rights remain with their producers. Attribute Workspace369’s compilation and calculations; consult each source’s reuse terms.