When AI Becomes a Teammate
A P&G field experiment reveals how generative AI changes team performance, expertise integration and decision-making in innovation.
Can one employee working with AI produce work as strong as a two-person human team?
A new field experiment published in Organization Science suggests that this is possible for certain innovation tasks. Yet the most useful conclusion is not that AI makes teams obsolete.
The study shows that generative AI can reproduce some benefits of teamwork, help employees cross functional boundaries and improve the quality of ideas. It also reveals an important limit: producing better options and selecting the best option are not the same capability.
Fabrizio Dell’Acqua and his coauthors describe AI in this role as a cybernetic teammate—not simply a tool being operated, but an artificial participant in the collaborative process.
Four ways of working at P&G
The preregistered field experiment involved 791 R&D and commercial professionals at Procter & Gamble. Participants worked on real challenges involving products, packaging, communications and retail execution within their own business units.
They were randomly assigned to four conditions:
- An individual working without AI
- A two-person team working without AI
- An individual working with AI
- A two-person team working with AI
Each human team paired an R&D professional with a commercial professional. Participants in the AI conditions received one hour of training before working with a GPT-4-based tool delivered through Microsoft Azure.
Independent experts who were blind to the experimental conditions evaluated the solutions for quality, novelty and feasibility. Each solution received more than three evaluations on average.
The setting gives the research unusual practical relevance. These were experienced employees working on genuine organizational problems, and the strongest proposals were suitable for entry into P&G’s innovation pipeline.
An individual with AI matched a human team
Human teams without AI produced solutions 0.24 standard deviations above the individual control group—an improvement the authors estimate at approximately 6.3%.
Individuals using AI performed 0.37 standard deviations, or approximately 9.6%, above the control group. Their output quality was comparable to that of two-person teams working without AI.
AI-enabled teams performed 0.39 standard deviations above the individual control. However, the average-quality difference between teams with and without AI was not statistically conclusive.
The evidence therefore does not show that “AI is better than teams.” It indicates something more specific: on a defined, early-stage innovation task, AI gave individual employees access to some of the performance benefits traditionally associated with collaboration.
AI helped employees cross functional boundaries
Commercial and R&D professionals naturally approach innovation from different directions. One tends to emphasize customer and market value; the other gives greater weight to technical feasibility.
Without AI, participants’ proposals reflected these professional backgrounds. Commercial employees generated more commercially oriented solutions, while R&D employees produced more technical ones.
When participants used AI, this difference largely disappeared. Both groups generated a more balanced mix of technical and commercial ideas.
AI did not necessarily teach employees another profession. It gave them temporary access to perspectives and knowledge outside their usual functional frame.
That distinction matters. Access to expertise is not the same as development of expertise. A commercial manager may produce a more technically informed proposal with AI without acquiring the judgment of an experienced engineer.
Better ideas did not mean better selection
The most consequential finding comes from separating the stages of the innovation process.
Participants first generated five ideas, selected one and then developed that choice into a detailed proposal. This allowed the researchers to distinguish idea generation from evaluative selection.
AI increased the average quality of the initial ideas. It lifted the entire pool rather than merely producing one exceptional option or eliminating variation.
The pattern changed at the selection stage. Human teams without AI selected the highest-quality idea from their five options approximately 50% of the time. In the AI-enabled conditions, the figure was about 37%.
AI users still produced stronger final proposals because they were choosing from a better starting pool. But the underlying mechanism was clear:
AI operated as a quality amplifier, not as a decision enhancer.
The implication is that organizations should not design ideation and evaluation in the same way. AI can broaden and strengthen the option set. Independent human judgment may remain particularly valuable when the organization must decide which option deserves investment.
Human-AI teams showed promise at the top of the distribution
Average performance is not the only objective in innovation. A small number of exceptional ideas can create disproportionate value.
Compared with the individual control group, AI-enabled teams were 9.2 percentage points more likely to produce a proposal ranked in the top 10% of all submissions. Against a control mean of 5.8%, this represented roughly three times the probability of a top-decile result.
The direct difference between teams with and without AI was not statistically decisive, so this should be treated as a promising signal rather than definitive proof of human-AI synergy.
It nevertheless suggests two different organizational objectives:
- If the goal is a consistently higher quality floor, an individual with AI may be a strong configuration.
- If the goal is a rare, outstanding solution, combining complementary human expertise with AI may offer greater potential.
The experience of work also changed
Participants using AI reported greater increases in enthusiasm, energy and excitement, alongside reductions in anxiety and frustration. Individuals working with AI experienced a larger positive-emotion increase than members of human teams working without it.
This does not mean that lower friction is always better. Constructive disagreement can expose weak assumptions and stimulate creative exploration. An agreeable artificial teammate may make work feel easier while removing some of the tension that supports rigorous evaluation.
Organizations should therefore look beyond satisfaction and adoption. They also need to examine whether AI-supported work preserves challenge, dissent and independent judgment.
Five principles for redesigning innovation work
The findings suggest five practical principles.
1. Separate generation from selection
Do not ask a team to generate options with AI and immediately choose among them in the same flow. Build the option set first. Evaluate it later against explicit criteria.
2. Surface functional perspectives before involving AI
Ask technical and commercial contributors to record their initial views independently. AI can then help connect those perspectives without prematurely smoothing away meaningful differences.
3. Do not treat human teams as a cost category alone
Matching the output quality of a two-person team does not mean reproducing every function of that team. Human collaboration also provides challenge, accountability, learning and organizational memory.
4. Train for the workflow, not merely the prompt
Participants achieved meaningful gains after a standardized one-hour session. Sustainable organizational capability, however, requires more than prompt technique. Employees need to know when to use AI, how to test its output and where decision ownership remains human.
5. Measure work quality instead of tool adoption
Useful indicators include:
- The quality and diversity of options
- Integration of knowledge across functions
- Accuracy in identifying the strongest option
- The reasoning behind the final decision
- Development of independent employee capability over time
Important boundaries on the findings
The experiment took place in one consumer-goods company, used one AI model and involved one-day virtual workshops. Human teams were pairs of largely unfamiliar participants rather than established groups with shared history and trust.
The task also concerned early-stage product development. The results should not automatically be generalized to investment decisions, hiring, risk management or crisis response.
The opportunity lies in work design
This research suggests that generative AI is becoming more than an individual productivity tool. It can perform some collaborative functions, bridge professional boundaries and improve the raw material of innovation.
But better ideas do not automatically produce better decisions. Organizational value will come from redesigning the division of labour among generation, evaluation, dissent and decision ownership—not from inserting an AI licence into an unchanged process.
Stratify’s AI-Empowered Leadership program helps leadership teams examine this division through realistic innovation challenges and measurable experiments.
The defining question is no longer simply whether employees use AI. It is which parts of collective thinking the organization wants AI to perform—and which must remain distinctly human.
Source: Fabrizio Dell’Acqua et al., “The Cybernetic Teammate: A Field Experiment on Generative AI and Teamwork”, Organization Science, 2026.
← All posts
