Over the past two years, most professionals have seen some version of this sentence: “A paper in Nature proves that hybrid work cuts attrition by a third without affecting performance.” It shows up in executives’ reposts, in HR proposals, and in countless articles about the future of work.
The claim comes from a paper published in June 2024 in Nature, by Stanford economist Nicholas Bloom, Ruobing Han of the Chinese University of Hong Kong, Shenzhen, and James Liang, co-founder and chairman of Trip.com Group (formerly Ctrip). The title leaves little room for interpretation: “Hybrid working from home improves retention without damaging performance.”
As an organization design consultant, my research and project experience left me unable to take that conclusion at face value. So I went through the Nature paper, its 52-page NBER working-paper predecessor, and the public data.
I do not dispute the experimental results. The experiment itself is not complicated, and the statistical treatment is, on the whole, sound.
What I question is this: how far do the results actually sit from the claim now in circulation?
From “this group’s quit rate fell by 2.4 percentage points over a few months” to “hybrid work improves retention”; from “individual performance grades did not change” to “without damaging performance”; and from an experiment run in Shanghai in 2021 across two departments to a Nature title with no qualifier for population, place, duration, or role. Each step seems to say only a little more, but the claim the title ends up making is far broader than what the experiment can actually answer.
This article does not challenge the data. It challenges the scale of the inference.
What did the experiment actually do?
In the summer of 2021, Trip.com piloted hybrid work in two business units, Airfare and IT. A total of 1,612 employees were eligible, all with a bachelor’s degree or higher; interns and new hires still on probation were excluded in advance. Trip.com had roughly 35,000 employees at the time, so the experiment covered about 4.6% of them.
The company first called for volunteers. After one email and two rounds of reminders, only 518 signed up. Management suspected that employees feared volunteering for work from home would be read as a lack of ambition, so the remaining 1,094 “non-volunteers” were brought into the experiment as well.
Assignment was by odd or even birthdate. Odd-birthdate employees received an option: they could work from home on Wednesdays and Fridays. Even-birthdate employees continued to come in five days a week. Note the wording: odd-birthdate employees “received the option of two days at home per week,” not “were required to work from home two days per week.” Actual take-up fell below the ceiling; converted, the treatment group averaged roughly one day at home per week, mostly Friday.
The two cohorts started at different times. Volunteers began the week of August 9; the rest began the week of September 13. The experiment was planned for six months, but in January 2022 Covid appeared at Trip.com’s Shanghai headquarters and the company immediately allowed everyone to work from home. The original control condition vanished, the experiment effectively ended early on January 21, and the attrition count closed on January 23.
So the first cohort was under randomized control for about five and a half months, the second for a little over four. Not a full six months, and certainly not two years.
During the experiment, the cumulative quit rate was 7.2% in the control group and 4.8% in the treatment group, a difference of 2.4 percentage points; measured against the control base, quits “fell by a third.” Trip.com runs performance reviews every six months, with letter grades from D to A. The four review rounds and promotion data reported in the paper show no clear difference between the groups, and lines of code submitted by the 653 engineers whose main work is coding show no adverse change.
Up to this point, the results present no obvious problem. The problems begin with the interpretation.
The first leap: from a short-term quit gap to “improves retention”
“Not quitting” within four to five and a half months is not deciding to stay for the long term
Both 4.8% and 7.2% are cumulative quit rates over four to five and a half months, not annualized rates. Quitting is also a low-frequency, lagged, and seasonal event: bonus payouts, performance reviews, promotion decisions, the Chinese New Year, and hiring season all affect when an employee actually hands in a resignation.
The experiment can show that people who received the hybrid option quit less over four to five and a half months. It cannot answer a different question: were these people retained for the long run, or did they merely postpone leaving by a few months? If hybrid work became a permanent policy, would the gap persist, widen, or disappear? The experiment cannot say.
If the goal is a conclusion about a long-term working arrangement, I would require randomized control to cover at least two consecutive full years or performance cycles, with a contemporaneous control group maintained throughout. One full year passes through the bonus, review, promotion, and hiring cycle only once; two consecutive years make it possible to compare the same month and the same management checkpoint across different years, to see whether the between-group gap recurs, and to begin distinguishing “fewer quits” from “delayed quits.”
The paper did track reviews and promotions for two years afterward, but that is not the same as extending the experiment to two years. After January 2022, Covid had already broken the original office condition; after February 14, hybrid work was rolled out company-wide and the former control group received the same option. The follow-up data can check whether the early randomized experience left delayed consequences such as a promotion penalty; it cannot keep comparing “long-term hybrid” against “long-term office.” Two years of tracking is not two years of randomized control.
“Experiencing the same environment at the same time” is not “being affected the same way”
There is a more critical timing problem. The experiment landed precisely in a period when Covid, the travel industry, and labor-market expectations were all shifting, and it ultimately lost its control group early to an unexpected event.
We cannot assume, just because both groups lived through the same economic and Covid shifts at the same time, that those shifts moved the two groups’ mindsets in the same direction and by the same amount.
The treatment group may have read hybrid work as an extra benefit or a sign of the company’s goodwill. The control group may have reacted differently to being temporarily denied the option, or may have postponed quitting in anticipation of the policy being extended. A shifting economy may have made both groups more risk-averse, but not necessarily to the same degree.
What the experiment observed, therefore, is the combined effect of “the hybrid option × the external environment at the time.” Randomization can identify the overall policy effect within that specific environment; it cannot tell us whether the effect would still be 2.4 percentage points in an environment of plentiful jobs, active job-hopping, and no Covid risk, still less prove that “a third” is a stable property carried by hybrid work itself.
The experiment measured the price of an umbrella in the rainy season, not its year-round average.
“A third” has no uniform meaning across industries
What a business decision actually has to weigh is absolute impact, replacement cost, and implementation cost.
Take a company with a 30% quit rate over the same window. Transplant Trip.com’s absolute difference of 2.4 percentage points and you get 30% falling to 27.6%; transplant the relative “one third” and you get 20%. The two arithmetics mean completely different things for the business, and this experiment offers no evidence about which one to transplant. The safer answer: neither can be moved over directly.
Moreover, in industries or roles where turnover is already high, both 4.8% and 7.2% may already be excellent short-term numbers, and a marginal 2.4 points may not deserve a place on management’s agenda. For scarce specialist roles with very high replacement cost, the same 2.4 points may be worth a great deal. Statistical significance does not automatically mean managerial significance; managerial significance depends on the company’s own baseline attrition, headcount, cost of turnover, and the coordination cost that hybrid work brings.
The second leap: from “grades didn’t change” to “performance isn’t damaged”
The experiment measured individual reviews, not organizational performance
When a manager reads “without damaging performance,” what usually comes to mind is the performance of the company as a whole: revenue, project delivery, speed of innovation, customer experience, and cross-functional collaboration.
What the experiment mainly measured, however, was individual performance grades, individual promotion outcomes, and lines of code submitted by a subset of engineers. These are certainly related to job performance, but they are not organizational performance itself.
Organizational losses often occur in the gaps between individual reviews: more time spent coordinating, more friction across teams, harder development of newcomers. None of these may be enough to move any one person’s semiannual grade, yet they genuinely affect a team’s time, quality, and cost. Organizational performance has never been the simple sum of individual performance.
Furthermore, the experiment randomized by individual, not by team. Treatment and control employees were spread across the same teams. When a treatment-group employee worked from home, control-group colleagues may also have had to switch to messaging or video. Team-level communication costs could therefore land on members of both groups at once and be partially cancelled out in a comparison of individual scores.
What this design identifies well is whether, within a team where some people already work hybrid, receiving the option changes an individual’s own grade. What it cannot identify as clearly is what happens to an organization’s net performance when an entire team or company moves from office-based work to hybrid.
The yardstick may not be sensitive enough
The main performance metric is Trip.com’s semiannual review on a five-grade scale of D, C, B, B+, and A. Before the experiment, these employees’ prior review average was 3.81 out of 5.
This does not prove that managers were going easy, and a mean alone cannot establish that the reviews are distorted. The real question is whether the instrument’s resolution is fine enough to capture every productivity change the title implies.
Corporate reviews serve compensation, promotion, talent calibration, and performance conversations first. Even when the process is taken seriously, it can still be a relatively coarse classification scale rather than a high-sensitivity output instrument built for an experiment. An employee in the 4 band may stay in the same band even if real output slips by a few percent; a slightly slower delivery, a little more coordination time, a little less innovation will not necessarily drop a semiannual grade by a full band.
The company’s guidance to employees stated that work targets stayed the same during home days, the review method stayed the same, and employees continued in the existing review. No new output targets or more sensitive instruments were set up for those four to five and a half months. The first review round even covered July through December 2021, while the two cohorts only started in August and September, so the review results blend in performance from before the experiment began.
“Grades did not change” is a real result. “Productivity did not slip a little” requires the additional assumption that the review scale is sensitive enough. The experiment did not test that assumption.
Equivalence testing tightened the statistics but never calibrated the yardstick
To the paper’s credit, it uses TOST equivalence testing. A non-significant result in an ordinary test only says that no difference was detected; an equivalence test requires the researcher to first define a range “small enough to ignore” and then ask whether the data can rule out differences beyond it.
And that is exactly where the problem lies: the range has to be defined by a person.
The authors set the equivalence bound for grades at plus or minus 0.5 points, explained as half a letter grade; the bound for lines of code is 10% of the control mean. But the paper does not say how much real output, customer value, or organizational cost half a grade corresponds to. The bounds explain the size on the scale; they do not establish that the difference is negligible in business terms.
I will not go to the other extreme here and accuse the authors of manufacturing equivalence by loosening the bounds. The conventional 95% confidence intervals for the four review-round differences all fall roughly between −0.11 and +0.15 grade points; under the 90% interval convention that TOST uses, the four results would still sit inside the bounds even if the equivalence margin had been narrowed in advance to ±0.15. So the real issue is not the statistical procedure but measurement validity: what an equivalence test can support is that “the two groups are close enough on the observed metric of performance grades.” It cannot automatically yield “the two groups’ real productivity is equivalent.” That further conclusion requires performance grades to measure real productivity sensitively and validly, and the paper does not verify this.
The third leap: from a local experiment to an unconditional title
Written out concretely and completely, the conclusion the experiment can directly support looks roughly like this:
In 2021, in Trip.com’s Airfare and IT departments in Shanghai, 1,612 employees with a bachelor’s degree or higher who had passed probation were randomly offered the option to work from home on Wednesdays and Fridays; actual use averaged about one day a week. Over the following four to five and a half months of randomized control, the cumulative quit rate was 4.8% in the treatment group and 7.2% in the control group; no adverse change was detected in individual performance grades or in lines of code submitted by a subset of engineers.
That is a valuable, local conclusion.
The title Nature ran, however, is: “Hybrid working from home improves retention without damaging performance.” Company, department, city, year, population, actual dose, observation window, and metric have all vanished.
Three substitutions in particular happened along the way.
First, what was actually randomized was “receiving a hybrid option”; the title says “hybrid working.” The treatment group averaged about one day at home per week, not a steady two. A low-dose, optional hybrid arrangement and a mandatory, fixed hybrid regime cannot be lumped together.
Second, the experiment measured whether people quit within four to five and a half months; the title says “improves retention.” Between “did not quit within the observation window” and “was retained for the long term” lie time and seasonality.
Third, the core of what the paper measured is individual performance grades; the title generalizes it to “performance,” and in circulation it is routinely translated further into “productivity.” Individual review scores, individual performance, team performance, and enterprise productivity are four different levels of concept. Each time the word widens, a layer of measurement boundary disappears.
Trip.com rolled the policy out company-wide as soon as the experiment ended. As a business decision, that is not unreasonable: a company is entitled to bet on limited evidence. But the business arithmetic described in the paper takes a “one third” reduction in quits obtained in two departments, among about 4.6% of employees, over four to five months, and applies it directly to roughly 35,000 employees to estimate millions of dollars in annual savings. That step assumes similar policy elasticity across different roles, different baseline attrition rates, and different implementation costs, and the experiment tested none of those assumptions.
A company’s decision to roll out is a management decision, not an independent verification that the conclusion generalizes.
Why a good experiment still needs a narrower title
I do not think this study did not deserve publication. Large randomized experiments inside real companies are rare, and Trip.com’s willingness to release its data deserves credit. It supplies a useful piece of evidence: for this group of professional employees, a hybrid option used about one day a week on average improved retention in the short term, without a visible deterioration in existing individual performance metrics.
But one of the most important qualities of rigorous research is that the conclusion does not exceed the evidence.
A title like the following would match the experiment better:
In two departments at Trip.com, offering a hybrid work option reduced short-term quits, with no detectable change in individual performance reviews.
It is less striking than the original, but it accurately preserves the treatment, the time scale, the level of measurement, and the scope.
The next time someone cites a workplace study, the first question worth asking is not “did it get into a top journal,” but these three:
What exactly was randomized, and for how long was the control maintained?
Does the outcome measure individual scores, team output, or business results?
Does the relative percentage being circulated still carry managerial meaning against your own company’s baseline, costs, and external environment?
Statistical significance can only answer: “Did the two groups differ during this period?” A business decision also has to answer: “How long will the difference last? Does it survive a change of environment? Is it worth adopting and scaling in my organization?”
The paper answers the first question well. Nature’s title makes it sound as if the rest have been answered too.
The data are not wrong. What needs tightening is the distance from data to conclusion, and from conclusion to title.
Sources and verification note
The experimental design, sample, timeline, quit rates, performance reviews, promotions, lines of code, and statistical tests discussed in this article are drawn primarily from the following original materials. Headcount and percentage-point conversions are the author’s.
[1] Nicholas Bloom, Ruobing Han & James Liang, “Hybrid working from home improves retention without damaging performance,” Nature, 630, 920–925 (2024). DOI: 10.1038/s41586-024-07500-2. Open access: https://pmc.ncbi.nlm.nih.gov/articles/PMC11208135/
[2] Nicholas Bloom, Ruobing Han & James Liang, “How Hybrid Working From Home Works Out,” NBER Working Paper No. 30292, July 2022, revised January 2023. DOI: 10.3386/w30292. Full text: https://www.nber.org/papers/w30292
[3] Bloom, Han & Liang, “Replication Data for: Hybrid working from home improves retention without damaging performance,” Harvard Dataverse, 2024. Data and replication code DOI: 10.7910/DVN/6X4ZZL
[4] “Is Hybrid the Future of Work: Evidence from a Chinese Experiment,” American Economic Association RCT Registry, trial no. 8075. Registration DOI: 10.1257/rct.8075-1.0
Author’s note
This article does not dispute the statistical results the experiment obtained within its randomized assignment and observation window. The discussion of experiment duration, external environment, measurement validity, organizational performance, cross-industry applicability, and managerial value is the author’s independent analysis based on the public materials above, not the original conclusions of the Nature paper’s authors. The author has no financial or other interest in Trip.com Group or in any of the paper’s authors.

