← Library The Experiment That Wanted to Be Right Sign in

The Experiment That Wanted to Be Right

The greenhouse had one healthy plant and twelve suspicious ones. Leah was assigned to test a new soil treatment before the palace gardens adopted it. The gardener said the treatment had cured a wilted mint. The cook said the treated plants looked greener. A traveling merchant said the same thing about his own field, though he could not name a comparison.

Leah loved the idea. A small experiment could settle the matter. She divided the plants into two groups, gave one group the soil treatment, and left the other with the ordinary soil. Then she watered only the treated plants and placed them beside the sunny window. The treated plants perked up. The untreated ones drooped. Leah recorded “treatment works” before the gardener could suggest that thirst had been part of the result.

The palace gardener, Mara, stopped her. “You gave the treatment and the extra water together. What happened without the treatment is unclear.” “I can try again,” Leah said. “For a fair comparison, try a plant that could have responded without the treatment. A control is what would happen anyway.”

Leah added untreated plants with the same watering and light. She was surprised that some perked up too. The result was smaller than the first story. She moved the treated plants to a dimmer shelf, where their leaves curled. Then she moved the controls to the sunny window, where they recovered. Every plant wanted the experiment to be right because Leah did.

She began selecting the plants that had grown best. The brightest leaves made a persuasive display. The pale, damaged ones went into a drawer. When Mara asked for the original records, Leah showed only the successful pots. The result was not a mistake in arithmetic. It was a mistake in attention. Seeking and weighting confirming evidence made the desired answer easier to see. Ignoring disconfirming evidence made the other answer disappear.

The cook suggested adding a little fertilizer to the treated group. Leah did, but only there. A brighter lamp appeared above the controls. She wrote a new page, then forgot to number the old one. By dusk, the greenhouse held not one test but a collection of accidental advantages. Mara called for a reset. “We need a design before we touch the soil.”

They planned the fair experiment together. First came the question: does the treatment affect plant growth under the conditions we can measure? Then the plants. Leah selected equal seedlings with similar size and health. To reduce the influence of known differences, she assigned them by a random method rather than placing the largest plants in one group. The process could not balance every unknown feature, but it balanced known and unknown differences on average.

A single extra pot of water or a brighter lamp would not automatically be tied to the treatment. Some enrollment choices were still possible. Leah could select only plants that had already responded in an earlier trial, or recruit only plants from one bed. Random assignment among the chosen plants would not repair an unfair choice about which plants entered the experiment. The frame mattered. A treatment could look effective because the chosen plants were unusually ready to grow.

Leah folded labels so she could not see the treatment. Mara prepared the pots with the same soil, water, and spacing. The labels were shuffled. A separate assistant, who did not know which treatment each plant received, measured height and counted leaves. This was blinding: the person recording the result could not be guided by seeing which pot had been treated.

The first fair run produced a modest difference. The treated plants averaged slightly more growth, but the gap was small enough that seedlings varied widely. Leah wanted to stop because the result did not sparkle. Mara reminded her that a smaller real result is more useful than a dramatic result purchased by design.

A single run is not the end of inquiry. Leah and Mara repeated the experiment with a new batch, a different week, and another person preparing the pots. Some results favored the treatment. Some barely differed. One run favored the controls. The smaller result survived replication better than the original spectacular claim. The team did not pretend that a handful of plants proved a universal law.

It could say that under these conditions and with these measurements, the treatment produced a modest, uncertain improvement. The cook asked whether the controls were unfair because they received no treatment. Mara explained that a control was not a neglected plant. It was a way to estimate what would have happened otherwise. A placebo, when appropriate, and the same handling, soil, light, and watering could keep expectations from shaping the result. The control made the comparison meaningful.

Leah wondered how anyone could believe anything when so many influences existed. Mara pointed to the blind recorder, the random assignment, the repeated runs, and the original notes. Those choices did not make an experiment invulnerable. They reduced avoidable ways for a wish to steer the evidence. The conclusion remained open to revision, but it had a fairer chance to be right.

At the presentation, the gardener praised the mint and the merchant raised his own field story. Leah felt the old pull to use those claims as support. Instead, she showed the first biased trial, the fair design, and the range of outcomes. She included the run that favored the controls. A listener asked whether the treatment should be adopted. Leah said the experiment suggested a small benefit worth testing, not a promise for every garden.

The next experiment began with a question Leah had been avoiding: what if the mint had been unusually vigorous from the start? She planted seeds from the same parent plant in four similar pots, then assigned two to treatment and two to ordinary soil. She did not choose the best-looking seedlings for one group. A clerk drew slips from a hat while Leah watched, and the labels were folded before the pots were watered. The slips balanced the known differences on average, but they could not promise that every hidden trait was balanced in this small group.

Blinding changed the work in an ordinary but important way. The assistant measuring leaf length saw only pot numbers. Leah prepared the treatment, but she did not announce which pot had received it. When the first measurements came back, the treated group was not the obvious winner. One treated plant had been damaged by a snail. The assistant marked the missing leaf rather than quietly replacing it with a neighboring leaf. The damage became part of the record, not a reason to choose a more convenient plant.

Replication gave the team a pattern to ask about. In a second tray, the same treatment produced almost no difference. In a third, the controls grew better. Leah’s first hope was to repeat the favorable run, but Mara asked her to repeat the procedure and publish the result that appeared. The team kept the soil moist, yet the treated pots dried faster because they sat beside a warm wall. That was a new condition, not a failure of the method. It showed why a result depended on the environment and why one small trial could mislead.

The cook still believed the treatment was useful. “The mint can see it,” he said. Leah did not dismiss the observation. She separated it from the experiment: the cook’s observation could suggest a useful question, while the randomized and blinded comparison could estimate whether growth changed under the tested conditions. Confirmation bias was not the presence of a belief. It was the hidden work of seeking, weighting, and ignoring evidence because a favorite answer felt better. The experiment could be fair even when the conclusion was small.

The palace’s first request for a larger trial nearly repeated the enrollment mistake. A clerk offered only the mint plants that the gardener had praised. Leah stopped him. Those were selected after the first biased result, so they could carry the very bias the new experiment was meant to test. The team returned to the seed tray and enrolled plants before assigning the treatment.

Random assignment could balance the known and unknown differences among the plants it was given, but it could not repair the decision to recruit only plants from a famous bed.

Leah also drew a small table of every seedling, including the plants that died. The table looked less impressive than the gardener’s memory. It showed that the treatment group had two damaged leaves, the control group had one, and neither group had been protected from the snail by magic. She marked the damage before the next round and added a feeding schedule that would not depend on what she hoped to see.

A small sample can still be useful when the question and the design are clear, but it can mislead when natural variation is large or when a few unusual plants decide the average.

Mara asked Leah to repeat the treatment test with a second kind of plant. The new result was even smaller, and the team recorded that difference instead of hiding it. The treatment might help under a narrow set of conditions. That was a more humble conclusion than the original one, but it was also more useful. A fair experiment did not force the world to agree.

It changed the question so the answer could survive after the gardener’s excitement and the apprentice’s wish had gone home.

The palace gardener asked what he should do with the mint on his desk. Leah suggested that he try a small, labeled comparison outdoors, with ordinary mint nearby, and keep watering decisions the same for a few weeks. The gardener disliked the plan because it gave him fewer dramatic plants to show visitors. He agreed after seeing that the fair test could still tell him whether to keep the treatment. A credible result did not need to be large.

It needed to survive a person’s wish to be right.

Before leaving, Leah placed the first biased records in a folder marked “lessons learned.” The brightest plants were still there, the dead leaves still counted, and the extra watering still visible. The record was not embarrassing enough to throw away. It was the beginning of a method: seek the evidence that agreed, seek just as carefully for the evidence that did not, and let a fair comparison have a chance to disappoint the experimenter.

The palace bought enough treatment for a larger trial, and Mara asked Leah to write down every decision before the next planting season. The page began with the question and ended with a warning: if the result is surprising, check the design before blaming the world. The warning was not a promise of certainty. It was a way to keep a favorite answer from becoming a hidden experimenter.