Multiple comparisons
Scanning thousands of segments guarantees some look significant by chance.
The problem
At the scale WHYSE operates — thousands to tens of thousands of segments per question — a handful will show a dramatic-looking swing purely by chance, even with no real effect present.
The correction
Results are corrected for multiple comparisons and filtered by effect size, so a tiny segment with a large but noisy swing doesn't out-rank a smaller, more reliable one.