You slice your record by sport, by book, by market, by day of week. One slice is clearly your best. It is probably nothing.
A 95% bar means a coin-flip record clears it one time in twenty per test. Run the test on several slices and the chance that at least one looks good by luck stops being small:
| Slices tested | Chance one looks real by luck | Bar it should clear |
|---|---|---|
| 1 | 5.0% | z = 1.96 |
| 5 | 22.6% | z = 2.58 |
| 10 | 40.1% | z = 2.81 |
| 20 | 64.2% | z = 3.02 |
By 14 slices it is a coin flip that something in your record looks significant when nothing is. A bettor who invents tags until one of them looks profitable is not running a test. They are running a search, and a search needs a higher bar than z = 1.96 — up to z = 3.02 at 20 slices.
What to do instead
Decide which slices matter before looking, and count every slice you tested, including the ones you abandoned. A slice under 30 settled bets should not be characterised at all — not "slightly negative", not "promising", nothing.
This is the correction essentially nobody in this category applies, and tag breakdowns are exactly where it is needed. A tool that shows you twenty segments and highlights the green ones is selling you the search results and calling them a finding.