×

    Two versions of a practice-area page have been live for a week. Version B is ahead. The testing tool paints the result green, and the team wants to ship it.

    Before calling a winner, ask a less exciting question: was this experiment capable of resolving the business decision?

    An A/B test can estimate the effect of showing eligible people one acceptable experience or another. It cannot compensate for a tiny eligible audience, an outcome that rarely occurs, inconsistent assignment, broken tracking, or a rule that stops whenever the preferred version moves ahead.

    The issue is conditional. This article presents no evidence that “most” law firm tests fail; that unsupported prevalence claim has been removed from the working title. The useful question is whether this test can support this conclusion with its audience, outcome, design, and time.

    Define the decision before the test

    “Which headline wins?” is usually incomplete. Write the decision the firm will make:

    If a clearer consultation explanation produces a practically meaningful increase in accepted inquiries from independently assigned visitors to the Colorado employment-service page, without lowering the service-fit share or increasing intake beyond capacity, roll it out on that journey.

    That statement identifies:

    • the eligible audience and page;
    • the two acceptable experiences;
    • the primary outcome;
    • downstream guardrails;
    • the practical effect the firm cares about; and
    • the action a supported result would trigger.

    If the team would not act on a difference, it should not spend traffic estimating it. If one experience is inaccurate, broken, inaccessible, or misleading, repair it; do not expose a control group to a known defect.

    Count the audience that can actually enter

    A firm may report 8,000 monthly website visits while the tested service page receives 500 eligible visitors. The other 7,500 visits do not make the comparison more informative.

    Define who can enter, when assignment occurs, how long assignment persists, and what unit is analyzed. If assignment is by person, returning sessions from one person should not masquerade as independent participants. If the outcome is counted per person, the denominator should also be people assigned to that experience.

    Do not combine unrelated informational and hiring-intent pages to inflate sample size. The resulting average may answer no service decision the firm can implement.

    Then account for real losses: consent choices, bot filtering, cross-device duplication, unavailable outcomes, and the time between contact and qualification. A sample plan based on raw site sessions can be badly optimistic.

    Choose the smallest effect worth detecting

    Suppose the current accepted-inquiry rate is 5%. A move to 6% is a one-percentage-point increase and a 20% relative increase. Both descriptions are correct; the first makes the absolute business change easier to see.

    The minimum detectable effect is not a forecast. It is the smallest difference the planned test is designed to detect with its stated power and assumptions. Choosing a larger effect reduces the required sample, but it does not make a large effect more likely.

    Set the threshold from the decision. If implementation costs $12,000 and the firm would not change its process for less than a two-point absolute improvement, plan around that fact. Do not choose a dramatic target solely to make the calendar shorter.

    Translate the rate into counts. At 500 eligible people a month, a 5% rate means about 25 accepted inquiries in expectation. A 6% rate means about 30. The proposed one-point change is roughly five additional accepted inquiries per month before qualification, consultation, or retained-matter outcomes are known.

    Calculate the sample before promising an end date

    Here is a reproducible planning illustration. Assume:

    • independent binary outcomes per eligible visitor;
    • equal assignment to control and variant;
    • baseline rate (p_0 = 0.05);
    • two-sided significance level (\alpha = 0.05);
    • power of 80%; and
    • a fixed-horizon comparison under stable conditions.

    Let (p_1) be the alternative rate and (\bar{p}=(p_0+p_1)/2). The approximate sample per arm is the ceiling of:

    [
    \left[\frac{1.96\sqrt{2\bar{p}(1-\bar{p})}+0.842\sqrt{p_0(1-p_0)+p_1(1-p_1)}}{p_1-p_0}\right]^2
    ]

    The formula follows the dominant-tail normal approximation described in statsmodels’ two-proportion sample-planning documentation. It omits the far rejection tail, uses no continuity correction, and is a planning illustration. Other methods and testing products can yield different requirements.

    The results are:

    Scroll sideways to review every column.Each row is shown as a labeled card.

    Change designed to be detected Approximate visitors per version Total eligible visitors Months at 500 eligible visitors/month
    5% to 6% 8,158 16,316 32.632
    5% to 7.5% 1,471 2,942 5.884
    5% to 10% 435 870 1.740

    The first row comes from an unrounded per-arm result of approximately 8,157.731, rounded up to 8,158. At 500 eligible visitors a month, 16,316 ÷ 500 = 32.632 months. Even at 2,000 eligible visitors a month, it takes 16,316 ÷ 2,000 = 8.158 months to collect the planned sample.

    Those calendar estimates exclude ramp-up, outcome maturation, data loss, and any minimum coverage for weekly or seasonal variation. A 32.6-month result is a feasibility warning. It is not a recommendation to freeze a page and business environment for nearly three years.

    Interpret significance and power without sales language

    A 5% significance level describes the planned procedure’s false-positive behavior under the null and its assumptions. It does not mean that a reported winner has a 95% probability of being better.

    Eighty percent power means that, if the specified alternative really holds and assumptions are met, the planned procedure has about an 80% chance of rejecting the null. It does not mean the website change has an 80% chance of succeeding.

    These settings do not rescue a poorly chosen metric. A test can detect more form opens while suitable inquiries remain unchanged. More scrolling may mean interest or confusion. An intermediate action is useful when it answers the stated product question; it is not a proxy for retained clients simply because it occurs more often.

    Optimizely’s low-traffic testing guidance discusses testing larger changes and using earlier actions as ways to work with limited traffic. For a law firm, the boundary is essential: an easier-to-measure interaction does not establish an improvement in qualified or retained matters.

    Know how peeking changes the procedure

    In a conventional fixed-horizon design, repeatedly checking an ordinary significance threshold and stopping the first time it favors a version can increase false-positive risk. A dashboard refreshing every hour is not automatically a valid sequential design.

    Write the analysis and stopping rule before launch. If the selected platform uses a sequential or Bayesian method, understand the method and follow its allowed decision procedure rather than layering an improvised rule on top.

    Monitor technical failures and harm throughout. A broken form, bad routing, inaccessible interaction, or severe guardrail change should not remain live in the name of statistical purity. Microsoft’s trustworthy-experimentation guidance distinguishes data-quality and guardrail monitoring from inferential decisions and discusses methods that account for repeated checks.

    Changing the primary metric, baseline, practical threshold, or audience after seeing results also weakens the conclusion. Label exploratory findings as exploratory and plan a new confirmation if the decision requires it.

    Audit the experiment before reading the winner

    An experiment can collect enough people and still be invalid. Check:

    Scroll sideways to review every column.Each row is shown as a labeled card.

    Failure Question to ask
    Assignment Did eligible people have a stable, random chance of receiving either version?
    Sample ratio Did each group receive the allocation the design expected?
    Implementation Did both versions load and function across supported devices?
    Measurement Did the same accepted outcome fire once in both versions?
    Interference Could one person see both versions or influence another participant’s outcome?
    Concurrent change Did traffic sources, fees, eligibility, or intake handling change unequally?
    Maturity Did both groups receive the same time for qualification and engagement outcomes?

    A mobile defect isolated to one version tests the defect and the intended message together. A tracking event that fires before form acceptance measures button behavior. A mid-test intake-script change may alter downstream outcomes. Document these conditions before presenting a rate difference.

    Work through a law firm’s go/no-go decision

    Consider a fictional two-office PI firm. The website receives 6,000 visits a month, and marketing proposes testing two consultation explanations. The relevant campaign page receives only 500 eligible independent visitors monthly. Its reconciled accepted-inquiry baseline is 5%. Intake can handle 35 qualified inquiries weekly, so capacity is not currently binding.

    Diagram showing eligible traffic, baseline rate, minimum useful effect, sample per variant, runtime, power/significance plan, and a go/no-go alternative.
    Planning illustration under the article’s stated assumptions: at 500 eligible visitors a month, detecting a one-point absolute change can require about 32.6 months; this is a feasibility warning, not a recommendation to run that long.

    The $7,500 implementation would be justified only if the firm could detect at least a one-percentage-point absolute increase in accepted inquiries without lowering qualification share. Marketing initially schedules a six-week test because the full website receives substantial traffic.

    The correct calculation uses the page’s 500 eligible visitors, not 6,000 site visits. Detecting 5% versus 6% under the assumptions above requires approximately 16,316 visitors total, or 32.632 months before maturation and losses. Conditions would be unlikely to remain stable enough for that comparison to govern the same business decision.

    The firm declines the proposed A/B test. It does not conclude that the new explanation is ineffective. It concludes that this experiment is not a practical way to resolve a one-point decision.

    The team then checks what it can learn:

    • Twelve task participants receive fictional injury scenarios; seven cannot tell whether the first call is attorney advice or intake screening.
    • The current page says “speak with our legal team,” while the intake script says “non-attorney intake specialist.”
    • Controlled form and call tests work under the specified conditions.
    • Qualification coding is complete for only 31 of 47 recent accepted inquiries, so the proposed guardrail is not yet dependable.

    The firm first aligns the page with the real intake experience and repairs qualification coding. It rechecks understanding with new task participants and observes mature cohorts. Because the copy correction is required for accuracy and the comparison is not randomized, it makes no causal lift claim.

    A different decision might be testable. If the firm were deciding between two accurate, substantially different paths and would act only on a change from 5% to 10%, the illustrated plan calls for about 870 eligible visitors total, or 1.74 months at 500 per month, plus maturation and calendar coverage. The team would still need valid assignment, stable implementation, a prewritten rule, and a business reason to care about such a large difference.

    The worked example shows why total traffic, a green dashboard, and an arbitrary six-week calendar are not feasibility evidence.

    Keep improving when randomization is impractical

    Low traffic does not make the firm powerless. Use methods matched to the uncertainty:

    • repair reproduced failures and inaccurate statements;
    • reconcile accepted inquiries and downstream stages;
    • observe people performing realistic tasks with fictional scenarios;
    • code the misunderstandings and mismatches intake hears;
    • make coherent, evidence-backed changes; and
    • read later cohorts with explicit limits on causal interpretation.

    Juris Digital’s historical website focus-group account illustrates how task observation can expose functional and information problems. Its historical performance claim is not needed here. Qualitative work identifies a mechanism; it does not estimate a population lift.

    If the measurement audit shows that the proposed change requires new templates or service architecture, the appropriate buying question may be a rebuild rather than an experiment. Juris Digital’s current website design service provides public detail about direct builds, migration protection, a four-phase delivery path, and named responsibilities. It does not sell the sample-size calculation as an A/B-testing offer or promise a lift.

    Bring the eligible-audience count, baseline definition, sample calculation, measurement audit, and practical decision threshold so a proposal can separate the build requirement from the claim that still cannot be tested. Recognizing that a test cannot answer the question is evidence arriving before the firm spends months pretending to learn.

    Casey Meraz Casey Meraz is an entrepreneur, SEO expert, investor, creator, husband, father, friend, and CEO of Juris Digital. Casey is a frequent speaker at industry events and the author of two books on digital marketing, including "Local Marketing for Personal Injury Lawyers" and “How to Perform the Ultimate Local SEO Audit”
    X - Close
    👋 Questions? Fire away...
    X - Close