IntelliJAMS

A/B testing can feel high-stakes, especially in e-commerce. You launch a test, check it an hour later, and either think you've broken your store or discovered a goldmine. Neither is true yet. This episode is a therapy session for anyone who's felt the anxiety of running A/B tests on their Shopify store.

In this episode of IntelliJAMS, Alex McEachern and experimentation expert Ally Petretti walk through the most common anxieties of A/B testing and how to manage them with better process instead of more stress.

Should you check your A/B test results every day?
Checking your test daily is like weighing yourself multiple times a day on a diet. The number isn't wrong, but it's not ready to be meaningful yet. Check to make sure the test is collecting data and firing correctly, but don't read into the directional results until you've hit your planned runtime and traffic thresholds.

What is statistical significance and when should you end an A/B test?
Statistical significance is often treated as a finish line, but reaching stat sig on day one doesn't mean you have a winner. Stat sig only looks at the math. You also need time (to account for day-of-week behavior changes, promotions, and the novelty effect) and volume (small sample sizes are fragile and can skew results dramatically). A test typically needs to run for at least one to two full weeks to account for these variables.

What is the novelty effect in A/B testing?
The novelty effect happens when returning visitors see something different on your site and react to the change itself rather than the actual experience. In price testing, this can look like a customer treating a different price as a discount and converting out of urgency. The novelty effect fades after a week or two as your sample includes more new visitors and repeat visitors adjust.

What is metric shopping in CRO?
Metric shopping is when you pick a winning metric after a test is already running, instead of sticking to the success metric defined in your hypothesis. For example, if your hypothesis targeted conversion rate but AOV happened to spike, declaring a win based on AOV is metric shopping. It's what Ally calls "the silent killer of CRO." The right move is to note the unexpected metric change, form a new hypothesis around it, and run a separate test.

How should you communicate A/B test results?
Don't just drop a list of metrics. That's left open to interpretation. Data storytelling means framing results around your original hypothesis, explaining what customers did and why, and tailoring the narrative for your audience. A CFO needs a different story than a CMO. Use tools like the Intelligems Slack bot or MCP integration to help frame results for different stakeholders.

How do you handle test ideas that aren't backed by data?
Ideas from teammates or Twitter threads aren't bad, but they need a hypothesis before they become tests. A strong hypothesis includes a data-backed problem, a proposed solution (the test idea), and a success metric. If an idea doesn't have that, either formulate the hypothesis yourself from existing data or add it to the backlog until you can.

How do you coordinate A/B testing across teams?
CRO teams and paid media teams often test independently, which creates noise in each other's data. The key is sharing hypotheses (not just test plans), coordinating testing calendars, and recognizing that every team is working toward the same growth goal. Regular cross-team syncs where you discuss challenges, upcoming campaigns, and testing angles can turn conflict into collaboration.

Join GEM Academy for free courses and a community of brands sharing what works: https://www.skool.com/intelligems-aca...

Connect with Ally Petretti:
https://www.linkedin.com/in/ally-petretti-kuhn-47a27014/ 


Connect with Intelligems:

Website: https://intelligems.io
Blog: https://intelligems.io/blog

----
0:00 - Intro: Why we need a CRO therapist
1:05 - The #1 stressor in testing: urgency and timing
1:44 - "I'm a genius" vs. "I'm killing the business" (early test reactions)
2:33 - Peeking at tests: checking vs. reacting
4:51 - Pre-test analysis: setting guardrails before you launch
5:30 - Statistical significance isn't the finish line
6:13 - The novelty effect and why it matters for price testing
7:47 - Using the time series chart to read test stability
8:39 - When to call a test early (the right way)
10:12 - No test losers: every result is a learning
10:59 - Hypothesis discipline and metric shopping
12:18 - What metric shopping is and why it's the silent killer of CRO
13:25 - Communicating test results: data storytelling
15:57 - The three E's framework: Explore, Experiment, Extend
17:25 - Understanding your customers through testing
18:20 - "Feelings over data": handling test ideas without a hypothesis
21:19 - Cross-team testing: getting paid media and CRO on the same page
25:42 - Wrap-up and where to connect with Ally

What is IntelliJAMS?

We got a real jam going down, welcome to IntelliJAMS.

Hosted by Alex and Adam, each episode serves up quick insights and real-world takeaways, all packed into a fast and fun format.

Whether you’re experimenting with new ideas or want to learn what other brands are already up to, Intellijams has everything you need to stay in the loop. Tune in for short, sweet, and totally actionable insights!