効果測定

AI Store Agent ROI Calculator for Shopify

Download an editable Shopify AI Store Agent ROI calculator and learn how to estimate incremental gross profit without treating assisted revenue as causal lift.

A balance scale weighing a parcel and service bell against unmarked value tokens
Illustration: weigh verified value against the full program cost instead of mistaking activity for impact.

The core rule: Calculate what changed because the agent was available—not how much revenue happened after a conversation. Use a treatment-control comparison where the store has enough data, add only support costs that were genuinely avoided, and subtract the full program cost.

A shopper sees help choosing a product or resolving a question. A merchant also needs evidence that the operating cost produces incremental value.

This guide is for ecommerce leaders building the business case, finance and analytics reviewers validating the measurement, and support operations teams estimating cash savings.

The BuyScout® AI Store Agent is designed to handle sales and support conversations using a merchant's store information. But the existence of a conversation, an attributed order, or a resolved question does not prove incremental value. A defensible ROI model separates activity from impact.

Use the spreadsheet to replace every input with data from your own store. It accepts separate treatment and control conversion rates, return-adjusted order values, and margins. It then calculates each group’s gross profit per eligible session, the incremental difference, verified cash savings, program cost, and ROI while surfacing checks designed to expose double counting.

Download the editable AI Store Agent ROI calculator

Quick start: Make a copy, choose the strongest measurement tier your store can support, replace the yellow cells with same-period store records, and review every audit check before using the result in a budget decision.

The spreadsheet and this measurement plan are decision tools, not a promise of results.

In this guide

Start with verifiable value streams

Start with value streams that can be observed and independently checked. Most Shopify merchants will consider four:

Value streamWhat shoppers doWhat the merchant measuresHow it belongs in ROI
ConversionAsk questions, compare options, then buy or leaveDifference in purchase rate between treatment and controlCount incremental orders only
Order valueAccept a relevant add-on, upgrade, or bundleDifference in realized order value between treatment and controlCapture through group-level profit per eligible session when order economics change
Cart recoveryReturn and complete a purchase after helpIncremental completed orders versus controlKeep inside conversion lift when already captured there
Support capacityGet an answer without staff involvementVerified paid hours or outsourced units avoidedCount cash savings; record redeployed capacity separately

Engagement is useful, but it is not a value stream by itself. Conversations, recommendation clicks, answer ratings, and handoffs help diagnose the experience. They should not be converted directly into dollars.

Shopify's analytics field reference defines metrics such as average order value, gross profit, gross margin, product conversion rate, and add-to-cart rate. Use those definitions consistently. The calculator's separate control and treatment return-adjusted order values per included order are merchant-defined planning inputs, not Shopify AOV; label their deductions and use the same basis for both groups.

The causal ROI formula

The simplest defensible sales model uses the difference between a treatment group, where the AI store agent is available, and a holdout group, where it is not. The simplified version below assumes a common return-adjusted order value and margin. The downloadable workbook implements the more general gross-profit-per-eligible-session method shown immediately afterward, so control and treatment economics can differ.

incremental orders at modeled rollout
  = modeled eligible sessions
  × (treatment conversion rate - control conversion rate)

incremental return-adjusted sales
  = incremental orders
  × common return-adjusted order value per included order

incremental gross profit
  = incremental return-adjusted sales
  × gross margin rate

verified support cost savings
  = paid support hours no longer scheduled or purchased
  × the applicable loaded or vendor hourly cost

total verified benefit
  = incremental gross profit
  + verified support cost savings
  + other independently verified cash savings

net benefit
  = total verified benefit
  - total AI store agent program cost

ROI
  = net benefit
  ÷ total AI store agent program cost

For an experiment-period estimate, use treatment sessions in the first line. For a full-rollout projection, use all sessions expected to be eligible and label the extrapolation. Never multiply the rate difference by the combined test sample and call that the observed treatment effect.

Estimate return-adjusted order value after discounts and expected returns if possible, and do not relabel that custom input as Shopify AOV. Shopify defines gross profit as net sales minus cost of goods sold. If fulfillment, payment, or other order-level costs matter, use contribution margin instead and document those deductions.

If treatment changes order value, returns, or product mix, avoid separate conversion and order-value add-ons. Prefer one group-level outcome:

incremental gross profit at modeled rollout
  = modeled eligible sessions
  × (treatment gross profit per eligible session
     - control gross profit per eligible session)

For a fast directional model, gross margin is a reasonable starting point. For a finance-grade model, replace it with contribution margin and document exactly which costs are included.

Operator note: Keep the raw inputs beside the result. A reviewer should be able to change any assumption and see the answer update immediately.

Assisted revenue is not incremental revenue

An assisted order happened after a shopper interacted with the agent. An incremental order happened because the agent was available. Those groups overlap, but they are not identical.

A shopper who asks a shipping question and buys may have purchased anyway; another may return later on a different device. Attribution describes a path, while a control comparison estimates causality.

Shopify supports several marketing attribution models and notes that any-click attribution can allocate more credit than orders received. Attributed sales aid investigation but are unsafe as the sole ROI numerator.

MetricQuestion it answersGood useUnsafe use
ConversationsDid shoppers use the agent?Adoption and capacity planningMultiplying every chat by AOV
Assisted ordersDid an order occur after interaction?Journey analysis and segmentationCalling all assisted revenue incremental
Attributed revenueWhich touchpoint received credit?Comparing attribution viewsTreating credit allocation as causal lift
Treatment-control liftWhat changed when the agent was available?Incremental sales estimateIgnoring imbalance or missing data
Gross profitHow much value remained after product cost?ROI numeratorReporting revenue and gross profit as two benefits

Keep assisted revenue on the dashboard. Just label it correctly: descriptive, not causal.

Run a holdout before assigning causal credit

The strongest practical design randomly assigns eligible shoppers to one of two experiences:

  • Treatment: the AI store agent is available.
  • Control: the AI store agent is not available.

Assign the experience before the shopper decides whether to engage. Comparing people who chatted with people who did not is self-selection: people with difficult questions, high intent, or complex carts may be more likely to open chat.

Keep assignment persistent and price, promotions, traffic mix, theme, and inventory as stable as operations allow. Record unavoidable changes.

Use an intent-to-treat view: measure everyone assigned, not only people who opened the agent.

Before launch, write down:

  • The eligible audience and primary outcome, usually purchase rate or gross profit per eligible session.
  • Guardrail metrics such as refund rate, cancellation rate, handoff rate, response latency, and shopper complaints.
  • The assignment method, split, planned duration, and decision rule.
  • Events that can invalidate the test, such as a major promotion, stockout, or checkout outage.

Do not stop at the first positive swing. If a randomized holdout is infeasible, use a phased rollout or matched period and label the result as less causal.

Measurement tiers for stores with limited data

Choose the measurement tier from the number of eligible sessions and purchases you expect—not from labels such as “small” or “large.” A store with few sessions but frequent repeat purchases may have more usable information than a higher-traffic store with a very low purchase rate.

Measurement tierWhen it fitsWhat to measureWhat you may conclude
1. Operational pilotToo few purchases to estimate a stable conversion differenceAnswer accuracy, unresolved-question rate, handoffs, shopper complaints, latency, and paid support units actually avoidedWhether the experience is safe and useful enough for a larger test; not incremental sales ROI
2. Phased comparisonRandom assignment is unavailable, but comparable periods, products, or markets existGross profit per eligible session plus the same guardrails before and after the rolloutA directional association, with promotion, inventory, seasonality, and traffic changes disclosed
3. Randomized holdoutThe store expects enough outcome events to evaluate a predeclared effectIntent-to-treat conversion or gross profit per eligible session, with persistent assignment and guardrailsA causal estimate within the tested population, subject to data quality and statistical uncertainty

Before choosing a sample size, ask an analyst to use the store's baseline purchase rate, minimum decision-worthy effect, desired confidence, and planned allocation. A generic traffic threshold cannot answer that question. When Tier 1 is the only responsible option, record operational value and keep incremental revenue at zero in the calculator.

Calculator inputs and source records

Use merchant data wherever possible. Industry averages are poor substitutes for the economics of a particular catalog.

InputPreferred sourceDefinition to lock
Modeled eligible sessionsAssignment log plus rollout forecastTreatment sessions for an experiment estimate, or all expected rollout sessions for a labeled projection
Control conversion rateHoldout orders ÷ holdout sessionsSame eligibility and order rules as treatment
Treatment conversion rateTreatment orders ÷ treatment sessionsInclude everyone assigned, not only chatters
Control return-adjusted order valueShopify orders plus finance assumptionsMerchant-defined value per included control order after the stated treatment of discounts, refunds, cancellations, tax, and shipping; do not label it Shopify AOV
Treatment return-adjusted order valueShopify orders plus finance assumptionsThe same merchant-defined basis applied to each included treatment order
Control gross or contribution marginFinance or product cost recordsCosts included for the control population and time period
Treatment gross or contribution marginFinance or product cost recordsThe same cost definition applied to the observed or projected treatment mix
Paid support hours avoidedStaffing, vendor, or ticket recordsHours no longer scheduled or purchased, not theoretical handle time
Loaded hourly costPayroll/vendor dataWage plus applicable employer and tooling costs
Other verified cash savingsInvoice, staffing, or operating recordIndependently evidenced cash cost no longer incurred and not counted in another row
Program costInvoice plus internal timeSoftware, setup, monitoring, content, and integration
GuardrailsReturns, complaints, latency, QAThresholds that can override a positive ROI result

Shopify's order conversion summary helps investigate visits and referrers, but can be limited by blocked cookies, headless storefronts, or other order paths. It does not replace experiment assignment.

If the agent is being evaluated as a build, buy, or hybrid project, use the Build vs. Buy guide and TCO worksheet to establish the program-cost input before calculating ROI.

Use the editable ROI calculator

Download the ROI calculator, make a copy, and complete the yellow input cells from Shopify, finance, assignment, staffing, and vendor records. Its starter values are explicitly illustrative—not industry benchmarks, BuyScout® AI Store Agent outcomes, or promises—and must be replaced before making a decision.

  1. Set the period and eligible population. Use treatment sessions for an experiment-period estimate. Use all expected eligible sessions only for a clearly labeled rollout projection.
  2. Enter group economics. Use consistent, merchant-defined return-adjusted order values for control and treatment, then enter each group’s gross or contribution margin. The workbook calculates control, treatment, and incremental gross profit per eligible session directly.
  3. Enter cash support savings. Include only paid hours or outsourced units that are no longer scheduled or purchased.
  4. Add other cash savings only when independently verified. Do not use this row for theoretical capacity, assisted revenue, or value already captured elsewhere.
  5. Enter the full program cost. Include software, setup, integration, monitoring, content maintenance, and internal operating time for the same period.
  6. Review the audit checks. Confirm that assisted revenue, recovered carts, and redeployed salaried capacity have not been added again.

The workbook separates software and vendor subscriptions, internal setup and monitoring, and other program costs. Use Other program costs for integration, contractors, external services, content maintenance, or another period cost not captured in the first two rows.

The workbook calculates control and treatment gross profit per eligible session, incremental gross profit, verified support cash savings, total benefit, net benefit, and ROI. A blank input is not evidence of zero cost or zero effect; resolve material blanks before using the result for a budget decision.

When a randomized experiment is unavailable, the same spreadsheet can record a directional estimate. Label the method and limitations beside the result rather than presenting it as causal ROI.

Hypothetical worked example

Illustrative arithmetic only: These inputs are not a benchmark, a BuyScout® AI Store Agent result, or a forecast for another store.

Suppose a merchant models 10,000 eligible sessions. The control converts at 2.0% and treatment at 2.2%. Both groups use an $80 return-adjusted order value and a 50% gross margin. Control gross profit per eligible session is $0.80; treatment is $0.88. The $0.08 difference across 10,000 sessions produces $800 in incremental gross profit.

If staffing records also verify $300 in paid support cost avoided, total verified benefit is $1,100. If the same period includes $600 in software and vendor cost, $200 in internal setup and monitoring, and $0 in other program costs, total program cost is $800. Net benefit is therefore $300 and ROI is 37.5% ($300 ÷ $800). Change any assumption or evidence status and the conclusion should update with it.

Six rules that prevent double counting

Use these rules to prevent double counting:

  1. Count an order once. If recovered carts are included in the treatment-control conversion difference, do not add recovered-cart revenue again.
  2. Choose revenue or profit. Incremental revenue is an intermediate step. The value used in ROI is the resulting gross or contribution profit.
  3. Separate conversion from order value. If the test directly estimates gross profit per eligible session, do not also add independent conversion and order-value benefits.
  4. Do not monetize every automated answer. Loaded labor cost is valid only for hours no longer paid or purchased. Treat reassigned salaried time as capacity; monetize its output only when measured independently and not captured elsewhere.
  5. Subtract discounts and returns. A sale generated through a costly incentive or followed by a return is not worth its checkout value.
  6. Keep attribution outside the causal total. Assisted and attributed revenue can explain the mechanism, but they are not added to experimental lift.

A useful audit question is: “Could the same order, hour, or dollar appear in another row?” If yes, consolidate it.

Handle tracking gaps explicitly

Shopify analytics and Google Analytics can differ because they use different definitions and are affected by JavaScript, cookies, blockers, time zones, privacy settings, and connectivity. Review Shopify's discrepancy guidance, choose one assignment system as the experiment source of truth, and reconcile completed orders against Shopify's transactional records.

Record the share of eligible traffic that could not be measured and describe the population the result represents. Do not bypass consent or send raw conversation text, email addresses, phone numbers, or other recognizable personal information to analytics tools; Google Analytics prohibits collecting PII.

Failure modes that can invalidate the result

After the double-counting audit, check for design and operating failures that can still make the result unreliable:

  • Using different eligibility, order, or margin definitions for treatment and control.
  • Letting promotions, stockouts, traffic mix, theme changes, or outages affect one group without recording the imbalance.
  • Stopping at the first positive swing or selecting only favorable dates or segments.
  • Applying one category's margin to a materially different treatment product mix.
  • Leaving missing assignment coverage or analytics-definition conflicts unquantified.
  • Ignoring slower pages, poor answers, complaints, unsafe claims, or excessive handoffs because the financial result is positive.

ROI is not the only decision criterion. A treatment that raises short-term gross profit while creating unsafe claims or a worse support experience should not ship unchanged.

Monthly review checklist

Use this operating checklist after the initial experiment:

  • Recalculate conversion and gross profit using consistent definitions.
  • Compare treatment and control or retain a small ongoing holdout where appropriate.
  • Reconcile discounts, refunds, cancellations, and product costs.
  • Validate support hours with staffing or vendor records.
  • Review answer accuracy, handoffs, latency, and complaints.
  • Check whether catalog, inventory, or policy changes affected results.
  • Record consent coverage and tracking gaps.
  • Update program cost with internal operating time.
  • Update the calculator's assumptions, evidence links, and analysis period.

Operator note: A monthly review should produce decisions, not just charts. Improve weak knowledge, change a guardrail, expand a winning use case, or pause a risky one.

Calculate before you claim ROI

Start with a narrow use case, preserve a control when the data supports one, and let gross profit and shopper outcomes—not assisted-revenue headlines—decide what to expand.

Download the editable ROI calculator

To test the underlying experience with your own catalog, explore the BuyScout® AI Store Agent or review current pricing. The BuyScout® AI Store Agent does not guarantee a particular ROI; outcomes depend on the store, catalog, traffic, margins, implementation, and measurement design.

その他の記事

すべての記事