# Creative test planner and experiment log

The Remarkable • Version 1.0 • September 29, 2026

Use one copy per test. Replace every bracketed field before launch. Copy the variant record for each control or challenger. Open this editable Markdown file in a text editor, or paste it into your team's document tool.

Companion guide: https://theremarkableagency.com/blog/creative-testing-framework-paid-social/
Completed hypothetical example: https://theremarkableagency.com/downloads/creative-test-worked-example.md

## 1. Write the question

- Test ID: [example naming pattern: product_market_channel_concept_hook_round]
- Owner and reviewer: [names]
- Product, audience, market, and channel: [details]
- Decision this test must inform: [what you will do differently]
- Hypothesis: Changing [one creative element] from [control] to [challenger] will improve [primary metric] because [reason].
- Evidence behind that belief: [customer research, prior test ID, or observation]
- Asset links: [control] / [challenger]
- What changes: [one element]
- What stays fixed: [audience, offer, body, format, landing page, attribution, placement, and bidding settings]

## 2. Define measurement before spending

- Primary outcome: [precise event; for example, first qualified activation]
- Event definition and deduplication: [what counts, what does not, unique event or customer ID]
- Primary metric and formula: [for example, media spend / qualified activations]
- Secondary diagnostic metric: [for example, outbound clicks / impressions]
- Business guardrail: [for example, paid-customer acquisition cost or activation-to-paid conversion]
- Data source for each metric: [ad platform, analytics, CRM, billing]
- Attribution window and model: [settings]
- Conversion lag to allow before final review: [duration and reason]
- Baseline period and value: [dates, currency, metric]
- Minimum worthwhile improvement: [absolute or relative change, selected before launch]
- Test method: [platform experiment / randomized split / observational comparison]
- Evidence requirement: [sample-size plan, method, confidence threshold, and reviewer]

A delivery comparison is not automatically a randomized experiment. Mark observational results as directional. A planning budget or fixed impression count does not establish statistical significance.

## 3. Check whether the budget can answer the question

- Number of cells, including the control: [n]
- Planning CPA for the chosen outcome: [amount and currency]
- Expected outcomes per cell for budget planning: [n]
- Media budget per cell = planning CPA × expected outcomes: [amount]
- Total media budget = budget per cell × cells: [amount]
- Production and review cost, kept separate: [amount]
- Planned delivery window: [start / end / account time zone]
- Planned average daily media spend = total media budget / delivery days: [amount]
- Final read date after conversion lag: [date]
- Maximum authorized spend: [amount]

Use the same outcome in the CPA and outcome-count fields. These equations estimate affordability; they do not forecast a winner or calculate sample size. If the budget is too small, reduce the number of cells or reconsider the question before launch.

## 4. Set decision and stop rules

- Adopt a challenger only if: [primary outcome, evidence requirement, and business guardrail all pass]
- Continue collecting data only if: [pre-agreed extension conditions and spend cap]
- Record an inconclusive result if: [insufficient evidence, uneven delivery, or conflicting signals]
- Stop immediately for: [tracking failure, incorrect claim, broken landing page, or agreed loss limit]
- Stop-rule approver: [name]
- If the budget ends without a clear result: [retain control / revise question / plan another test]

Do not change the winning metric after seeing the results. Do not extend a test indefinitely until a preferred result appears. Log any material change and separate results collected under different conditions.

## 5. Copy this record for each variant

### Variant [ID] — [control / challenger]

- Parent creative ID and asset link: [details]
- Concept / hook / format tags: [details]
- Exact change from control: [details]
- Production cost and human review time: [amount / minutes]
- Brand and claims approval: [reviewer / date / evidence link]
- Platform campaign / experiment / ad IDs: [IDs]
- Delivery dates and time zone: [details]
- Media spend and currency: [amount]
- Impressions: [n]
- Outbound clicks: [n]
- Primary outcome count: [n]
- Paid customers or other business guardrail: [n / value]
- CTR = outbound clicks / impressions × 100: [%]
- Primary CPA = media spend / primary outcomes: [amount]
- Paid-customer media CAC = media spend / attributed new paid customers: [amount]
- Test method's reported interval or confidence result: [value / unavailable]
- Tracking gaps, delivery imbalance, or other confounders: [details]
- Source export and extraction date: [link / date]

Use “not available” when a denominator is zero or a metric is missing. Media CAC excludes production and agency fees; add those separately if calculating fully loaded CAC. Do not mix platform-attributed and CRM-attributed counts without documenting the difference.

## 6. Close the loop

- Decision: [adopt challenger / retain control / iterate / inconclusive / invalid test]
- Evidence supporting the decision: [primary result and limits]
- What the test does not establish: [limits]
- What stays in the next brief: [concept or execution element]
- Next single change: [hypothesis]
- Next test ID, owner, and review date: [details]

### Change log

- [Date / person / change / reason / effect on interpretation]

### Before launch

- [ ] Hypothesis, primary outcome, and guardrail are written down.
- [ ] Conversion events and the landing page work.
- [ ] Claims, assets, usage rights, and reviewer are recorded.
- [ ] Spend cap, read date, and decision rules are agreed.
- [ ] Control and challenger differ only as intended.

### Before calling a winner

- [ ] The planned window and conversion lag have elapsed.
- [ ] Delivery, tracking, and attribution are comparable.
- [ ] The primary metric and evidence requirement pass.
- [ ] The downstream business guardrail passes.
- [ ] The decision and next hypothesis are recorded.
