GutCheck: AI Ad Creative Testing
AI incremental ad testing across Meta and Google—find winning creatives faster, measure true lift, and stop killing ads on gut instinct.
Performance marketers live in a permanent state of creative anxiety. You launch five ad variants on Meta, watch CPMs spike on day three, pause the “loser” based on a gut feeling, and never learn whether the winner actually drove incremental conversions—or just stole credit from another campaign. Platform dashboards show clicks and ROAS, but they cannot tell you which headline, hook, or visual element caused the lift. Traditional A/B tests take weeks, burn budget on statistically invalid sample sizes, and require a data analyst to interpret.
Built for Solo Founders, Side Hustlers, Weekend Builders.
Global digital ad spend is projected to exceed $700 billion by 2025, with AI-driven optimization cited as a primary growth driver (Ideabrowser opportunity analysis on idea #38). Budget exists; waste is the pain merchants pay to eliminate.
Suggested stack: Next.js 14 + Vercel, Supabase (Postgres + Auth), Meta Marketing API + Google Ads API, OpenAI GPT-4o (multimodal), Inngest or BullMQ + Redis, Stripe Billing. Weekend scope: about 10 hours.
The Problem
Performance marketers live in a permanent state of creative anxiety. You launch five ad variants on Meta, watch CPMs spike on day three, pause the “loser” based on a gut feeling…
The Solution
GutCheck is an AI-powered incremental conversion testing tool for marketers and lean agencies. Instead of launching one big A/B test and waiting, it runs a sequence of small…
Market Research
Global digital ad spend is projected to exceed $700 billion by 2025, with AI-driven optimization cited as a primary growth driver (Ideabrowser opportunity analysis on idea #38).…
Competitive Landscape
Google Ads Experiments — Native A/B and geo experiments, incrementality uplift analysis, bid and creative testing. Deep integration, trusted analytics, zero incremental SaaS…
Business Model
Free — Ad Optimization Audit ($0) — Paste campaign exports or connect read-only; instant report highlighting wasted spend, untested variables, and quick-win test suggestions. Lead…
Recommended Tech Stack
Next.js 14 + Vercel — App Router for marketer dashboard, Edge routes for OAuth callbacks and webhook receivers from ad platforms. Vercel Cron for nightly sync jobs.
AI Prompts to Build This
Copy these build prompts into Claude, Cursor, or your AI coding tool. Create a free account to unlock the full research behind them.
1. Project Setup
Build the weekend MVP of "GutCheck": generate single-change ad variants from a base ad, paste in how each one performed, and see which ones truly beat the baseline and which are still noise. It does not connect to ad platforms. Stack: Next.js (App Router), TypeScript, Tailwind, Supabase (Postgres, Row Level Security, Auth with email magic link), the OpenAI API (a small model) to write variants, Zod to validate its output. The statistics are plain code. Deploy on Vercel. Tables (Row Level Security on, each user reads only their own rows): - experiments(id, user_id, name, variable_type, status, min_conversions, created_at) variable_type is headline, primary_text or cta - variants(id, experiment_id, is_baseline boolean, headline, primary_text, cta, mutated_field, mutation_rationale, hypothesis) - results(id, variant_id, recorded_on, impressions, clicks, conversions, spend_cents) - lift_results(id, experiment_id, variant_id, lift, confidence, verdict, computed_at) verdict is winner, loser or keep_running Screens: /login, /experiments/new, /experiments/[id] (variants, results upload and verdicts). Env vars (names only): NEXT_PUBLIC_SUPABASE_URL, NEXT_PUBLIC_SUPABASE_ANON_KEY, SUPABASE_SERVICE_ROLE_KEY (server only), OPENAI_API_KEY. Do not build: billing, plans or usage meters, Meta or Google Ads connections, pushing ads live, image variants, a job queue or Redis, automatic pausing of ads. Done when: npm run dev starts, you can sign in, and the four tables exist with Row Level Security on.
2. Core Feature
Build the one feature that proves GutCheck: variants that change one thing, and a verdict you can trust. 1. An experiment starts from a base ad (headline, primary text, CTA) and a variable_type. Save the base as the baseline variant. 2. A server action generates 6 variants. It calls the OpenAI API with the base ad and the variable type, and asks for a JSON array of { headline, primary_text, cta, mutation_rationale, hypothesis } that changes only the chosen field and keeps the brand voice and every factual claim. 3. Validate with Zod, then compare each variant to the baseline in code. Reject the whole batch if any variant differs in more than one field or adds a claim (a price, a number or a guarantee) that the base ad does not have. Set mutated_field from the comparison, not from the model. 4. The user runs the variants elsewhere and uploads a CSV with columns variant_id, date, impressions, clicks, conversions, spend. 5. For each variant against the baseline, compute the conversion rates, the relative lift and a two-proportion z-test. Report confidence as 1 minus the p-value. Mark winner at 95 percent confidence or more with a positive lift, loser at 95 percent or more with a negative lift, and keep_running otherwise. 6. Never mark a winner or loser until both the variant and the baseline have at least min_conversions conversions (default 100). Show "not enough data yet" with how many conversions are missing. Rules: label the result "lift against the baseline in the same period", never incremental sales lift. Unit test the statistics with known counts. Empty state: with no experiments, show the form and a sample base ad. Done when: a batch with a two-field change is rejected, an uploaded sample gives the expected z-score in the unit test, a variant under the conversion minimum shows keep_running with a count, and a clear winner shows its lift and confidence.3. Landing Page
Build a one-page landing site for GutCheck, an ad creative testing tool for lean marketing teams. Hero: "Stop killing ads on gut instinct." Sub: "Test one change at a time. See which creative actually beat the baseline and which is still noise." One button: Join the waitlist for the free ad audit. Sections: the problem (platform dashboards show correlation, not a clean test), how it works in four steps (base ad, variants, results, verdict), a sample verdict table with a winner, a loser and one "not enough data yet", a placeholder for real lift numbers to replace with yours, and an FAQ on the statistics (a 95 percent threshold and a conversion minimum) and data privacy. Waitlist: store the email in a waitlist table in Supabase. No other service. Style: Geist, a near-black background, one emerald accent. Done when: the page renders on a phone and a submitted email appears in the waitlist table.
4. Branding Package
Use a design or image tool for this one. A coding agent cannot draw a logo. Brand for GutCheck: a wordmark and an icon that suggest a check mark built from two bars of different height. Colors: near-black, off-white and one emerald. Type: Geist, with tabular numbers for lift and confidence. Deliverables: wordmark, icon, three verdict badges (winner, loser, keep running) that differ by shape as well as color, and one launch graphic built from a verdict table. Done when: each deliverable is saved in one folder and the verdict badges are distinguishable without color.