CRO EngineArchitecture & Engineering Design

Engineering Documentation

Comprehensive system architecture, design decisions, implementation details, and engineering trade-offs for the CRO Engine — written at senior engineer level.

Version 1.0.0ProductionNext.js 15 · Gemini AI · MongoDB Atlas

1. Problem Statement

1.1 Background

Conversion Rate Optimization (CRO) is the systematic process of increasing the percentage of website visitors who take a desired action — purchasing a product, adding items to cart, or completing checkout.

For Shopify merchants, the difference between a 1% and 3% conversion rate on a store generating $1M/year is $20,000 in incremental revenue — without acquiring a single additional visitor.

Professional CRO audits from agencies cost $5,000–$25,000 per engagement, take 2–4 weeks, and require specialized knowledge of e-commerce UX heuristics, consumer psychology, and Shopify-specific patterns.

1.2 The Problem

There is no accessible, automated tool that: understands Shopify-specific storefront patterns, applies established CRO heuristics, produces actionable prioritized recommendations, and runs at internet scale in seconds rather than weeks.

SolutionGap
Google PageSpeed / GTmetrixTechnical performance only — no conversion psychology
Hotjar / VWO / OptimizelyRequires existing traffic data and weeks of A/B testing
Agency CRO AuditsExpensive ($5K–$25K), slow (2–4 weeks), not scalable

2. Goals & Non-Goals

GoalPriorityDescription
URL-to-report in <30 secondsP0Core product promise: instant analysis
Structured JSON recommendationsP0Machine-readable, typed output
SSRF protectionP0No internal network access via URL input
MongoDB persistenceP1Audits stored and retrievable by ID
Dashboard for past auditsP1Browse and compare historical results
PWA supportP2Manifest, theme color, icon set

Non-Goals

Non-GoalRationale
Real-time A/B testingRequires traffic instrumentation — separate product domain
Browser-rendered SPA scrapingPuppeteer overhead 10x — most Shopify stores render SSR
Multi-language supportEnglish-only in v1
User authenticationSingle-user portfolio project in v1

3. Requirements

3.1 Functional Requirements

FR-01 — URL Input: Accept a public HTTPS URL, normalize it, validate against SSRF protection, reject malformed inputs before any I/O.

FR-02 — Storefront Scraping: Fetch and parse the HTML, extracting headings, CTAs, navigation, and product metadata without executing JavaScript.

FR-03 — DOM Minification: Strip scripts, styles, SVG, and redundant layout — reduce HTML token count by ≥70%.

FR-04 — AI Audit: Use a versioned prompt with Gemini to produce a structured JSON `AuditReport`.

FR-05 — Self-Correction: Validate AI responses via Zod. Retry with a self-correction prompt up to 3 times on failure.

FR-06 — Persistence: Persist each `AuditReport` to MongoDB Atlas with idempotent upsert on audit ID.

3.2 Non-Functional Requirements

NFRTarget
URL validation latency<50ms
DOM extraction latency<5 seconds
AI analysis latency<20 seconds
Total end-to-end (P95)<30 seconds
AI retry limit≤3 attempts
Test pass rate (CI)100%
TypeScript errors (CI)0

4. System Architecture Overview

CRO Engine is a full-stack Next.js 15 application. The frontend, backend API routes, and server-side services all live within a single Next.js project — eliminating the need for a separate API server.

4.1 Request Lifecycle

1. User submits URL → POST /api/v1/analyze
2. requestId generated: "req_" + 8-char alphanumeric
3. Input parsed → Zod validation (400 on failure)
4. URL normalized → SSRF check (422 on block)
5. SnapshotOrchestrator.generateSnapshot(url)
   a. WebsiteFetcher.fetch(url)  → raw HTML
   b. DomMinifier.minify(html)   → ~85% smaller
   c. Returns WebsiteSnapshot
6. AnalysisOrchestrator.analyze(snapshot)
   a. ContextBuilder.build()     → prompt context string
   b. PromptBuilder.build()      → full system+user prompt
   c. GeminiClient.generate()    → raw AI text
   d. JsonParser.extract()       → raw JSON object
   e. SchemaValidator.validate() → typed AuditData
      → if fail: retry (max 3 attempts)
   f. DomainMapper.toDomain()    → AuditReport
7. AuditRepository.save(report) → MongoDB upsert
8. Return { id } → client redirects to /audits/:id

5. Frontend Architecture

TechnologyVersionRationale
Next.js15.xApp Router, Server Components, collocated API routes
React19.xConcurrent rendering, improved hydration
TypeScript5.6Strict mode, full type safety
Tailwind CSS4.xCSS custom properties design token system
react-hook-form7.xUncontrolled forms with Zod resolver
CVA0.7Type-safe component variants
lucide-react0.468Consistent, tree-shakeable SVG icons

5.1 App Router Structure

src/app/
├── layout.tsx         # Root layout: Header, Footer, global metadata
├── page.tsx           # Landing page (/)
├── error.tsx          # Global error boundary
├── not-found.tsx      # 404 page
├── sitemap.ts         # Dynamic sitemap.xml generator
├── robots.ts          # robots.txt generator
├── manifest.ts        # PWA manifest generator
├── dashboard/page.tsx # Past audits list (/dashboard)
├── audits/[id]/page.tsx  # Audit detail (/audits/:id)
├── docs/page.tsx      # Documentation (/docs)
└── api/v1/            # Backend API routes

5.2 Design System

The design system is defined entirely in src/styles/globals.css using CSS custom properties. Tailwind v4 references these tokens via the @theme directive:

/* Typography */
--text-primary: #0f172a;
--text-secondary: #475569;
--text-muted: #94a3b8;

/* Backgrounds */
--bg-primary: #ffffff;
--bg-secondary: #f8fafc;
--bg-card: #ffffff;

/* Accents */
--accent-violet: #1e40af;  /* Primary brand */
--accent-emerald: #059669; /* Success / good scores */
--accent-amber: #d97706;   /* Warning / medium scores */
--accent-rose: #e11d48;    /* Error / low scores */

6. Backend Architecture

CRO Engine uses Next.js 15 App Router API routes (route.ts files) as the backend layer. This eliminates the need for a separate Express.js server, reducing operational complexity and enabling TypeScript sharing across client and server.

6.1 Error Hierarchy

class ApiError extends Error {
  constructor(
    public code: ErrorCodes,
    public message: string,
    public statusCode: number,
    public cause?: unknown,
  ) {}
}

class ValidationError extends ApiError { /* 400 */ }
class NotFoundError   extends ApiError { /* 404 */ }
class ScrapingError   extends ApiError { /* 422 */ }
class AiAnalysisError extends ApiError { /* 503 */ }

6.2 Structured Logging

// Every log entry is structured JSON
{
  "timestamp": "2026-07-04T16:00:00.000Z",
  "level": "INFO",
  "message": "Website extraction completed",
  "meta": {
    "requestId": "req_abc123",
    "durationMs": 2847
  }
}

7. AI Pipeline Design

The AI analysis pipeline transforms variable, messy real-world HTML into a structured, typed domain model through a series of deterministic transformations.

7.1 Self-Correction Retry Loop

The self-correction loop is the most critical reliability mechanism. When Gemini output fails Zod validation, the system retries with a prompt that includes the exact validation errors — giving the model precise instructions on what to fix.
for (let attempt = 1; attempt <= 3; attempt++) {
  // First attempt: standard prompt
  // Subsequent: self-correction prompt with error context
  const prompt = attempt === 1
    ? promptBuilder.build(context)
    : promptBuilder.buildCorrection(context, lastErrors);

  const rawText = await geminiClient.generate(prompt);
  const rawJson = jsonParser.extract(rawText);
  const result  = schemaValidator.validate(rawJson);

  if (result.success) {
    return domainMapper.toDomain(result.data, snapshot);
  }

  lastErrors = JSON.stringify(result.error.format());
  logger.warn('Retrying with self-correction...', { attempt });
}

throw new AiAnalysisError('Max retries exceeded');

7.2 Gemini Configuration

ParameterValueRationale
temperature0.2Low = deterministic JSON, fewer hallucinations
maxOutputTokens4096Enough for 8–12 detailed recommendations
timeout25 secondsHard ceiling on AI response time
modelgemini-1.5-flashSpeed/cost optimized; Pro as fallback

8. Website Intelligence Pipeline

8.1 DOM Minification

The DOM Minifier is the most important performance engineering decision. Raw Shopify HTML frequently exceeds 400–800KB. The minifier applies Cheerio-based transforms to reduce this to ~15% of original size:

StoreRaw HTMLMinifiedReduction
Gymshark487 KB72 KB85.2%
Allbirds312 KB48 KB84.6%
ColourPop698 KB104 KB85.1%
// Phases of minification
$('script, style, noscript, svg, iframe').remove();
$('link[rel="stylesheet"]').remove();
$('[data-reactroot], #__NEXT_DATA__').remove();

// Strip data-* and on* attributes
$('*').each((_, el) => {
  Object.keys(el.attribs).forEach(attr => {
    if (attr.startsWith('on') || attr.startsWith('data-')) {
      $(el).removeAttr(attr);
    }
  });
});

// Collapse whitespace
return $.html().replace(/\s+/g, ' ').trim();

8.2 Brand Identity Extraction

To populate storefront visual elements dynamically, the pipeline crawlers scrape brand visual assets. Icons and favicons are resolved in order of priority (rel="icon" &rarr; rel="shortcut icon" &rarr; rel="apple-touch-icon"), and DOM image elements are inspected using Alt/Class/ID selectors to identify logos. All resolved assets are verified as active and reachable asynchronously using a 1.5-second fetch HEAD/GET validation cycle.

9. MongoDB Database Design

9.1 Collection: audits

{
  id: string,           // "aud_" + nanoid(8), unique
  url?: string,         // Fully qualified store URL
  storeUrl: string,     // Normalized store URL
  domain?: string,      // Scraped store domain
  storeName?: string,   // Scraped store brand name
  title?: string,       // Store home page title
  description?: string, // Store meta description
  logoUrl?: string,     // Verified storefront logo
  faviconUrl?: string,  // Verified storefront favicon
  appleTouchIcon?: string, // Mobile launcher touch icon
  themeColor?: string,  // HTML Meta themeColor value
  brandColor?: string,  // Computed brand color value
  platform?: string,    // Engine: 'shopify' | 'unknown'
  status?: string,      // Execution state: 'completed' | 'running' | 'failed'
  analysisTime?: number, // Execution latency in seconds
  createdAt?: string,   // Record ISO timestamp
  updatedAt?: string,   // Update ISO timestamp
  overallScore: number, // 0–100 CRO composite score
  analyzedAt: string,   // ISO 8601 timestamp
  pageScores: {
    homepage: number,
    pdp: number,
    collection: number,
    cart: number,
  },
  recommendations: [{
    id: string,
    pageType: 'homepage' | 'pdp' | 'collection' | 'cart',
    category: 'copywriting' | 'layout' | 'cta' | 'trust'
             | 'mobile' | 'performance',
    finding: string,
    rationale: string,
    actionSteps: string[],
    impact: 'HIGH' | 'MEDIUM' | 'LOW',
    effort: 'HIGH' | 'MEDIUM' | 'LOW',
  }],
}

9.2 Indexes

IndexFieldsPurpose
ux_audits_id{ id: 1 } uniqueO(1) audit lookup by ID
ix_audits_store_url{ storeUrl: 1 }Filter audits by store
ix_audits_analyzed_at{ analyzedAt: -1 }Chronological sort for dashboard
ix_audits_overall_score{ overallScore: -1 }Sort dashboard by CRO performance
ix_audits_store_name{ storeName: 1 }Alphabetical sorting by merchant name

10. API Design & Contracts

All API responses use a consistent ApiResponse<T> envelope for predictable client-side handling:

// Success
{ "success": true, "data": { /* T */ } }

// Error
{ "success": false, "error": {
    "code": "VALIDATION_ERROR",
    "message": "URL is required",
    "details": { /* Zod error format */ }
  }
}
MethodEndpointDescriptionAuth
POST/api/v1/analyzeSubmit URL for AI audit and extract brand assetsNone
GET/api/v1/auditsList paginated, filtered, and sorted audits with statsNone
GET/api/v1/audits/:idGet specific audit details by IDNone
POST/api/v1/extractDebug: raw DOM snapshot and branding parsingNone
GET/api/v1/healthHealth checkNone

11. Folder Structure

src/
├── app/              # Next.js App Router (pages + API routes)
├── components/       # UI component library
│   ├── common/       # Container, PageWrapper, ErrorBoundary
│   ├── feedback/     # Alert, Badge, Progress, Skeleton, Toast
│   ├── forms/        # Input, Textarea, Form
│   ├── layout/       # Header, Footer
│   └── ui/           # Button, Card, Badge (atomic)
├── config/           # env.ts, config.ts, routes.ts
├── lib/              # analytics.ts (vendor-agnostic)
├── server/           # All server-side services (never imported by pages directly)
│   ├── ai/           # AI pipeline (client, context, mapper, parser, prompts)
│   ├── crawler/      # WebsiteFetcher
│   ├── db/           # mongodb-client.ts, audit-repository.ts
│   ├── errors/       # ApiError hierarchy
│   ├── logger/       # StructuredLogger
│   ├── normalizer/   # URL normalization + SSRF
│   ├── snapshot/     # DOM extraction orchestration
│   └── validation/   # Zod input schemas
├── styles/           # globals.css (design tokens)
└── utils/            # cn, retry, timeout, safe-json-parse, assert-never

12. Prompt Engineering

12.1 Versioning Strategy

All system prompts are stored as flat markdown files under prompts/v1.0.0/. Prompt versions follow semantic versioning — a major bump signals a breaking change to the output schema.

12.2 Prompt Structure

SYSTEM PROMPT:
"You are an expert Conversion Rate Optimization specialist with 15 years
of e-commerce experience. Evaluate the storefront against:
- Hick's Law (decision fatigue)
- Social proof visibility
- CTA clarity and visual hierarchy
- Above-the-fold value proposition
- Cart abandonment friction
- Mobile-first layout
Return ONLY the JSON object conforming to: {schema}"

USER PROMPT:
"Analyze: https://gymshark.com
HEADINGS: H1: 'Built for Champions', H2: 'Shop Men's Training'...
CTAs: 'Shop Now' → /collections/mens, 'Add to Bag' → (cart)...
CONTENT SAMPLE: [2000 chars of minified body text]"

13. Security Architecture

SSRF (Server-Side Request Forgery) is the most critical security risk. An attacker could supply internal URLs like http://169.254.169.254/ (AWS metadata) or http://localhost:27017 (MongoDB) to pivot into internal infrastructure.

13.1 SSRF Protection — 4-Stage Pipeline

StageCheckBlocks
1Blocklisted hostnameslocalhost, 127.0.0.1, ::1, metadata.google.internal
2Private IP ranges (post-DNS)10.x, 172.16–31.x, 192.168.x, fc00:, fe80:
3Protocol enforcementOnly https: accepted
4URL parsing hardeningMalformed URLs that throw TypeError

13.2 Security Headers

'X-Frame-Options': 'DENY'
'X-Content-Type-Options': 'nosniff'
'Referrer-Policy': 'strict-origin-when-cross-origin'
'Permissions-Policy': 'camera=(), microphone=(), geolocation=()'
'Strict-Transport-Security': 'max-age=31536000; includeSubDomains'

14. Performance Architecture

14.1 Token Cost Optimization (AI)

The DOM Minifier is the primary cost lever. Gemini Flash is priced per input token:

Scenario100 req/dayCost/day
Without minification48.7 MB HTML tokens~$9.74
With minification (85% reduction)7.2 MB HTML tokens~$1.44

14.2 Rendering Strategies per Page

PageStrategyRationale
/ (landing)StaticNo dynamic data; CDN cached
/dashboardClient-side fetchData changes frequently
/audits/[id]Client-side fetchPersonalized audit data
/docsStaticMarkdown content; fully cacheable

15. Testing Strategy

Core principle: "Test the contract, not the implementation." Each test verifies behavior from the outside, making tests robust to internal refactoring.
Test SuiteCasesCoverage
url-normalizer.test.ts6HTTPS enforcement, SSRF blocking
schemas.test.ts8Zod validation schemas
snapshot-validator.test.ts2Minimum content thresholds
context-builder.test.ts1Prompt context formatting
prompt-builder.test.ts1Template rendering
domain-mapper.test.ts2AI output → AuditReport mapping
json-parser.test.ts5Fence stripping, trailing commas
app-error.test.ts4Error class hierarchy
audit-repository.test.ts4save, findById, findRecent
analysis-orchestrator.test.ts5Retry loop, error bubbling
extract-route.test.ts2API route integration
analyze-route.test.ts2API route integration
Total42100% pass rate

16. Deployment Architecture

LayerServiceConfig
HostingVercelNext.js first-party — zero config
DatabaseMongoDB Atlas M0512MB free tier
AIGoogle AI StudioGemini 1.5 Flash API
CDNVercel Edge NetworkAutomatic via Vercel deployment
CI/CDGitHub ActionsLint → typecheck → test on every push

16.1 CI Pipeline

# .github/workflows/ci.yml
jobs:
  ci:
    runs-on: ubuntu-latest
    steps:
      - npm ci
      - npm run lint          # ESLint
      - npm run check-types   # tsc --noEmit (strict)
      - npm run test          # vitest run
# All three must pass — PR merge blocked otherwise

17. Engineering Trade-offs

Cheerio vs Puppeteer

AspectCheerio ✅Puppeteer ❌
Speed<1 second5–15 seconds (browser startup)
Memory~50 MB~300 MB per instance
JS executionNoYes
Shopify SSR contentAvailableAvailable
Cost (serverless)LowHigh (memory/time billing)

Gemini vs GPT-4o

AspectGemini 1.5 Flash ✅GPT-4o
CostLowerHigher
Context window1M tokens128K tokens
Free tierGenerousLimited
JSON reliabilityGood (with Zod retry)Good (with function calling)

MongoDB vs PostgreSQL

AspectMongoDB ✅PostgreSQL
Schema flexibilityNo migrations neededSchema changes require migrations
Nested arraysNative document modelRequires JSONB or JOIN table
Type safety (TS)Good via typed genericsExcellent via Prisma

18. Future Work & Roadmap

FeaturePriorityDescription
User AuthenticationP1Clerk/NextAuth — multi-tenant audit history
PDF Report ExportP2Professionally formatted PDF download
Scheduled AuditsP2Weekly automated re-analysis per store
Recommendation TrackingP1Mark implemented, track score delta
Competitor ComparisonP3Side-by-side audit of two stores
Puppeteer FallbackP2Render JS-heavy SPA stores when Cheerio misses content
Redis CachingP2Cache identical URL analyses for 24h
Shopify App IntegrationP3Native Shopify App Store listing with OAuth
Fine-tuned ModelP3Improve schema reliability with domain-specific training

Open Architecture Questions

1. Multi-page crawling: Should we crawl 3–5 pages (homepage + PDP + collection + cart) simultaneously rather than just the homepage for deeper coverage?

2. Vector embeddings: Could we embed past recommendations and semantically search them to provide "similar findings" context — improving AI quality without increasing prompt length?

3. Fine-tuning: Could a fine-tuned model on historical CRO audit data outperform prompt-based Gemini on structured output reliability and domain-specific recommendation quality?

CRO Engine Engineering Documentation · Version 1.0.0 · Sprint 7 · July 2026

This document is maintained alongside the codebase. Update when architecture decisions change.