Engineering Documentation
Comprehensive system architecture, design decisions, implementation details, and engineering trade-offs for the CRO Engine — written at senior engineer level.
1. Problem Statement
1.1 Background
Conversion Rate Optimization (CRO) is the systematic process of increasing the percentage of website visitors who take a desired action — purchasing a product, adding items to cart, or completing checkout.
For Shopify merchants, the difference between a 1% and 3% conversion rate on a store generating $1M/year is $20,000 in incremental revenue — without acquiring a single additional visitor.
1.2 The Problem
There is no accessible, automated tool that: understands Shopify-specific storefront patterns, applies established CRO heuristics, produces actionable prioritized recommendations, and runs at internet scale in seconds rather than weeks.
| Solution | Gap |
|---|---|
| Google PageSpeed / GTmetrix | Technical performance only — no conversion psychology |
| Hotjar / VWO / Optimizely | Requires existing traffic data and weeks of A/B testing |
| Agency CRO Audits | Expensive ($5K–$25K), slow (2–4 weeks), not scalable |
2. Goals & Non-Goals
| Goal | Priority | Description |
|---|---|---|
| URL-to-report in <30 seconds | P0 | Core product promise: instant analysis |
| Structured JSON recommendations | P0 | Machine-readable, typed output |
| SSRF protection | P0 | No internal network access via URL input |
| MongoDB persistence | P1 | Audits stored and retrievable by ID |
| Dashboard for past audits | P1 | Browse and compare historical results |
| PWA support | P2 | Manifest, theme color, icon set |
Non-Goals
| Non-Goal | Rationale |
|---|---|
| Real-time A/B testing | Requires traffic instrumentation — separate product domain |
| Browser-rendered SPA scraping | Puppeteer overhead 10x — most Shopify stores render SSR |
| Multi-language support | English-only in v1 |
| User authentication | Single-user portfolio project in v1 |
3. Requirements
3.1 Functional Requirements
FR-01 — URL Input: Accept a public HTTPS URL, normalize it, validate against SSRF protection, reject malformed inputs before any I/O.
FR-02 — Storefront Scraping: Fetch and parse the HTML, extracting headings, CTAs, navigation, and product metadata without executing JavaScript.
FR-03 — DOM Minification: Strip scripts, styles, SVG, and redundant layout — reduce HTML token count by ≥70%.
FR-04 — AI Audit: Use a versioned prompt with Gemini to produce a structured JSON `AuditReport`.
FR-05 — Self-Correction: Validate AI responses via Zod. Retry with a self-correction prompt up to 3 times on failure.
FR-06 — Persistence: Persist each `AuditReport` to MongoDB Atlas with idempotent upsert on audit ID.
3.2 Non-Functional Requirements
| NFR | Target |
|---|---|
| URL validation latency | <50ms |
| DOM extraction latency | <5 seconds |
| AI analysis latency | <20 seconds |
| Total end-to-end (P95) | <30 seconds |
| AI retry limit | ≤3 attempts |
| Test pass rate (CI) | 100% |
| TypeScript errors (CI) | 0 |
4. System Architecture Overview
CRO Engine is a full-stack Next.js 15 application. The frontend, backend API routes, and server-side services all live within a single Next.js project — eliminating the need for a separate API server.
4.1 Request Lifecycle
1. User submits URL → POST /api/v1/analyze
2. requestId generated: "req_" + 8-char alphanumeric
3. Input parsed → Zod validation (400 on failure)
4. URL normalized → SSRF check (422 on block)
5. SnapshotOrchestrator.generateSnapshot(url)
a. WebsiteFetcher.fetch(url) → raw HTML
b. DomMinifier.minify(html) → ~85% smaller
c. Returns WebsiteSnapshot
6. AnalysisOrchestrator.analyze(snapshot)
a. ContextBuilder.build() → prompt context string
b. PromptBuilder.build() → full system+user prompt
c. GeminiClient.generate() → raw AI text
d. JsonParser.extract() → raw JSON object
e. SchemaValidator.validate() → typed AuditData
→ if fail: retry (max 3 attempts)
f. DomainMapper.toDomain() → AuditReport
7. AuditRepository.save(report) → MongoDB upsert
8. Return { id } → client redirects to /audits/:id5. Frontend Architecture
| Technology | Version | Rationale |
|---|---|---|
| Next.js | 15.x | App Router, Server Components, collocated API routes |
| React | 19.x | Concurrent rendering, improved hydration |
| TypeScript | 5.6 | Strict mode, full type safety |
| Tailwind CSS | 4.x | CSS custom properties design token system |
| react-hook-form | 7.x | Uncontrolled forms with Zod resolver |
| CVA | 0.7 | Type-safe component variants |
| lucide-react | 0.468 | Consistent, tree-shakeable SVG icons |
5.1 App Router Structure
src/app/
├── layout.tsx # Root layout: Header, Footer, global metadata
├── page.tsx # Landing page (/)
├── error.tsx # Global error boundary
├── not-found.tsx # 404 page
├── sitemap.ts # Dynamic sitemap.xml generator
├── robots.ts # robots.txt generator
├── manifest.ts # PWA manifest generator
├── dashboard/page.tsx # Past audits list (/dashboard)
├── audits/[id]/page.tsx # Audit detail (/audits/:id)
├── docs/page.tsx # Documentation (/docs)
└── api/v1/ # Backend API routes5.2 Design System
The design system is defined entirely in src/styles/globals.css using CSS custom properties. Tailwind v4 references these tokens via the @theme directive:
/* Typography */
--text-primary: #0f172a;
--text-secondary: #475569;
--text-muted: #94a3b8;
/* Backgrounds */
--bg-primary: #ffffff;
--bg-secondary: #f8fafc;
--bg-card: #ffffff;
/* Accents */
--accent-violet: #1e40af; /* Primary brand */
--accent-emerald: #059669; /* Success / good scores */
--accent-amber: #d97706; /* Warning / medium scores */
--accent-rose: #e11d48; /* Error / low scores */6. Backend Architecture
CRO Engine uses Next.js 15 App Router API routes (route.ts files) as the backend layer. This eliminates the need for a separate Express.js server, reducing operational complexity and enabling TypeScript sharing across client and server.
6.1 Error Hierarchy
class ApiError extends Error {
constructor(
public code: ErrorCodes,
public message: string,
public statusCode: number,
public cause?: unknown,
) {}
}
class ValidationError extends ApiError { /* 400 */ }
class NotFoundError extends ApiError { /* 404 */ }
class ScrapingError extends ApiError { /* 422 */ }
class AiAnalysisError extends ApiError { /* 503 */ }6.2 Structured Logging
// Every log entry is structured JSON
{
"timestamp": "2026-07-04T16:00:00.000Z",
"level": "INFO",
"message": "Website extraction completed",
"meta": {
"requestId": "req_abc123",
"durationMs": 2847
}
}7. AI Pipeline Design
The AI analysis pipeline transforms variable, messy real-world HTML into a structured, typed domain model through a series of deterministic transformations.
7.1 Self-Correction Retry Loop
for (let attempt = 1; attempt <= 3; attempt++) {
// First attempt: standard prompt
// Subsequent: self-correction prompt with error context
const prompt = attempt === 1
? promptBuilder.build(context)
: promptBuilder.buildCorrection(context, lastErrors);
const rawText = await geminiClient.generate(prompt);
const rawJson = jsonParser.extract(rawText);
const result = schemaValidator.validate(rawJson);
if (result.success) {
return domainMapper.toDomain(result.data, snapshot);
}
lastErrors = JSON.stringify(result.error.format());
logger.warn('Retrying with self-correction...', { attempt });
}
throw new AiAnalysisError('Max retries exceeded');7.2 Gemini Configuration
| Parameter | Value | Rationale |
|---|---|---|
| temperature | 0.2 | Low = deterministic JSON, fewer hallucinations |
| maxOutputTokens | 4096 | Enough for 8–12 detailed recommendations |
| timeout | 25 seconds | Hard ceiling on AI response time |
| model | gemini-1.5-flash | Speed/cost optimized; Pro as fallback |
8. Website Intelligence Pipeline
8.1 DOM Minification
The DOM Minifier is the most important performance engineering decision. Raw Shopify HTML frequently exceeds 400–800KB. The minifier applies Cheerio-based transforms to reduce this to ~15% of original size:
| Store | Raw HTML | Minified | Reduction |
|---|---|---|---|
| Gymshark | 487 KB | 72 KB | 85.2% |
| Allbirds | 312 KB | 48 KB | 84.6% |
| ColourPop | 698 KB | 104 KB | 85.1% |
// Phases of minification
$('script, style, noscript, svg, iframe').remove();
$('link[rel="stylesheet"]').remove();
$('[data-reactroot], #__NEXT_DATA__').remove();
// Strip data-* and on* attributes
$('*').each((_, el) => {
Object.keys(el.attribs).forEach(attr => {
if (attr.startsWith('on') || attr.startsWith('data-')) {
$(el).removeAttr(attr);
}
});
});
// Collapse whitespace
return $.html().replace(/\s+/g, ' ').trim();8.2 Brand Identity Extraction
To populate storefront visual elements dynamically, the pipeline crawlers scrape brand visual assets. Icons and favicons are resolved in order of priority (rel="icon" → rel="shortcut icon" → rel="apple-touch-icon"), and DOM image elements are inspected using Alt/Class/ID selectors to identify logos. All resolved assets are verified as active and reachable asynchronously using a 1.5-second fetch HEAD/GET validation cycle.
9. MongoDB Database Design
9.1 Collection: audits
{
id: string, // "aud_" + nanoid(8), unique
url?: string, // Fully qualified store URL
storeUrl: string, // Normalized store URL
domain?: string, // Scraped store domain
storeName?: string, // Scraped store brand name
title?: string, // Store home page title
description?: string, // Store meta description
logoUrl?: string, // Verified storefront logo
faviconUrl?: string, // Verified storefront favicon
appleTouchIcon?: string, // Mobile launcher touch icon
themeColor?: string, // HTML Meta themeColor value
brandColor?: string, // Computed brand color value
platform?: string, // Engine: 'shopify' | 'unknown'
status?: string, // Execution state: 'completed' | 'running' | 'failed'
analysisTime?: number, // Execution latency in seconds
createdAt?: string, // Record ISO timestamp
updatedAt?: string, // Update ISO timestamp
overallScore: number, // 0–100 CRO composite score
analyzedAt: string, // ISO 8601 timestamp
pageScores: {
homepage: number,
pdp: number,
collection: number,
cart: number,
},
recommendations: [{
id: string,
pageType: 'homepage' | 'pdp' | 'collection' | 'cart',
category: 'copywriting' | 'layout' | 'cta' | 'trust'
| 'mobile' | 'performance',
finding: string,
rationale: string,
actionSteps: string[],
impact: 'HIGH' | 'MEDIUM' | 'LOW',
effort: 'HIGH' | 'MEDIUM' | 'LOW',
}],
}9.2 Indexes
| Index | Fields | Purpose |
|---|---|---|
| ux_audits_id | { id: 1 } unique | O(1) audit lookup by ID |
| ix_audits_store_url | { storeUrl: 1 } | Filter audits by store |
| ix_audits_analyzed_at | { analyzedAt: -1 } | Chronological sort for dashboard |
| ix_audits_overall_score | { overallScore: -1 } | Sort dashboard by CRO performance |
| ix_audits_store_name | { storeName: 1 } | Alphabetical sorting by merchant name |
10. API Design & Contracts
All API responses use a consistent ApiResponse<T> envelope for predictable client-side handling:
// Success
{ "success": true, "data": { /* T */ } }
// Error
{ "success": false, "error": {
"code": "VALIDATION_ERROR",
"message": "URL is required",
"details": { /* Zod error format */ }
}
}| Method | Endpoint | Description | Auth |
|---|---|---|---|
| POST | /api/v1/analyze | Submit URL for AI audit and extract brand assets | None |
| GET | /api/v1/audits | List paginated, filtered, and sorted audits with stats | None |
| GET | /api/v1/audits/:id | Get specific audit details by ID | None |
| POST | /api/v1/extract | Debug: raw DOM snapshot and branding parsing | None |
| GET | /api/v1/health | Health check | None |
11. Folder Structure
src/
├── app/ # Next.js App Router (pages + API routes)
├── components/ # UI component library
│ ├── common/ # Container, PageWrapper, ErrorBoundary
│ ├── feedback/ # Alert, Badge, Progress, Skeleton, Toast
│ ├── forms/ # Input, Textarea, Form
│ ├── layout/ # Header, Footer
│ └── ui/ # Button, Card, Badge (atomic)
├── config/ # env.ts, config.ts, routes.ts
├── lib/ # analytics.ts (vendor-agnostic)
├── server/ # All server-side services (never imported by pages directly)
│ ├── ai/ # AI pipeline (client, context, mapper, parser, prompts)
│ ├── crawler/ # WebsiteFetcher
│ ├── db/ # mongodb-client.ts, audit-repository.ts
│ ├── errors/ # ApiError hierarchy
│ ├── logger/ # StructuredLogger
│ ├── normalizer/ # URL normalization + SSRF
│ ├── snapshot/ # DOM extraction orchestration
│ └── validation/ # Zod input schemas
├── styles/ # globals.css (design tokens)
└── utils/ # cn, retry, timeout, safe-json-parse, assert-never12. Prompt Engineering
12.1 Versioning Strategy
All system prompts are stored as flat markdown files under prompts/v1.0.0/. Prompt versions follow semantic versioning — a major bump signals a breaking change to the output schema.
12.2 Prompt Structure
SYSTEM PROMPT:
"You are an expert Conversion Rate Optimization specialist with 15 years
of e-commerce experience. Evaluate the storefront against:
- Hick's Law (decision fatigue)
- Social proof visibility
- CTA clarity and visual hierarchy
- Above-the-fold value proposition
- Cart abandonment friction
- Mobile-first layout
Return ONLY the JSON object conforming to: {schema}"
USER PROMPT:
"Analyze: https://gymshark.com
HEADINGS: H1: 'Built for Champions', H2: 'Shop Men's Training'...
CTAs: 'Shop Now' → /collections/mens, 'Add to Bag' → (cart)...
CONTENT SAMPLE: [2000 chars of minified body text]"13. Security Architecture
http://169.254.169.254/ (AWS metadata) or http://localhost:27017 (MongoDB) to pivot into internal infrastructure.13.1 SSRF Protection — 4-Stage Pipeline
| Stage | Check | Blocks |
|---|---|---|
| 1 | Blocklisted hostnames | localhost, 127.0.0.1, ::1, metadata.google.internal |
| 2 | Private IP ranges (post-DNS) | 10.x, 172.16–31.x, 192.168.x, fc00:, fe80: |
| 3 | Protocol enforcement | Only https: accepted |
| 4 | URL parsing hardening | Malformed URLs that throw TypeError |
13.2 Security Headers
'X-Frame-Options': 'DENY'
'X-Content-Type-Options': 'nosniff'
'Referrer-Policy': 'strict-origin-when-cross-origin'
'Permissions-Policy': 'camera=(), microphone=(), geolocation=()'
'Strict-Transport-Security': 'max-age=31536000; includeSubDomains'14. Performance Architecture
14.1 Token Cost Optimization (AI)
The DOM Minifier is the primary cost lever. Gemini Flash is priced per input token:
| Scenario | 100 req/day | Cost/day |
|---|---|---|
| Without minification | 48.7 MB HTML tokens | ~$9.74 |
| With minification (85% reduction) | 7.2 MB HTML tokens | ~$1.44 |
14.2 Rendering Strategies per Page
| Page | Strategy | Rationale |
|---|---|---|
| / (landing) | Static | No dynamic data; CDN cached |
| /dashboard | Client-side fetch | Data changes frequently |
| /audits/[id] | Client-side fetch | Personalized audit data |
| /docs | Static | Markdown content; fully cacheable |
15. Testing Strategy
| Test Suite | Cases | Coverage |
|---|---|---|
| url-normalizer.test.ts | 6 | HTTPS enforcement, SSRF blocking |
| schemas.test.ts | 8 | Zod validation schemas |
| snapshot-validator.test.ts | 2 | Minimum content thresholds |
| context-builder.test.ts | 1 | Prompt context formatting |
| prompt-builder.test.ts | 1 | Template rendering |
| domain-mapper.test.ts | 2 | AI output → AuditReport mapping |
| json-parser.test.ts | 5 | Fence stripping, trailing commas |
| app-error.test.ts | 4 | Error class hierarchy |
| audit-repository.test.ts | 4 | save, findById, findRecent |
| analysis-orchestrator.test.ts | 5 | Retry loop, error bubbling |
| extract-route.test.ts | 2 | API route integration |
| analyze-route.test.ts | 2 | API route integration |
| Total | 42 | 100% pass rate |
16. Deployment Architecture
| Layer | Service | Config |
|---|---|---|
| Hosting | Vercel | Next.js first-party — zero config |
| Database | MongoDB Atlas M0 | 512MB free tier |
| AI | Google AI Studio | Gemini 1.5 Flash API |
| CDN | Vercel Edge Network | Automatic via Vercel deployment |
| CI/CD | GitHub Actions | Lint → typecheck → test on every push |
16.1 CI Pipeline
# .github/workflows/ci.yml
jobs:
ci:
runs-on: ubuntu-latest
steps:
- npm ci
- npm run lint # ESLint
- npm run check-types # tsc --noEmit (strict)
- npm run test # vitest run
# All three must pass — PR merge blocked otherwise17. Engineering Trade-offs
Cheerio vs Puppeteer
| Aspect | Cheerio ✅ | Puppeteer ❌ |
|---|---|---|
| Speed | <1 second | 5–15 seconds (browser startup) |
| Memory | ~50 MB | ~300 MB per instance |
| JS execution | No | Yes |
| Shopify SSR content | Available | Available |
| Cost (serverless) | Low | High (memory/time billing) |
Gemini vs GPT-4o
| Aspect | Gemini 1.5 Flash ✅ | GPT-4o |
|---|---|---|
| Cost | Lower | Higher |
| Context window | 1M tokens | 128K tokens |
| Free tier | Generous | Limited |
| JSON reliability | Good (with Zod retry) | Good (with function calling) |
MongoDB vs PostgreSQL
| Aspect | MongoDB ✅ | PostgreSQL |
|---|---|---|
| Schema flexibility | No migrations needed | Schema changes require migrations |
| Nested arrays | Native document model | Requires JSONB or JOIN table |
| Type safety (TS) | Good via typed generics | Excellent via Prisma |
18. Future Work & Roadmap
| Feature | Priority | Description |
|---|---|---|
| User Authentication | P1 | Clerk/NextAuth — multi-tenant audit history |
| PDF Report Export | P2 | Professionally formatted PDF download |
| Scheduled Audits | P2 | Weekly automated re-analysis per store |
| Recommendation Tracking | P1 | Mark implemented, track score delta |
| Competitor Comparison | P3 | Side-by-side audit of two stores |
| Puppeteer Fallback | P2 | Render JS-heavy SPA stores when Cheerio misses content |
| Redis Caching | P2 | Cache identical URL analyses for 24h |
| Shopify App Integration | P3 | Native Shopify App Store listing with OAuth |
| Fine-tuned Model | P3 | Improve schema reliability with domain-specific training |
Open Architecture Questions
1. Multi-page crawling: Should we crawl 3–5 pages (homepage + PDP + collection + cart) simultaneously rather than just the homepage for deeper coverage?
2. Vector embeddings: Could we embed past recommendations and semantically search them to provide "similar findings" context — improving AI quality without increasing prompt length?
3. Fine-tuning: Could a fine-tuned model on historical CRO audit data outperform prompt-based Gemini on structured output reliability and domain-specific recommendation quality?
CRO Engine Engineering Documentation · Version 1.0.0 · Sprint 7 · July 2026
This document is maintained alongside the codebase. Update when architecture decisions change.