Almarai Scan & Win — AI-Powered Receipt Verification
An AI pipeline that reads Arabic and English supermarket receipts, identifies qualifying Almarai products from abbreviated till codes, and rewards shoppers automatically — routing uncertain receipts to a human review queue. It evolved across three 2026 campaigns: Ice Cream, Back-to-School and Lulu Ice Cream.
!Problem Statement
Almarai's promotions required proof of purchase, but manual review couldn't keep up with a national campaign. Saudi till receipts are thermal-printed, bilingual, right-to-left and heavily abbreviated — Almarai products appear as codes like "ALM MOZ 450G" — which defeats simple text matching.
- Every receipt previously needed a person to approve it
- Arabic letters differing only by dots, blurred by thermal print
- Till abbreviations that can never be listed exhaustively
- Duplicate submissions, non-receipts and edited images
- Per-call AI costs that one account must not run up
- Recognition issues needing same-day fixes without a developer
✓The Solution
We built a staged pipeline: a vision model transcribes each receipt verbatim against a strict JSON schema, a classifier maps every line to Almarai's product list with a confidence score, and a measured 0.90 threshold decides between instant approval and a human review queue. Prompts live in Firestore so operations staff can retune them without a deployment, and every change is regression-tested against labelled receipt lines.
Key Features
Verbatim Receipt Transcription
- Gemini vision with a fixed JSON response schema
- Merges bilingual line pairs and wrapped rows, strips barcodes
- Checks item count against the receipt's printed count
- Images downscaled for AI; full resolution kept for audit
Confidence-Gated Classification
- Each line mapped to a fixed enum of Almarai products, or none
- Per-line confidence score
- 0.90 threshold chosen by measurement on real receipts
- Everything below the threshold routed to people
Human Review Queue
- Receipt image shown alongside extracted lines
- Approve through the same award path, or reject with a reason
- Ops dashboard using server-side aggregate counts
Fraud & Duplicate Controls
- Not-a-receipt detection
- Date | total | tax fingerprint dedupe against approved receipts
- Image-type guard after hallucinations were observed on PNG input
- Transactions that re-check status before any award
Cost & Abuse Controls
- Server-side AI proxies with ID-token verification
- Transactional Firestore rate limiter shared across function instances
- Per-user daily limits with separate buckets per engine
- Retries with backoff for transient upstream errors only
Operable by Non-Developers
- Prompts editable in Firestore, live on the next upload
- Prompt candidates compared against labelled line sets before release
- Handover playbooks for daily checks, rollback and incidents
- Python tooling for seeding, backups and points reconciliation
Tech Stack
Frontend
Backend
Services
Architecture
Multi-Tenancy
Per-campaign Firebase projects; the classify-and-award decision moved fully server-side in one transaction by the Lulu edition
Realtime
Request/response pipeline: upload → vision → classify → gate → award or review
Data Model
Users → Receipts → Recipes → Gifts → Lucky Draw Entries → Prompts (metadata) → Rate Limits → Activity Logs
My Responsibilities
- 1Two-stage vision + classification pipeline design
- 2Prompt engineering for verbatim Arabic/English transcription
- 3Confidence gating and human-in-the-loop review flow
- 4Server-side AI proxies with per-user rate limits
- 5Transactional, server-authoritative points awards
- 6Evaluation harness with labelled Arabic ground truth
- 7Ops dashboard, Python tooling and handover documentation
Challenges Overcome
Outcome
Improved across three campaignsRoadmap
- Re-measure the confidence threshold per product category
- Deploy regression-tested prompt candidates
- Extend server-side decisions to every edition
Impact
- ✓Confident receipts approved and paid instantly — only uncertain ones need a person
- ✓Product recognition no longer depends on an exhaustive abbreviation list
- ✓Operations staff retune the AI in minutes through Firestore, with no deployment
- ✓AI behaviour changes regression-tested against labelled Arabic and English lines
- ✓Each campaign improved on the last, ending with a fully server-authoritative decision