Skip to content
All work

Previa: receipts to reconciled records

A Monash team prototype that turns receipts and bank statements into matched, tax-ready records using Mistral OCR and GPT-4o-mini. Demo offline.

Status
Status: Prototype
Role
Team project, AI and full-stack engineering
When
2025
Stack
  • React
  • TypeScript
  • Supabase
  • n8n
  • Mistral AI
  • OpenAI

Previa is a prototype our team built at Monash (FIT3195, Team Ivory) for Australian households and freelancers. Most finance apps start at the budget. Previa starts earlier: it turns receipts and bank statements into matched, categorised records that are ready for tax time. The public demo is offline.

The problem

Before anything reaches a tax return, someone has to do the pre-accounting:

  • Receipts are scattered across paper, email and payment apps.
  • Bank exports are cryptic. A line like "AMZN MKTP US*2A4BD" means nothing until it is matched to a receipt.
  • Matching is manual. Amounts, dates and merchant names have to be cross-checked by hand.
  • Australian deductions need specific categories that budgeting apps don't offer.

The pipeline

From a photo of a receipt to a matched record

Each model does the job it's suited to, and nothing reaches the database without passing a schema check.Select a step to see what it does.

Connections: Upload to Mistral OCR; Mistral OCR to Schema check; Schema check to GPT-4o-mini match; GPT-4o-mini match to Review.

Supabase holds the data, with row-level security on every financial table, and pushes results back to the React app through Realtime, so nobody waits on a refresh.

Two models, two jobs

We split the AI work by the kind of task. Reading a receipt is structured extraction from an image, with a fixed output, so it went to Mistral's vision model. Matching a cryptic bank line to the right receipt is a reasoning task, so it went to GPT-4o-mini. Two providers meant two clients and two sets of error handling. Running both models against the same test set is what settled the split.

Validate before you write

Early testing showed models sometimes return data that parses but is wrong: amounts as strings, dates in odd formats, missing fields from a blurry photo. So every response passes a Zod schema before it touches the database. Amounts are stored as integer cents to avoid floating-point drift in yearly totals.

typescript
import { z } from "zod";

const ReceiptExtractedSchema = z.object({
  merchant: z.string().min(1),
  amount_cents: z.number().int().positive(),
  gst_cents: z.number().int().nonnegative().optional(),
  date: z.string().regex(/^\d{4}-\d{2}-\d{2}$/),
  receipt_number: z.string().optional(),
});

export async function processReceiptOCR(fileUrl: string) {
  const response = await callMistralVision(fileUrl);
  const parsed = ReceiptExtractedSchema.safeParse(response.structured);
  if (!parsed.success) return { data: null, needsReview: true };
  return { data: parsed.data, needsReview: false };
}

The schema doubles as the contract for the OCR prompt. If a prompt change breaks the output, the schema fails at once instead of corrupting records downstream.

Where it got to

MVP stories done
24 of 30
About 83% of the planned MVP
Source: Project board at the time of writing
Automated tests
545
Vitest and React Testing Library
Source: At the time of writing; to be re-run
Test coverage
85%
As reported by the test run
Source: At the time of writing; to be re-run

Financial software earns its tests. A wrong amount or a silent matching failure has tax consequences, and writing the tests early caught negative amounts, duplicate receipts and timezone date shifts before they shipped.

Working on something like this?

I build agent workflows, document pipelines and the apps around them. If this looks like your problem, book a call.

Book a call (opens in a new tab)

wihithat@gmail.com

  • Status: ProductionVia Silvatron Pty Ltd

    For an Australian state government client

    Document pipeline

    A 12-node LangGraph pipeline extracts compliance documents row by row and links each value to its source page.

    Evidence: 90% F1 on one named benchmark; larger document sets varied

    Stack: LangGraph · Python · Pydantic · FastAPI

  • Status: ProductionBuilt at Silvatron Pty Ltd

    Claude Code, n8n, Bitbucket, Jira and Discord run first-pass reviews and diagnose failed pipelines. A person approves every fix.

    Evidence: About 20–40 developer hours a month (estimate)

    Stack: Claude · n8n · Node.js · Bitbucket