Skip to content
Back to all work
AI Systems / Data Engineering2025–2026Client work

Procurement Intelligence

Every supplier offer read, normalized and scored — instead of the handful a buyer had time to open.

Role
Architect & sole engineer
Client
a European B2B cosmetics distributor
Scale
970 commits · ~102,000 lines (Sept 2026) · built solo
PythonFastAPIpandasMicrosoft GraphrapidfuzzPostgreSQLNext.jsTypeScriptAzure AD

Client system. Described here at the level of the problem, the architecture and my contribution. No source, no screenshots, no client data.

The same architecture — the confidence gates, the weighted scoring, the review queue — is running and screenshotted in OfferGrade, a demo built over invented data.

The problem

A distributor buys from around a hundred suppliers. Every one of them emails price lists — daily, weekly, whenever stock moves. The formats have nothing in common: Excel workbooks with the header row four rows down, PDFs, prices typed into the body of the message, columns labelled in German, French, Italian or Polish, and no agreement anywhere on whether a number is a net price, a recommended retail price, or something quoted only above a €75,000 order.

A buyer opens the ones they have time for. That is the whole system. Everything they don't open is a decision made by accident.

The question was not "can this be automated" — it was whether the messy half could be automated reliably enough to be trusted, because a normalizer that is right 80% of the time is worse than no normalizer at all.

What I built

A two-phase ingestion pipeline and a scoring engine behind an operator dashboard.

Phase 1 — ingestion. MSAL device-flow authentication into Microsoft Graph, incremental mailbox scanning with last-run state so a re-run doesn't reprocess a year of mail, SHA-256 content deduplication, and supplier attribution inferred from the sender domain — including the awkward case where supplier A forwards supplier B's list and the attribution has to survive the forward.

Phase 2 — normalization. A 3,245-line multilingual normalizer that finds the header row (or works without one), maps arbitrary column names onto a fixed schema via fuzzy matching, extracts EAN / REF / brand / price / currency / pack size, and classifies which of several price columns is actually the one you can buy at.

The scoring engine. Every normalized line is weighted against supplier price rank, delta from the last purchase, resulting margin, and a data-confidence score, then emits one of four actions — BUY, NEGOTIATE, REVIEW, IGNORE — with an override rule for offers confirmed as cheapest. A separate matcher cross-references live client requests against fresh offers, so when a customer asks for four hundred units of something, the buyer sees who can fill it and at what margin that same day.

The part that mattered most

Confidence, not extraction.

The easy failure mode is a pipeline that produces a clean-looking number from a column it misread. A price that is really a weight, a figure that is really RRP, a currency nobody converted — each of those reads as the best deal on the board until a human checks. So every extracted field carries a confidence value, low-confidence lines are routed to review rather than scored, and anomalies are surfaced explicitly instead of being silently averaged away.

That single decision is why the output gets used.

Architecture and delivery

  • Backend — Python 3.11 / FastAPI, modular routes over a layered pipeline orchestrator; pandas, openpyxl and xlsxwriter for the spreadsheet layer; rapidfuzz for header and brand matching; Pydantic throughout.
  • Auth — JWT with bcrypt, Azure AD for SSO, role-based access control over the operator surface.
  • Storage — started on SQLite, migrated to PostgreSQL as volume grew; S3 via boto3 for the raw file archive.
  • Frontend — Next.js 14 App Router, TypeScript, Tailwind, TanStack Table for the offer grid, Recharts for the analytics.
  • Deployment — Railway and Vercel, with CI running ruff, mypy, pytest and eslint.

Designed, built, deployed and documented single-handedly across 970 commits and roughly 102,000 lines, as of September 2026.

Where it got to

Running daily. As of 6 September 2026 it ranks 371,982 supplier offers across 165 suppliers, and flags 64,738 of them as buy signals — offers that clear both a confidence floor and a margin threshold. At the launch measurement on 31 July 2026 those figures were 188,000 offers across 112 suppliers, so the corpus has roughly doubled in five weeks without the ingestion changing shape.

The number that matters more is smaller. On that same day 106 offers were held out of the headline entirely — around €898,500 of apparent upside — because each was priced more than 30% below the last purchase for that product. A price that good is more often a unit mismatch, a currency error or an RRP column read as a net price than it is a bargain. The pipeline does not suppress them or quietly average them away; it parks them, says why, and asks a human. Being visibly unsure is the feature.

The buyer's morning changed from "which of these forty emails do I open" to "here are the eleven lines worth acting on".

You have a process thatshould be a system.

Tell me what arrives, who has to act on it, and where it currently falls over. That conversation is usually enough to scope the build.