I build software and AI systems, and deliver programmes. End to end.
20+ years across fintech, payments and financial services — VP at JP Morgan, CTO of a regulated fintech, PRINCE2 project / programme manager. I build and ship production software and applied-AI (LLM) systems end to end — with real eval harnesses, not vibes — and lead complex delivery. Most recently I shipped BTCBitByBit single-handedly across web, iOS and Android, and a growing set of live AI tools.
Live, working products — not just repos. A flagship fintech product shipped solo across three platforms, and a growing set of small AI tools.
Consumer fintech platform · self-custody wallet · web, iOS & Android
A consumer fintech product (Bitcoin education + self-custody) — and a real engineering build. At its core: a multi-protocol self-custody wallet (Lightning BOLT11 + BOLT12, Liquid, on-chain) on the Breez SDK, plus an on-device web wallet running the SDK in the browser via WebAssembly, so users hold their own keys straight from a tab.
Shipped end-to-end across web (Next.js / React / TypeScript), native iOS and native Android, on AWS / Terraform — designed, built, tested (Jest + Playwright) and operated solo, deployed to production daily.
Role: Architect, Lead Engineer & Product — solo, end to end
A suite of small, live, open-source AI tools — each built around one idea: measure LLM systems, don’t vibe them, so every one ships a real eval harness. Grouped by area below — expand any category to explore. Built with FastAPI + Next.js on the Anthropic Claude API.
A patent-claim structure & antecedent-basis linter — a deterministic engine decides the §112 drafting defects, and the LLM only explains each and suggests a fix. Recall / precision / grounding evals, plus a judge that keeps it to drafting help.
Maps a patent claim’s limitations to supporting text in a prior-art reference (a claim chart) and flags what isn’t disclosed — a grounding filter drops any ‘disclosed’ quote that isn’t verbatim in the reference. Planted-disclosure evals.
Classifies an invention into CPC patent codes via retrieval — the LLM only picks from retrieved candidates, so it can’t invent a code, and abstains when nothing fits. Top-1 / recall@k + no-hallucinated-code evals.
Drafts a patent office-action response — a deterministic parser decides which limitations the rejection actually charted (verbatim-quote-verified), the LLM drafts the argument & amendment, and a no-new-matter gate verifies every added word traces to the specification. Includes the real Amazon 1-Click claim.
Reviews a contract or terms-of-service and flags the risky clauses (auto-renewal, liability, IP assignment, data rights) with a verbatim quote — grounded, with planted-clause evals. Educational, not legal advice.
A safe demo of NHS primary-care monitoring & recall: a deterministic rules engine decides what’s overdue for a synthetic patient, and the LLM only drafts the recall — with recall / precision evals.
Cited Q&A over drug/condition monitoring guidance that abstains when the corpus doesn’t cover the question — retrieval, citation-validity and faithfulness evals.
Maps a clinical note to SNOMED CT codes via retrieval — the LLM only picks from retrieved candidates so it can’t invent a code, evidence must be quoted verbatim, negation & family-history mentions are filtered, and the SNOMED hierarchy subsumes parent codes. Abstains when nothing fits.
Ship a clinical rule change safely: shadow-runs the proposed ruleset against the baseline over the same synthetic cohort, diffs the flagged volumes, gates on a golden regression cohort, and alarms when monitoring silently vanishes — you can’t A/B test on patients.
Builds NHS-style disease registers from SNOMED-coded records and flags what’s keeping patients off them — miscoded free text (with the exact suggested code) and status conflicts — deterministically; the LLM only drafts the completeness report, count-verified against the computed results.
A tool-using agent that researches a company into a sourced interview brief — an orchestrated search + fetch loop, hard guardrails, and source-grounding evals.
Cited document Q&A that abstains when the docs don’t cover it — retrieval metrics, citation-validity checks, and an LLM faithfulness judge.
Turns a plain-English question into a read-only SQL query, runs it on a sample database, and returns the rows — measured by execution accuracy (did it return the right rows), not string similarity.
Reviews a diff or PR for real bugs and security issues — not style — with every finding pinned to a changed line and scored against a planted-bug test set.
A witty-but-fair roast of any public GitHub repo — funny lines, a grade, and 3 genuinely actionable improvements, all grounded in the repo’s real README & facts, with a fairness judge.
An AI architecture/design reviewer — paste a design or RFC and get severity-ranked findings across scalability, reliability, cost, security & operability, grounded in the design. Planted-weakness evals.
An architecture-decision trade-off advisor — weighs your candidate options across the trade-off axes and recommends one of them (never an invented one), with the trade-off it’s accepting. Decision evals.
A grounded LLM pipeline that tailors a CV to a job description — forced structured output, plus deterministic and LLM-as-judge evals.
A tiny evaluation harness as a product — run a prompt over test cases, an independent judge scores pass/fail, and get a pass-rate, the failure pattern, and a better prompt. Its own eval proves it discriminates.
A mock interviewer that scores your answers against a rubric (relevance, depth, clarity, correctness) with feedback and a sharper follow-up — the grader is eval-tested to rank strong answers above weak.
An autonomous red-teamer that adaptively attacks an LLM you own, detects breaches with an independent judge, and reports how to harden — recall / precision evals included.
A running training-plan engine (a mini-Runna) — training-science rules decide a safe, progressive plan with a taper and Riegel paces, and the LLM only drafts the athlete briefing (easing off when you’re not 100%). Safety-gated evals + an opus faithfulness/safety judge.
Paste a run’s splits → split, pace-fade and HR-decoupling analysis + coaching feedback that can only cite the real numbers — a grounding filter drops any hallucinated stat. Planted-pattern + grounding evals.
Ask about your training history in plain English → it writes a read-only SQL query, runs it on a sample activity database, and returns the rows — measured by execution accuracy (did it return the right rows), not string similarity.
I'm a senior software engineer and PRINCE2 project / delivery manager with 20+ years across fintech, payments, capital markets and regulated financial services — JP Morgan Chase, Dun & Bradstreet, Virgin Media O2 and Tesco Mobile — delivering multi-million-pound outcomes, regulated products, complex programmes and teams I led. I move fluently between building software end to end and leading project, programme and change delivery — with governance, stakeholders and vendors. Most recently I built and shipped BTCBitByBit end to end on my own across web, iOS and Android — designed, built, tested (Jest + Playwright) and operated solo, shipped daily. I'd rather show you than tell you: the work is at btcbitbybit.com.
Engineering is the core — but taking a product to market solo meant wearing every other hat too
Interested in working together or have a project in mind?
I'd love to hear from you.