Stretch.CodesLondon

Agentic systems built to run in production.

I build LLM applications and agent pipelines, along with the apps and infrastructure they need to run day to day.

I work as a contractor, usually embedded in an existing engineering team, and I stay through to production and handover.

Available for new contracts

Clients

  • RWS
  • Zoopla
  • Monzo
  • Which?
  • Eurostar
  • AND Digital
  • Reward Gateway

What I do

  • Agentic systems & LLM pipelines

    Agent loops, tool layers and multi-step pipelines, with a person checking the steps that need it. Most of this work is in places where a bad answer causes real damage.

    • Agent loops & LLM harnesses
    • Orchestration with LangGraph / LangChain
    • Human-in-the-loop checkpoints
    • Tool design & MCP servers
    • Prompt engineering & context engineering
    • RAG & retrieval pipelines
    • Evals, tracing & observability
  • Keeping the output honest

    Models make things up, and they agree with whoever is asking. I design around both: adversarial review passes where one agent’s job is to pull apart another’s output, sub-agents orchestrated by a stronger model that checks the work instead of rubber-stamping it, and evals that catch it when quality slips.

    • Adversarial review passes
    • Sub-agents orchestrated by a stronger model
    • Hallucination & sycophancy mitigation
    • Grounding & citation checks
    • Evals & regression testing on output quality
  • Building the product around it

    A model on its own is not a product. Someone has to build the app, the APIs and the infrastructure that put it in front of users and keep it up.

    • TypeScript · React · Next.js · Node
    • Python · LangGraph · AWS Bedrock
    • GraphQL & event-driven services
    • AWS · Terraform · CI/CD
  • Working out whether to build it, and where

    The question is whether the agent does something the existing interface can’t. If people are already pasting screenshots of your product into ChatGPT to get answers, they have answered that for you. If they aren’t, what you’re describing is often a form and a database, and I’d rather say so early than bill you for finding out.

    Then there’s where it lives. An MCP server puts you inside the assistants people already use, which is cheap to build and good for reach. An agent inside your own product costs more, and keeps the experience, the data and the revenue with you. The deciding question is usually which of those you can least afford to give up.

    • Working proofs of concept
    • Testing output quality on real data
    • Cost-per-use modelling
    • MCP server, in-product agent, or both
    • UX for chat and agent interfaces
    • A written recommendation at the end
  • Leading the team

    I have led teams of up to twelve engineers through replatforms and go-lives, and set up the testing and delivery process so it holds after I leave.

    • Architecture & delivery ownership
    • Testing & QA strategy
    • Agile / DevOps practices
    • Mentoring & upskilling

Selected work

  1. 01

    RWS

    Agentic patent drafting

    An in-house tool that takes an invention disclosure and produces a draft patent, with the attorney refining claims and scope as it goes. I handle document ingestion, the pipeline and prompt design, and the adversarial review passes that pull a draft apart before an attorney ever sees it.

    5 hrs → 10 min

    per patent draft

  2. 02

    Zoopla

    Agentic property search

    A working proof of concept on top of an MCP service, with the average cost per use measured so the business could decide whether to scale it. I also moved a legacy Perl estate onto Next.js, GraphQL and AWS.

    27M

    monthly visits replatformed

  3. 03

    Which?

    Personalised content

    Proof of concept through to production: AWS services serving personalised content to members, plus a set of experiments to find out what moved engagement.

    100k+

    users served per day

  4. 04

    Eurostar

    Booking service rebuild

    A two-year replacement of the legacy booking platform and agent portal, plus a team of six moving £80m of vouchers onto a new secure service. Front end shipped with over 90% test coverage.

    25%

    faster booking process

Also worked with

  • Monzo

    Internal tools that took a lot of manual operations work off the team, owned from scoping through to monitoring.

  • AND Digital

    Led two teams, around twelve engineers, on a point-of-sale app used in 300+ restaurants, then a restaurant ordering and payments app that cut order processing time by 30%.

  • Ember

    Accounting software for web and mobile in Node, React / React Native and Postgres.

  • Quander Digital

    Interactive technology and digital experiences for live events, including YouTube Brand Casts.

  • Young & Shand

    Social and marketing websites built with WordPress and React.

  • Reward Gateway

    Replatformed several separate web apps into one platform, and moved them onto a shared UI component library.

My own projects

  • Agentic Accountant

    An AI accounting agent I built for my own company, wired into FreeAgent, the HMRC APIs and Gmail. It answers questions about the accounts, VAT and corporation tax, and drafts replies to my accountant.

    Agents · Tool design · HMRC API · FreeAgent · Gmail API

  • ReturnSorted

    A UK self-assessment tax product I designed and built end to end, including the HMRC OAuth integration, magic-link auth and Stripe payments.

    Next.js · Vercel · AWS Lambda · Supabase · Stripe

I write up what I learn as I go — read the blog

What I work with

Ten years building full-stack systems for companies like Eurostar, Monzo and Which?. These days most of my work is on AI products in areas where the output has to be right: patent drafting, property search, consumer data.

AI & agentic systems

  • LLM application development
  • Agent loops & LLM harnesses
  • Agent orchestration (LangGraph / LangChain)
  • Tool design & MCP servers
  • Prompt engineering
  • Context engineering
  • RAG & retrieval pipelines
  • Human-in-the-loop pipelines
  • Adversarial review & verifier agents
  • Multi-agent orchestration
  • Hallucination & sycophancy mitigation
  • Evals & LLM-as-judge
  • Tracing, monitoring & observability
  • Token & cost instrumentation
  • Guardrails & output validation
  • Anthropic & OpenAI APIs · AWS Bedrock
  • Agentic dev tooling (Claude Code)

Core engineering

  • TypeScript / JavaScript
  • React / React Native
  • Next.js / Node.js
  • GraphQL
  • AWS / Terraform / DevOps
  • Python / Go / PHP
  • MySQL / Postgres / MongoDB
  • Unit & integration testing
  • Auth0 / OAuth 2.0

Contact

Tell me what you’re trying to build.

And where it’s got stuck. If I’m not the right person for it, I’ll tell you.