All work

Case study · 202605 / 06

Agentic dev pipeline

A team Claude Code setup for a multi-tenant Django SaaS. Eight scoped subagents plan, build and review; hooks block risky commands and gate every turn.

Role
Designer and main maintainer
Year
2026
Stack
  • Claude Code
  • Subagents
  • Skills
  • Hooks
  • Python
  • Django
Cover reading "Agentic dev pipeline" with the tagline "Eight scoped subagents, three hooks and a three-lens review gate on a multi-tenant Django SaaS"

Problem

Our team of 15–20 developers builds a multi-tenant restaurant management SaaS on Django REST Framework and PostgreSQL. It handles orders, invoices, VAT, payments and stock for many restaurants in one database, so the expensive bugs are specific: one restaurant’s query reaching another’s data, an invoice total that drifts from the report total, a model change shipped without its migration.

We already used coding agents on this codebase; the first version of this setup, in April 2026, was Copilot instructions plus hooks for the Django system check and a command blocklist. An agent working from generic defaults doesn’t know this codebase’s layering or tenancy rules, and nothing makes it run the checks before it says “done”. With 15–20 people prompting agents in their own way, every session was a fresh roll of the dice: one developer’s agent knew to scope by restaurant, the next one didn’t. I wanted a setup where an agent could take a feature from scope to a reviewed branch, and where the checks that matter for this codebase ran every time, whether or not the model remembered them.

My role

I designed the setup and wrote most of it: the project brain (CLAUDE.md), the layered rules, the subagent definitions, the orchestration skills and the hook scripts. It is committed to the backend repo, so the whole team works with it. Teammates commit with it and have contributed fixes to the rules. I wrote 29 of the 42 commits that touch .claude/ or CLAUDE.md. The same agent set now runs in our frontend and CMS repos too.

Architecture

The /feature skill runs in the main session and moves through Plan (explorer and architect, read-only), Build (main session or code-writer) and Verify (test-writer, security-checker, code-reviewer). /full-review then runs pr-reviewer, security-checker and qa-engineer in parallel and writes a PR-REVIEW.md verdict before a human merges into uat. PreToolUse, PostToolUse and Stop hooks act on the session throughout.
Skills orchestrate from the main session, because subagents can’t spawn subagents. Hooks are the part no prompt can argue its way past.

Project brain and rules. CLAUDE.md holds the stack, verification commands, tenancy rules and a completion checklist. .claude/rules/ splits the conventions into one always-on architecture.md (layering and import boundaries) plus path-scoped files that load only when the agent touches matching files:

---
paths: ["**/models.py", "**/models/*.py"]
---

Eight subagents, each with only the tools it needs. explorer and architect are read-only: one maps the patterns an app already uses, the other turns acceptance criteria into a plan with an API contract, permission codes, migration notes and a test matrix. code-writer and test-writer can edit and run commands. code-reviewer, security-checker, pr-reviewer and qa-engineer review. The read-only ones are restricted at the tool level:

---
name: security-checker
description: Validates multi-tenant isolation, permission safety and data integrity after code changes. Read-only.
tools: [Read, Grep, Glob]
---

Skills as orchestrators. /feature runs in the main session and walks scope → recon/design → build → self-verify → delegate (test-writer, then security-checker, then code-reviewer) → resolve. It closes only when a checklist with real command output is satisfied, including a test that proves one restaurant can’t touch another’s data. /review-changes is the lighter review. /full-review is the pre-merge gate: it runs verification once, sends the same package to pr-reviewer, security-checker and qa-engineer in one parallel batch, then merges their findings into a single severity-ranked PR-REVIEW.md with a verdict from Block to Approve. Workflow skills are manual-only (disable-model-invocation), so the model never starts a full pipeline on its own.

Three hooks in .claude/settings.json (trimmed):

{
  "PreToolUse":  [{ "matcher": "",                 "hooks": [{ "command": "… safety_check.py", "timeout": 5 }] }],
  "PostToolUse": [{ "matcher": "Edit|Write|MultiEdit", "hooks": [{ "command": "… django_check.py", "timeout": 30 }] }],
  "Stop":        [{ "matcher": "",                 "hooks": [{ "command": "… stop_gate.py", "timeout": 120 }] }]
}
  • PreToolUse denies shell commands that match a blocklist: rm -rf, git push --force, git reset --hard, --no-verify, migrate --fake, raw DROP/DELETE FROM, and any SSH or git traffic to the server remotes.
  • PostToolUse runs manage.py check after every Python edit and, if it fails, tells the agent to stop writing code until it passes.
  • Stop gates the end of every turn: makemigrations --check, ruff format and lint, and a regex secret scan of changed files. If any check fails it blocks the stop and hands back the reason. It respects stop_hook_active, so it can’t loop forever.

Key decisions & tradeoffs

Hooks for rules, prompts for judgement. Anything that must always happen (the system check, the migration check, the command blocklist) is a hook, because a hook runs every time. Prompts are kept for things that need judgement, like whether a queryset is scoped correctly.

A reviewer’s taxonomy comes from our own bugs. qa-engineer hunts 21 bug classes, each taken from defects this repo actually shipped, with the fixing commits cited: money drifting between invoice and report, rounding bases, cross-tenant querysets, NULL-matching filters, concurrent double-capture. A finding needs an entry state, a trigger, a wrong outcome and a file:line, or it is labelled unproven. I mirrored the frontend repo’s structure but rebuilt the list from this repo’s history rather than copying it.

Three overlapping lenses, verification run once. The reviewers overlap on purpose: two lenses finding the same issue raises confidence. The orchestrator runs check, the migration drift check and the tests once and gives the output to all three, so they don’t each re-run the suite.

Fewer agents, then more. The first version had an orchestrator agent, a requirement analyst and a code writer. In June I cut those, moved workflows into skills and made every agent tool-scoped, because subagents can’t spawn subagents and the main session is the only real orchestrator. architect and code-writer came back a month later in narrower forms. qa-engineer and /full-review followed in September.

Precise, not paranoid, guardrails. The blocklist avoids bare --force (it would block pip --force-reinstall) and bare truncate (it would block a harmless .truncate() call). When the secret scan flagged a stub value in a test settings override, I exempted only values that call themselves fixtures (test-, dummy-, fake-), so a real key pasted into a test still gets caught.

Outcome & metrics

  • Used by the backend team: teammates’ commits carry the Claude co-author trailer, and they have committed rule fixes themselves. The same agent set runs in the frontend and CMS repos, and parallel sessions run in separate git worktrees.

  • Review findings now show up in history as commits of their own, e.g. “fix zero-price order handling … (review findings B1, B3, S1, B7)”. 11 commits reference review findings by name.

  • A scoping bug in the review skills: they diffed against main, which trails our integration branch by about 320 commits, so one /full-review scoped 527 files and 82k lines instead of the branch’s ~73 files. After the fix, every review uses the integration branch as its base, includes uncommitted and mid-merge work, and stops itself when the file list passes 100 because the base must be wrong.

  • The gate earns its keep on the bugs that are hardest to see in a diff. Three it caught before merge:

    • A ×100 money error. Reviewing my own AI assistant branch, /full-review returned Block: a report query, inherited from older code, forgot to divide a flat discount by 100 on VAT-inclusive items with add-ons, so a £10 line with £5 off could come out as −£90. It had sat quietly on a dashboard screen; the assistant was about to state it in chat with full confidence.
    • Guest allergy notes going to the model next to the guest’s name. The reviewer flagged it as a privacy issue, and the tool now sends only a yes/no flag.
    • A forged payment webhook. A security review of the payments kernel found that an unsigned webhook’s reference reached the gateway URL unchecked, so a crafted value could make another payment’s status settle this one. References are now shape-checked and encoded, and a capture whose amount disagrees with its intent is refused and flagged for reconciliation.
  • Each PR-REVIEW.md ends with a resolution log: every finding marked fixed and pinned by a regression test (confirmed to fail before the fix), documented, or left open with the reason written down. Reviewers argue with code, not with impressions.

  • New developers get the codebase’s rules from the first prompt instead of from their first rejected PR. Conventions that used to live in a few senior heads now live in files the agent reads every time.

What I’d do differently

Start with the gate, not the builders. The hooks and the review skills paid off faster than the agents that write code. I’d build the Stop gate and /full-review first and add generators once the checks they must pass exist.

Put the gate in CI as well. Everything runs locally in the developer’s session today. Running the same migration, lint and secret checks in CI would cover commits made without the agent.

Log every verdict from day one. Each review writes a detailed report, but they live in working trees and get overwritten. If every /full-review had appended its verdict and finding counts to one log, I could say exactly how many blockers the gate stopped. Today I can show examples, not a rate.