Standards

The implementation methodology, in writing.
Operators, not titles.

The work behind the “checklist rigor, real testing, human review” line: the four named reviewer roles that sign off each implementation before it goes live, the pre-launch test plan run against the staging replica, the four consequential outbound categories that are deliberately kept human-reviewed, and the monthly retainer measurement that names what ran and what changed. Same vocabulary your operations lead will see in the retainer brief.

01 · Who reviews each implementation

Names, not titles. Each role signs the same artifact.

Every Northbench implementation lands with four reviewers named by role, not by team function. Each role reads a named artifact and signs the same artifact on the way out. The artifact list is in the audit brief; the names are in the go-day checklist.

The reviewer list is the same one your operations lead walks on launch day — no alternate roster, no fallback reviewer. If a reviewer can't be on the page that day, the launch moves, not the reviewer list.

Reviewer 01

Workflow lead

Who: Workflow lead (Northbench)

Artifact: Reviewer's checklist

What they read: The audit brief against the as-built workflow; verifies the same workflow, same tools, same escalation paths — last mile before the implementation close.

What they sign: Reviewer's checklist page — every line checked, named exceptions, the date and signature in the audit-trail appendix.

Reviewer 02

Integration reviewer

Who: Integration reviewer (Northbench)

Artifact: Smoke-test log

What they read: The smoke-test log printed against the staging replica — per-integration runs, replay timestamps, named test fixtures, no green ticks without a recorded artifact.

What they sign: An integration sign-off card per connector: target, fixture, run, outcome, reviewer — bound to the implementation brief.

Reviewer 03

Named Northbench operator

Who: Named Northbench operator (on call for the retainer)

Artifact: Integration sign-off card

What they read: Every artifact above plus the revert steps as rehearsed against the throwaway environment — the named operator has to be able to walk the steps blind.

What they sign: The named-operator page: who they are, their email, the same kill switch the client team will use, and a named escalation lane if the first reader is unreachable.

Reviewer 04

Client operations lead

Who: Client operations lead

Artifact: Go-day walk-through checklist

What they read: The same go-day walk-through checklist Northbench runs, line by line, in front of the client team — the reviewer's checklist, the smoke-test log, the sign-off card, the revert steps, in that order.

What they sign: A single sign-off card the client writes next to the named-operator page — same artifact, same name, no separate cover page.

02 · Pre-launch test plan

Five runs against the staging replica. Each carries a named target.

The test plan runs against the staging replica — never the production environment, never the demo. Each row in the plan ends with a named target, the same shape the pilot pause conditions are written in. A test that misses is logged by name; the next row does not paper over it.

Step 01

Per-integration smoke checks against the staging replica

Every integration the workflow touches — CRM, scheduling tool, inbox, calendar, sheets — runs a smoke check on the staging replica. No green ticks without a recorded artifact, no exceptions carved out after the fact.

Target: per-integration smoke run logged for every connector named in the audit brief.

Step 02

Replay runs against the staging replica

The same cases the audit brief walked are replayed end-to-end on the staging replica, against a named test fixture. The replay log lands next to the smoke-test log; the two artifacts have to read against each other.

Target: replay log bound to the smoke-test log by fixture name, not by gut feel.

Step 03

The named regression suite

The same regression suite the retainer will measure against — response time, completion rate, hours saved, error rate — runs once pre-launch against the staging replica as a full pass.

Target: full regression pass logged with the four named metrics, against the staging replica, before the pilot starts.

Step 04

Kill-switch rehearsal — one named toggle, flips in under a minute

A single named toggle in the workflow control panel flips the workflow to paused. The rehearsal is logged against a throwaway environment; the same target the pilot pause conditions will measure.

Target: workflow paused, end-to-end, in well under one minute.

Step 05

Operator-led walk-through with the client team on launch day

A Northbench operator walks the client team through the reviewer's checklist, the smoke-test log, the integration sign-off cards, the revert steps, and the named contact — line by line, on the launch call.

Target: client team walks the revert steps from a laptop, without Northbench on the line, on the end of the launch call.

Same five rows for every implementation
Each row ends in a Target line
Logged by name, not by green tick

What the test plan is anchored against

  • The staging replica — same fixtures, same integrations, same escalation paths as production. Never the demo, never the client environment.
  • The audit brief — the test plan walks the same workflows the brief walked; the brief is the page the test results bind to.
  • The pilot pause conditions — the same thresholds the pilot gate measures against, run once against the staging replica pre-launch as a full pass.
  • A throwaway environment for the kill-switch rehearsal — the steps have to work when it matters, not just on paper.

03 · Deliberately not automated

Consequential outbound stays on the page of a human.

This is not an aspiration. It is a written rule on the retainer cover page: anything in the four consequential categories below sits in a review queue, not on an auto-send path. The auto-send path exists for everything else. The four categories do not move into it on its own.

The rule is the same during the pilot as it is during the retainer — a pilot is not a license to skip the review pass, and a small pilot is not a carved-out exception. Renaming the list is a retainer change order, with a price attached before the work starts.

How a message is reviewed

  1. 01The AI draft is produced and sits in a review queue — no auto-send path exists for a consequential category.
  2. 02A Northbench operator reads the draft, edits if needed, and either approves or routes to the client's owner for sign-off on high-stakes items.
  3. 03A copy of the sent message, with the reviewer's name, lands in the monthly outcome report's audit-trail appendix.

The four consequential categories

Anything to a customer

A reply, a quote, a confirmation, a renewal reminder — any message that leaves your business under your name to someone who is paying you, or thinking about it.

Anything to a vendor

Work-order routing, partnership replies, statements of work — any message that obligates your business to a third party or spends money on your behalf.

Anything that moves money

Invoice creation, payment confirmations, refund acknowledgements, reconciliation entries — every step on the money path is a human-readable step before it commits.

Anything irreversible

Cancellations, deletions, opt-outs, scoping decisions that close a door. An automation cannot close a door without a human on the page first.

Deliberately not automated
Active during the pilot, not on retainer-only
Changes via retainer change order

04 · Retainer measurement & monthly reporting

Same four metrics, every month. Same definitions, no moving target.

The monthly outcome report covers the one workflow the retainer runs. Same four metrics every month, for the life of the retainer — no vanity dashboards, no logs to scroll. The report names what ran, what changed, and any incident that fired the kill switch, written so an operations lead can read it in two minutes.

The same four lines are the pilot gate metrics at day thirty — same names, same definitions, no rename between pilot and retainer. A trend that goes the wrong way for two consecutive months triggers a written review call, not a silent renewal.

Four metrics, every month

  1. 01

    Response time

    Minutes from intake to first reply on a typical case. Measured end-to-end against your inbox, CRM, or queue — not against the agent, not against the demo. The median is the number we report; the best case is what the demo showed.

  2. 02

    Completion rate

    Share of items the workflow carried from intake to a closing state without a human having to step in and re-run anything. Trend matters more than the absolute figure — the first month is the baseline, not the goal.

  3. 03

    Hours saved

    Wall-clock hours your team would have spent on the same workload without the workflow, minus the time they still spend on it. We count it from the start of the pilot against your pre-pilot baseline — never against an aspirational "fully automated" number.

  4. 04

    Error rate

    Share of items that needed a correction after the workflow finished — wrong field, wrong recipient, wrong amount, wrong state. An error routes back through the human-review queue and shows up by name in the same month it happened.

What the monthly PDF report contains

  • One page, sent the first of every month — the four named metrics for the month, bound to the same baseline the audit brief recorded.
  • Audit-trail appendix— every consequential message sent that month, with the reviewer's name; the appendix is the page the trust promise is read against.
  • A named root cause for any pause event — the pause latency, the revert latency, and the change shipped to keep it from recurring again, even when everything else is green.
  • The since-onboarding running totals: hours saved (rolling), error-rate trend, response-time trend, pilot → retainer switch date — the trend across months matters more than any single monthly figure.

What triggers a monthly review call

  1. 01

    Any pause event fires during the month

    A named pause event — error-rate ceiling, response-time regression, consequential-message error, second pause in a pilot week — auto-schedules a review call inside two business days.

  2. 02

    The trend line goes the wrong way for two consecutive months

    Flat or down on completion rate, up against the error-rate ceiling, response-time slippage — a written review call, not a silent renewal.

  3. 03

    A metric definition is about to change

    Renaming, scope-shifting, or moving a target is a retainer change order — the call is the page the change ships on.

Metric-rename rule

Renaming any of the four metrics, shifting the target, or moving the baseline is a written retainer change order — the same shape as the consequential category list. A silent override does not exist on this engagement.

Same trust language, already on the site

The four reviewer roles, the test plan, the consequential category list, and the monthly report format are not invented here. The same sentences, the same targets, and the same change-order rules already ship on /methodology, /how-it-works, and /comparison — the standards page is the page that holds them in one place.

The four promises, in plain English

Checklist rigor, real testing, human review — one engagement, in writing.

  • Review

    Named roles sign each implementation before go-live.

    The workflow lead, the integration reviewer, the named Northbench operator, and the client operations lead each sign the same artifact — same artifact, same name, no separate cover page.

  • Test

    Smoke + replay + regression + kill-switch rehearsed, in writing.

    Per-integration smoke checks, replay runs, the named regression suite, and the kill-switch rehearsal — all logged against the staging replica before the pilot starts.

  • Trust

    Consequential outbound kept on the page of a human.

    The four categories written into the brief — customer, vendor, money-path, irreversible — sit in a review queue, never on an auto-send path.

  • Measure

    One page. Four numbers. Every month. Same definitions.

    Response time, completion rate, hours saved, error rate — measured end-to-end against the same baseline the audit brief recorded, with no rename between pilot and retainer.

Next step

Get the deep methodology written up for your workflow.

Two weeks. A paid audit returns a written brief on one chosen workflow in your vertical: what to automate, what to leave alone, and what it should cost monthly — including the four reviewer names, the test plan walk, and the monthly report format the retainer will be measured against. $250 deposit, fully refundable against an implementation within sixty days.

Book a paid audit