40 checks before your n8n workflow goes live
A practical production-readiness review for teams moving an n8n workflow from 'it ran in a demo' to something an operator can trust, debug, recover, and own.
Production failures are usually boring
A workflow can be logically correct and still fail in production because the webhook arrived twice, a vendor rate-limited the API, a credential expired, a retry duplicated a record, or nobody saw the alert. This checklist reviews those operational edges before launch. It is not an n8n certification or an official n8n document; it is Mopshy's implementation checklist.
Trigger & input integrity
01- Every production trigger has a documented event source, expected payload, and owner.
- Webhook endpoints verify signatures or another authenticating secret where the sender supports it.
- Unexpected or malformed payloads fail safely instead of being coerced into the happy path.
- Duplicate events are expected and tested rather than assumed not to happen.
Credentials & secrets
02- No API keys, tokens, passwords, or private URLs are hard-coded inside workflow nodes.
- Credentials are scoped to the minimum systems and permissions each workflow needs.
- Credential ownership and rotation steps are documented for the client or operating team.
- Test and production credentials are separated where the connected system supports environments.
Idempotency & data consistency
03- Every external write has an idempotency key, existence check, or equivalent duplicate-write protection.
- Retries cannot create duplicate contacts, invoices, tickets, calendar events, or outbound messages.
- Records have a stable correlation identifier that follows the run across systems.
- Partial success is handled explicitly when one system writes successfully and the next one fails.
Timeouts, retries & rate limits
04- Every outbound HTTP or API call has a defined timeout.
- Transient failures use bounded retries with backoff instead of immediate infinite retry loops.
- Permanent failures are separated from transient failures so bad payloads are not retried forever.
- Vendor rate limits and concurrency limits are known for the busiest integrations.
Error handling & recovery
05- A dedicated error workflow or equivalent failure path captures production failures.
- Failure alerts include workflow, execution, record or correlation id, and the failing step.
- There is a dead-letter or review queue for items that need human repair.
- The operator knows how to safely replay a failed item without duplicating previous side effects.
AI steps & human handoff
06- Model output is validated before it is written to a CRM, database, ticket, or customer-facing channel.
- Low-confidence, ambiguous, or out-of-scope cases have a defined human escalation path.
- Prompts, model choices, and structured-output expectations are versioned with the workflow.
- Irreversible or accountable actions require human approval when the risk justifies it.
Monitoring & alerting
07- The team can see execution success, failure, retry volume, and latency over time.
- Alerts go to a channel that a named person actually watches.
- Silent failures are detectable even when n8n itself stays healthy but a downstream result is missing.
- Logs retain enough context to debug a failure without exposing unnecessary sensitive data.
Testing & release
08- Happy-path, invalid-input, duplicate-event, timeout, and downstream-failure cases are tested before launch.
- A representative production-like test payload exists for every important branch.
- Changes are reviewed or compared before replacing the current production workflow.
- The first live release has a rollback or disable path that does not require rebuilding the workflow from memory.
Backups & deployment
09- Workflow definitions are exported or otherwise backed up outside the live n8n instance.
- The n8n database and required persistent data are backed up on a documented schedule when self-hosted.
- The upgrade process is documented, including who tests version changes before production rollout.
- Environment variables, domains, queues, workers, and storage dependencies are documented outside one person's memory.
Ownership & handoff
10- A named owner is responsible for the workflow after launch.
- The runbook explains normal operation, common failures, credentials, alerts, and replay steps.
- The client or internal team has access to the n8n instance, connected accounts, documentation, and backups they are expected to own.
- There is a review cadence for removing dead workflows, rotating credentials, and updating integrations when vendor APIs change.
How to use it
Review it with the person who will own the failure
Do not treat this as a paperwork exercise. Walk each item with the builder and the operator. For every unchecked item, decide whether it is a launch blocker, an accepted risk, or scheduled hardening work with an owner.
30-minute AI systems assessment
Find the AI system worth building next
Bring one process that is slow, repetitive, expensive, or leaking opportunities. We will map the business problem, the systems involved, and what an end-to-end AI system would need to own.
No obligation. If automation is not the right answer, we will say so.