AI Global Academy Join the waitlist

Courses / Optimise, automate, operate

Lesson 8.5 · 50 minMonitoring and maintenance

Duration~50 min in the lesson + ~30 min homework
PrerequisitesCheckpoint lesson-8.4.
Checkpointlesson-8.5

What you will have

The live site reports its own errors, is checked from outside for uptime, runs a daily check of the lead path, and has a tested database backup. A deliberately triggered error produces an alert in the student's chosen channel, and a monthly maintenance checklist is in the repository.

Video

The video for this lesson is not recorded yet.

Prompts used in this lesson

Part 2 — Error tracking

Prompt to Claude
Purpose: I installed Sentry with its Next.js wizard so that I learn about errors
on the live site. Make the setup safe for a site that handles personal data, and
give me a way to test it.

Context: the wizard's files are in the project (uncommitted). Our lead form and
CRM handle names, emails, phones and messages. Consent handling is from lesson 6.5.

Before changing anything: read Sentry's current Next.js documentation on data
collected by default and on scrubbing data. Tell me which pages you read and
which options you will use; do not rely on memory for option names.

Do:
1. Configure Sentry so that request bodies, form field values, cookies, IP
   addresses and URL query strings are not sent. Tracing and Session Replay off.
2. Errors are reported from production only, not from local development and not
   from the end-to-end test environment (.env.e2e).
3. Remove the wizard's example page. Add instead an admin-only endpoint and a
   "Send test error" button on an admin settings page; the test error's message
   says clearly that it is a test.

Done when: the build passes, the end-to-end tests pass, and you have shown me the
final configuration and one captured local test event (with reporting enabled
temporarily) proving that no form data or IP address is attached. Report
anything you could not verify.

Part 4 — The daily lead-path check

Prompt to Claude
Purpose: tell me within a day if the live lead path stops saving leads, and tell
me if the site cannot reach its database.

Context: lead endpoint at /api/leads; Telegram bot from lesson 4.5; the scheduled
job from lesson 5.5 and its CRON_SECRET check; Sentry is installed.

Build:
1. /api/health: returns ok only if a trivial database read succeeds. No data in
   the response.
2. /api/cron/lead-path-check, protected by CRON_SECRET exactly like the 5.5 job:
   run the real validation and database insert with a clearly marked synthetic
   lead, read it back, delete it. Skip notifications, ad conversions and AI
   scoring. It must never leave a row behind, also when a step fails halfway.
3. On failure: a Telegram message naming the failed step, and a Sentry error. On
   success: call the address in the environment variable HEARTBEAT_URL if set.
4. Schedule it once a day in vercel.json.

Done when: the build and end-to-end tests pass; you have called both endpoints
locally against the test database and shown the results, including one forced
failure with its Telegram message; and you have confirmed no row was left behind.

Part 5 — Backups and a restore drill

Prompt to Claude
Purpose: a database backup I control, and proof that it can be restored.

Context: production is the Supabase project in .env.local; the test project from
lesson 8.4 is in .env.e2e. Read variable names only; never print values.

Before starting: read Supabase's current backup documentation, check whether
Docker is available on this machine, and tell me which method you will use.

Do:
1. Create a script "npm run backup" that dumps the production database
   (structure and data) into a dated file in a backups/ folder outside Git
   tracking. Add backups/ to .gitignore. The file contains personal data.
2. Restore drill: wipe the TEST project, restore the dump into it, and compare
   row counts for leads, lead_notes, lead_tasks and lead_status_history between
   production and the restored copy. Double-check the target is the test project
   before wiping anything, and show me the check.
3. Write the restore steps into docs/maintenance.md in plain language.
4. Afterwards remove the restored production data from the test project, and set
   it up again (migrations, test admin user) so the end-to-end tests pass.

Done when: the row counts matched, the test project holds no production data, the
end-to-end tests pass, and the steps are written down.

Do along

Work on your own project and pause the video where a step says so.

  1. Pause after Part 2. Run the Sentry wizard yourself; choose error monitoring only.
  2. Run the Sentry prompt from Part 2. Add SENTRY_AUTH_TOKEN and any other variable Claude names to Vercel. Add Sentry to the providers on /privacy. Deploy, then create the issue alert.
  3. Pause after Part 3. Choose an uptime service; create a monitor for your home page.
  4. Pause after Part 4. Run the prompt from Part 4. Deploy. Add the /api/health monitor and the heartbeat monitor; set HEARTBEAT_URL in Vercel; redeploy.
  5. Pause after Part 5. Run the backup prompt from Part 5. Store the backup file off your laptop.
  6. Pause after Part 6. Ask Claude to write docs/maintenance.md as described in Part 6, and to add .github/dependabot.yml if you want update pull requests.
  7. On the live site, click "Send test error" and wait for the alert.
  8. Commit and tag with the commands under "Recap and next".

Check your work

  1. Click "Send test error" on the live admin. Expected: an alert arrives in your chosen channel, and the issue appears in Sentry with no form data attached.
  2. In Vercel's cron job settings, check the job is listed; after its first scheduled run, the heartbeat monitor shows a received ping and your CRM shows no synthetic lead.
  3. Open docs/maintenance.md. Expected: the monthly list, the key inventory by name, and restore steps with the date of your drill.

Common problems

  • No alert after the test error → the error is not reaching Sentry (variables missing on Vercel, or reporting limited to another environment), or it arrived but no alert rule or notification channel is set → check the issue list first, then the alert rule, then your notification settings.
  • supabase db dump fails immediately → Docker is not installed or not running → start Docker, or ask Claude to use the PostgreSQL dump tool with the connection string from an environment variable.

Homework

About 30 minutes, on your own, after the lesson. Lesson 8.6 does not depend on it.

  1. Complete the key inventory. Open the key list in docs/maintenance.md. For each key, sign in to its service and confirm where it is renewed and whether it has an expiry date. Fill in what is missing: names and dates only, never values. Put the earliest expiry in your calendar two weeks ahead. Done when: no key has an empty "where to renew" or "last rotated" entry, and the reminder exists.
  2. Alert drill away from your desk. Later today, click "Send test error" from your phone and note how many minutes pass before you notice the alert. If you did not notice it within the hour, change the channel or its notification settings and repeat. Deliverable: a line "Alert drill" in docs/maintenance.md with the date, the channel and the minutes. Done when: the line records a drill you noticed within the hour.
  3. A backup without Claude, and a date for the next. Run npm run backup yourself, move the new file to your private place off the laptop, and record the date in docs/maintenance.md. Then create a monthly calendar event for the checklist. Done when: the new dated file is stored off your laptop, the date is in the file, and the event repeats monthly.

Commit without a tag: git add -A, git commit -m "Homework 8.5: maintenance notes", git push.

Save your work

git add -A
git commit -m "Lesson 8.5: error tracking, uptime and lead-path checks, backup and maintenance checklist"
git tag lesson-8.5
git push
git push --tags