# Operational runbook

Day-2 operations for the `demo` service. See `docs/deployment.md` for the CD
pipeline and `docs/ssl.md` for TLS specifics.

## Health & signals

- **Liveness/readiness**: `GET /actuator/health` (public). `UP` requires the DB and
  the `ssl` indicator to be healthy.
- **SSL expiry**: the `ssl` health component reports days-to-expiry per certificate;
  it goes `OUT_OF_SERVICE` under 30 days and `DOWN` under 14. Alert on the Micrometer
  gauge **`ssl.certificate.expiry.days`** (min across the chain) — e.g. warn < 30,
  page < 14.
- **Metrics**: `GET /actuator/prometheus` (authenticated) — scrape for JVM, HTTP,
  HikariCP (`hikaricp_connections_*`), and the SSL expiry gauge. All metrics carry
  an `application` tag.
- **Logs**: structured JSON (ECS) on stdout in prod; every line carries `requestId`.
  Correlate a client report by its `X-Request-Id` response header:
  `docker compose ... logs app | grep '"requestId":"<id>"'`.

## Common tasks

### Deploy / rollback
Deploys are automatic to staging (push `main`) and gated for prod (`vX.Y.Z` tag).
To roll back, run the **rollback** workflow (pick env; blank tag = previous), or on
the host: `cd $DEPLOY_DIR && ./rollback.sh`. See `docs/deployment.md`.

### Rotate TLS certificates
Replace the three PEM files on the host (same paths) and restart:
`docker compose -f docker-compose.yml -f docker-compose.prod.yml up -d app`.
No rebuild/image change. The app **validates the trio on startup** and refuses to
open the port on a bad/expired/mismatched set — check logs for the specific reason.
Verify the served chain: `openssl s_client -connect HOST:4455 -showcerts`.

### Rotate the JWT secret
Set a new `JWT_SECRET` (≥32 bytes) and restart. All existing **access tokens**
become invalid immediately (they were signed with the old key) — clients re-login
or refresh. Refresh tokens are opaque and DB-backed, so they survive; issue new
access tokens on the next `/auth/refresh`.

### Unlock an account
Lockout is automatic (5 failures → 15 min) and self-clears. To clear early:
`UPDATE users SET failed_login_attempts = 0, locked_until = NULL WHERE lower(email)=lower('…');`

### Database backup / restore
```bash
# backup
docker compose exec -T db pg_dump -U "$POSTGRES_USER" "$POSTGRES_DB" | gzip > demo-$(date +%F).sql.gz
# restore (into an empty DB)
gunzip -c demo-YYYY-MM-DD.sql.gz | docker compose exec -T db psql -U "$POSTGRES_USER" "$POSTGRES_DB"
```
Take a backup **before** any deploy that includes a migration.

### Failed migration
See `docs/deployment.md` → "A failed migration": roll the app back, repair
`flyway_schema_history`, fix forward. Migrations are forward-only and immutable.

## Incident playbook

| Symptom | First checks | Action |
|---|---|---|
| Health `DOWN` | which component? (`/actuator/health` details when authorized) | DB down → check `db` container / connectivity; `ssl` down → cert near/at expiry → rotate |
| 5xx spike | logs by `requestId`; `hikaricp_connections_pending` | pool exhaustion → check slow queries / raise `DB_POOL_MAX`; downstream (mail/SMTP) failures are isolated to auth flows |
| Won't start after deploy | startup logs | SSL trio invalid (readability/expiry/mismatch), missing `JWT_SECRET`, or a failed migration → roll back |
| Auth abuse / brute force | `/auth/**` 429 rate, failed-login logs | lockout + rate limiting already active; tighten `RATE_LIMIT_*` / `LOGIN_MAX_FAILED_ATTEMPTS`; block source at the edge |
| Cert expiring | `ssl.certificate.expiry.days` gauge | rotate the trio (above) before 14 days |

## Tunables (env)

`DB_POOL_MAX`, `TOMCAT_THREADS_MAX`, `LOGIN_MAX_FAILED_ATTEMPTS`,
`LOGIN_LOCK_DURATION`, `RATE_LIMIT_CAPACITY`, `RATE_LIMIT_REFILL_PERIOD`,
`JWT_ACCESS_TTL`, `JWT_REFRESH_TTL`, `APP_ACTIVATION_TTL`, `APP_RESET_TTL`.
Restart to apply.
