A single reference for the full project lifecycle, team capability model, tool ownership, global network topology, and the operational challenges of running IT across the US and United Kingdom.
Six stages from intake to retire, governed by the 60/40 capacity rule.
Five domains around a central Architects center of excellence.
Who owns each tool: internal, shared, or external vendor SaaS.
AWS + Azure core with six operational domains.
Five teams sharing one Infrastructure platform.
UniFi VLANs connect the sandbox with AWS and Azure sites.
Coverage gaps across PST → ACST and the 17.5 hr offset.
Resource management ties every view back to capacity.
Tracks utilization across all silos in real time, flags over-allocation at intake, and enforces the 60/40 split — escalating when thresholds are breached. Tech debt, support, enhancements, and maintenance all live inside the 40%.
| Domain | Internal — we own & author | Shared responsibility | External / vendor SaaS |
|---|---|---|---|
| BUILD | Container templates · deployment runbooks | Incus profiles · AWS/Azure site connectivity | OCI images · package repositories |
| SECURE | Vault policies · AppRole access · UFW rules | UniFi VLANs · Caddy TLS termination | AWS and Azure identity boundaries |
| PROTECT | Incus snapshots · recovery runbooks · retention | ZFS storage · persistent service volumes | Offsite backup targets |
| OBSERVE | Alert rules · Grafana dashboards · runbooks | Prometheus · Loki · Pyroscope · CheckMK | Email and webhook notifications |
| AUTOMATE | Ansible playbooks · Python · shell scripts | Semaphore · Flowise · Git repositories | Provider APIs |
| GOVERN | Service standards · architecture decisions | Vault audit records · container inventory | AWS and Azure service controls |
Model lifecycle, inference, observability
Custom apps, integrations, APIs
Laboratory technology systems
Reporting, analytics, dashboards
System interfaces & integrations
End-user support & ITSM
Limited capacity
40% break-fix / ops
Timelines slip when infra is over-allocated.
Internal roadmap deferred indefinitely.
Quality & morale drop under pressure.
Demand visibility, enforced 60/40, triage.
Backup & disaster-recovery failover for all primary sites.
International site (GMT, UTC+0) in the United Kingdom.
UniFi · Incus · Ollama · Vault · Prometheus · Grafana connect through defined boundaries.
Follow-the-sun rotation — London hands off to EST at 14:00 UTC, EST to Pacific at 23:00 UTC, Pacific to London at 09:00 UTC.
Grafana alert routing — urgent alerts notify on-call; lower priorities queue for the next business-hours team.
OpenProject tracks after-hours burden — after-hours hours count against the 40% break-fix allocation.
Runbooks live in Omaha DR, reachable across all time zones — auto-failover removes manual cross-TZ coordination.
Review UK data residency — confirm London data remains within approved regions; audit cross-border backup flows.
UTC 08:30–17:00 has no staffed team; Grafana alerts are the safety net.
London incidents can reach the US team outside core hours, creating the largest gap in the portfolio.
Business demand vs internal roadmap competes for the same 40% capacity.
London data must remain within approved UK regions and controls.
From intake at the ARB to retirement and review — lifecycle, capability, ownership, network, and coverage all governed through a single 60/40 capacity model.