agentic AI sandbox · R&D Sandbox
01 / 11
R&D Sandbox · Reference Platform

Agentic AI SandboxArchitecture Blueprint

A single reference for the full project lifecycle, team capability model, tool ownership, global network topology, and the operational challenges of running IT across the US and United Kingdom.

Reference platform · R&D Sandbox Multi-site · Global operations AWS Private Cloud · Microsoft Azure 60 / 40 Capacity Model
agentic AI sandboxR&D Sandbox · Reference Deck
Reference architecture deck
September 2026
What's inside
The Operating Model at a Glance
Seven connected views — from how work enters the pipeline to how it is run, secured, and supported worldwide.
01

Project & Service Lifecycle

Six stages from intake to retire, governed by the 60/40 capacity rule.

02

IT Capability Wheel

Five domains around a central Architects center of excellence.

03

Tool Responsibility

Who owns each tool: internal, shared, or external vendor SaaS.

04

Infrastructure Ecosystem

AWS + Azure core with six operational domains.

05

Team Structure & Contention

Five teams sharing one Infrastructure platform.

06

Network Topology

UniFi VLANs connect the sandbox with AWS and Azure sites.

07

Time Zones & On-Call

Coverage gaps across PST → ACST and the 17.5 hr offset.

OpenProject Control Plane

Resource management ties every view back to capacity.

Section 01
Proposed Project / Service Lifecycle
Every project moves through six stages from intake to retire. 50–60% of employee time is reserved for projects.
Stage 1

Architecture Review Board

  • Discuss technologies
  • Security & governance
  • High-level diagram
  • Silo ambassadors
Stage 2

Discovery

  • Scope & inventory
  • Build vs buy
  • Time estimates
  • Materials / cost
Stage 3

Enterprise PM and Procurement

  • Resource allocation
  • 60% projects
  • 40% break-fix / ops
  • Capacity visibility
Stage 4

Project

  • Resourcing
  • Timelines
  • Stakeholders
  • Tiger Team
Stage 5

Ops Excellence

  • Documentation
  • SLAs
  • Training
  • Maintenance
Stage 6

Enhance / Retire

  • Annual review
  • Enhance or sunset
  • Feeds back to ARB
60%
Total project allocation
40%
Break-fix & ops excellence

OpenProject CE — Resource Management Layer

Tracks utilization across all silos in real time, flags over-allocation at intake, and enforces the 60/40 split — escalating when thresholds are breached. Tech debt, support, enhancements, and maintenance all live inside the 40%.

Section 02
IT Capability Wheel
Five domains surround a central Architects team — each with a dedicated architect role and a clear set of tools.
AI Automations Architects Incus PostgreSQL · Qdrant Claude · Copilot Azure DevOps · VS Code GLPI · Snipe-IT Open WebUI LLM Guard · Garak MLflow · Marquez OpenProject · OpenEMR Snipe-IT · Power BI Infrastructure Application Dev Service Desk Business Apps Security
Infrastructure
AMD Threadripper · UniFi Gateway · UniFi Switches · VLANs · Incus · ZFS · PostgreSQL · Redis · Qdrant · MinIO · NVIDIA RTX 3090
Application Development
Claude · Azure DevOps · GitHub Copilot · VS Code · Python · Node.js · Playwright · REST APIs
Service Desk
GLPI · Snipe-IT · Open WebUI · Obsidian · Mailpit · Service Dashboard
Automations
Ansible · Semaphore · Flowise · k8n (Experimental) · Python · PowerShell · Shell scripts
Security & AI Governance
LLM Guard · Garak · LiteLLM · MLflow · Langfuse · Marquez · HashiCorp Vault · Caddy · UFW · Network Segmentation · AppRole · TLS
Business Applications
Experimental native Incus: ERPNext · OpenProject · OpenEMR · OpenMRS · Bahmni · GNU Health
Snipe-IT Inventory · Power BI · GLPI · Medical Equipment · Open WebUI · Services Dashboard

Shared Responsibility

PlatformIncus, ZFS, backupsAWS/Azure connectivity
SecurityVault, UFW, accessCaddy TLS, DNS, providers
OperationsGrafana, runbooksNetwork and site incidents
Section 03
Tool Responsibility by Domain
Every tool is categorized by ownership — internal policy & runbooks, shared config with vendors, or fully external SaaS.
Domain Internal — we own & author Shared responsibility External / vendor SaaS
BUILDContainer templates · deployment runbooksIncus profiles · AWS/Azure site connectivityOCI images · package repositories
SECUREVault policies · AppRole access · UFW rulesUniFi VLANs · Caddy TLS terminationAWS and Azure identity boundaries
PROTECTIncus snapshots · recovery runbooks · retentionZFS storage · persistent service volumesOffsite backup targets
OBSERVEAlert rules · Grafana dashboards · runbooksPrometheus · Loki · Pyroscope · CheckMKEmail and webhook notifications
AUTOMATEAnsible playbooks · Python · shell scriptsSemaphore · Flowise · Git repositoriesProvider APIs
GOVERNService standards · architecture decisionsVault audit records · container inventoryAWS and Azure service controls
Section 04
Infrastructure Ecosystem
Self-hosted agentic AI infrastructure connected to AWS and Azure sites through defined network boundaries.
Core Platform
Agentic AI Sandbox · AWS Site · Azure Site
UniFi networking · Incus containers · GPU inference · observability
RUN

Compute & network

AMD ThreadripperUniFi GatewayUniFi SwitchesVLANsIncusZFS
SECURE

Security boundaries

HashiCorp VaultCaddyUFWAppRole
DATA

State & retrieval

PostgreSQLRedisQdrantMinIO
OBSERVE

Telemetry

PrometheusGrafanaLokiLangfuse
AUTOMATE

Delivery & workflows

AnsibleSemaphoreFlowisePython
INFER

Agentic AI services

OllamaOpen WebUIOpenClawRTX 3090 GPUs
Section 05
Team Structure & Contention
Five teams serve the business — all depend on Infrastructure as a shared platform with limited capacity.
ML OPS

Model lifecycle, inference, observability

APP TEAM

Custom apps, integrations, APIs

LAB TECH

Laboratory technology systems

BI TEAM

Reporting, analytics, dashboards

INTERFACE

System interfaces & integrations

SERVICE DESK

End-user support & ITSM

Business demand
App team requests
LIS team requests
BI & Interface requests

Infrastructure

Limited capacity
40% break-fix / ops

Internal roadmap
Hardware refresh
Security patching
Platform upgrades · DR

Delayed projects

Timelines slip when infra is over-allocated.

Tech debt grows

Internal roadmap deferred indefinitely.

Team overloaded

Quality & morale drop under pressure.

OpenProject control

Demand visibility, enforced 60/40, triage.

Section 06
Global Network Infrastructure
UniFi-managed VLANs segment the sandbox, with controlled connectivity to AWS and Azure sites.
Sandbox network infrastructure map — UniFi network with AWS and Azure sites Network diagram showing the UniFi-managed sandbox, isolated VLANs, remote access, and controlled connectivity to AWS and Azure sites. United Kingdom Internet Hub Hybrid Cloud Azure · AWS External sites Sandbox Services Ollama · Open WebUI · Flowise Vault · Qdrant · Langfuse Boise Headquarters Boise, ID Portland Portland, OR Omaha Backup / DR Omaha, NE Louisville Louisville, KY Manchester Manchester, NH Raleigh Raleigh, NC Remote Users UniFi VPN London United Kingdom FW Legend Primary site (HQ / UK) Branch / secondary site Backup / DR site IPSec tunnels On-prem firewall UniFi Gateway / firewall Sites ● Boise — HQ · ID → Manchester · NH → Portland · OR → Louisville · KY → Raleigh · NC ★ London · UK AI sandbox R&D Sandbox · Global operations
7
Physical sites across the US & United Kingdom, plus remote VPN users.

Omaha DR

Backup & disaster-recovery failover for all primary sites.

London

International site (GMT, UTC+0) in the United Kingdom.

Sandbox Platform

UniFi · Incus · Ollama · Vault · Prometheus · Grafana connect through defined boundaries.

Section 07
Time Zone Convergence & On-Call
A seven-hour daily window has no staffed coverage; the seven-hour Boise ↔ London difference is the largest offset.
Time zone coverage, on-call challenges, and regionality across sandbox sites Three-section diagram: 24-hour coverage bar showing gaps between US and United Kingdom sites, on-call matrix by site, and regionality challenge cards. 24-Hour Coverage Map — Time Zone Convergence 0 2 4 6 8 10 12 14 16 18 20 22 UTC PST (CA/AZ) PST 9am – 6pm MST (Boise) MST 9am – 6pm CST (Omaha) CST 9am – 6pm EST (KY/NH/NC) EST 9am – 6pm Coverage gap ~4 hrs UTC 10:00 – 14:00 No staffed coverage Overlap ~1 hr On-Call Coverage Matrix by Site Site Time Zone Business Hrs On-Call After Hours DR / Failover Challenge Boise HQ MST (UTC-7) 9am–6pm Primary Grafana alerts Omaha DR Covers GMT gap London GMT (UTC+0) 9am–6pm Local team US team Limited 7 hr offset Louisville / Raleigh EST (UTC-5) 9am–6pm Secondary Grafana alerts Boise / Omaha 2hr offset from HQ Omaha (DR) CST (UTC-6) Automated Automated Monitoring Primary DR Failover site only Regionality & On-Call Challenges Coverage Gap ~7 hr window where no team is in business hours. Grafana alerts only coverage. GMT Escalation London after-hours incidents must wake US team at 3-5am MST. 7 hr offset is the largest gap. DR Coordination Omaha failover must be triggered across time zones. Runbooks must be clear for on-call staff. Compliance / Data UK data residency differs from US HIPAA. Cross-border data flows must be audited and logged. OpenProject Control Track on-call hours by region. Flag UK staff overallocation. Enforce rotation to prevent burnout. Recommendations 1. Establish a follow-the-sun rotation — London hands off to EST at ~14:00 UTC, EST to Pacific at ~23:00 UTC, Pacific to London at ~09:00 UTC. 2. Use Grafana alert routing — urgent alerts notify on-call; lower priorities queue for the next business-hours team. 3. OpenProject tracks after-hours burden — count after-hours hours against the 40% break-fix allocation. 4. Maintain runbooks in Omaha DR accessible to all time zones; Omaha auto-failover removes manual cross-TZ coordination for DR. 5. Review UK data residency — confirm London data remains within approved regions; audit cross-border backup flows.
Where this lands
Recommendations & Key Risks
Resolving capacity and coverage comes back to one control plane — OpenProject Community Edition — and disciplined runbooks.
  • 1

    Follow-the-sun rotation — London hands off to EST at 14:00 UTC, EST to Pacific at 23:00 UTC, Pacific to London at 09:00 UTC.

  • 2

    Grafana alert routing — urgent alerts notify on-call; lower priorities queue for the next business-hours team.

  • 3

    OpenProject tracks after-hours burden — after-hours hours count against the 40% break-fix allocation.

  • 4

    Runbooks live in Omaha DR, reachable across all time zones — auto-failover removes manual cross-TZ coordination.

  • 5

    Review UK data residency — confirm London data remains within approved regions; audit cross-border backup flows.

Top operational risks

8.5 hr coverage gap

UTC 08:30–17:00 has no staffed team; Grafana alerts are the safety net.

7 hr UK offset

London incidents can reach the US team outside core hours, creating the largest gap in the portfolio.

Infrastructure contention

Business demand vs internal roadmap competes for the same 40% capacity.

Data residency

London data must remain within approved UK regions and controls.

One blueprint, end‑to‑end.

From intake at the ARB to retirement and review — lifecycle, capability, ownership, network, and coverage all governed through a single 60/40 capacity model.

6
Lifecycle stages
5
Capability domains
7
Global sites
60/40
Capacity model
agentic AI sandboxR&D Sandbox · Reference Deck
Thank you
Questions & discussion
← → or Space to navigate · F for fullscreen