CloudTwin
Real-time digital twin platform for containerized infrastructure. Every Docker container gets a live virtual counterpart that mirrors its state, predicts failures, and self-heals — with your approval.
Overview
CloudTwin implements a full closed loop — not just monitoring and alerting. Every metric that arrives from the agent flows through a five-stage pipeline that ends with an audited, human-approved action on the live container.
Monitor
Go agent polls Docker Engine API every 10s and POSTs snapshots to your backend over HTTPS.
Predict
Rules engine evaluates CPU and memory thresholds. What-if simulator extrapolates trends at any load multiplier.
Decide
Decision engine proposes a restart with a 10-minute cooldown. You approve or reject it on the dashboard.
Act
Approved actions go into a Redis queue. The agent pops, executes via Docker API, and acks the result.
Quick start
Sign up at cloud-twin.vercel.app/signup, copy your API key and workspace ID from the onboarding page, then run the agent on any machine with Docker.
Start infrastructure
cd infra && docker compose up -d
Starts PostgreSQL (5432), Redis (6379), and Mosquitto (1883).
Apply schema and seed rules
cd backend && npm install && node src/db/init.js && npm start
Creates all tables and inserts 3 default alert rules.
Start the frontend
cd frontend && npm install && npm run dev
Runs on http://localhost:3000. Sign up to get your API key.
Run the agent
export CLOUDTWIN_BACKEND_URL=https://cloudtwin.onrender.com export CLOUDTWIN_API_KEY=ct_your_key_here export CLOUDTWIN_WORKSPACE_ID=your_workspace_id cd agent && go run main.go
Your containers appear in the dashboard within 10 seconds.
Architecture
Three independent planes communicate over HTTPS and a Redis queue. The backend never connects directly to the monitored host — the agent polls for approved commands and executes them locally via the Docker socket.
Agent (Go binary on host) ├── Metrics loop POST /api/snapshots every 10s └── Command loop GET /commands/claim every 5s → docker restart Backend (Node.js / Express) ├── Ingestor receives snapshots, upserts twin state ├── Rules engine evaluates CPU/memory thresholds ├── Decision engine proposes restart (10-min cooldown) ├── Redis queue approved actions queued per twin └── REST API /twins /alerts /actions /simulate /settings Frontend (Next.js) ├── Dashboard live twin grid with health gauge + sparklines ├── Actions panel approve / reject with full audit trail └── Simulation what-if load projection
Agent
A compiled Go binary that runs on any machine with Docker. It reads the Docker Engine API via Unix socket — no Docker account, no OAuth, no cloud dependency. Two goroutines run independently: one publishes metrics, one polls for commands.
Install
git clone https://github.com/AaryanBairagi/CloudTwin.git cd CloudTwin/agent go build -o cloudtwin-agent . ./cloudtwin-agent
How it works
The metrics loop calls GET /containers/json and GET /containers/:id/stats on the Docker socket, builds a snapshot, and POSTs it to the backend with an API key header. The command loop polls /commands/claim and executes any approved action via POST /containers/:id/restart.
Digital twins
Each container gets one row in the twins table. It is upserted on every snapshot so the dashboard always reflects live state. The health score (0–100) is a weighted penalty model:
score = 100 - (cpu_percent × 0.5) - (memory_percent × 0.3) healthy score ≥ 80 warning score ≥ 50 critical score < 50
The state column stores the full latest snapshot as JSONB. Metric history is stored separately in metric_snapshots and used by sparklines and the what-if simulator.
Alert rules
Rules are stored per workspace in Postgres and evaluated on every incoming snapshot. Three default rules are created at signup. You can toggle them from the Settings page.
CPU > 90%→ creates alert + proposes restartCPU > 75%→ creates alert + proposes restartMEMORY > 85%→ creates alert + proposes restartSelf-healing actions
When a rule trips, the decision engine checks for a recent action (10-minute cooldown) then inserts a pending action. A human approves or rejects it on the dashboard. Approved actions are pushed to a Redis queue keyed by workspace and twin ID.
pending → human approves approved → pushed to Redis queue queued → agent pops command executing → docker restart running completed → agent acked success failed → agent acked failure rejected → human rejected (terminal)
Every transition is timestamped in Postgres. Nothing is silent — rejected and failed actions are permanently recorded.
What-if simulation
The simulator fits a linear slope to the last 10 minutes of metric history and extrapolates 5 minutes forward at a chosen load multiplier. No ML — just linear regression against real metric_snapshots data.
slope = (latest - oldest) / elapsed_seconds projected = current × multiplier + slope × 300s time_to_90 = (90 - current × multiplier) / (slope × multiplier) recommendation: ADD_REPLICA projected CPU ≥ 90% or projected memory ≥ 88% MONITOR_CLOSELY projected CPU ≥ 70% or projected memory ≥ 75% SAFE otherwise
API Reference — Auth
Dashboard routes require Authorization: Bearer <jwt>. Agent routes require X-API-Key: <key>. JWTs expire after 7 days.
/auth/signupJWTCreate account — returns token + agentApiKey (shown once)
/auth/loginJWTSign in — returns JWT token
/auth/meJWTCurrent user info and workspace ID
/auth/api-keyJWTList API keys (hashed, not raw)
/auth/api-key/rotateJWTRotate agent API key — new key returned once
API Reference — Twins
/twinsJWTAll twins in workspace, ordered by last update
/twins/:idJWTSingle twin with full state snapshot
/twins/:id/historyJWTMetric snapshot history, last 100 entries
API Reference — Alerts
/alertsJWTAll alerts, most recent first, limit 100
API Reference — Actions
/actionsJWTAll actions in workspace, ordered by created_at
/actions/:id/approveJWTApprove pending action — pushes to Redis queue
/actions/:id/rejectJWTReject pending action — terminal state
API Reference — Agent endpoints
These routes use API key auth, not JWT. The agent sends these automatically.
/api/snapshotsAPI KeyIngest container metrics snapshot
/commands/claim?twinId=API KeyPop next approved command for a twin
/commands/:id/ackAPI KeyReport execution result — closes the loop
API Reference — Simulation
/simulate/:twinId?loadMultiplier=2JWTProject metrics at given load multiple
/settings/rulesJWTList workspace alert rules
/settings/rules/:idJWTToggle rule enabled/disabled