Documentation

CloudTwin

Real-time digital twin platform for containerized infrastructure. Every Docker container gets a live virtual counterpart that mirrors its state, predicts failures, and self-heals — with your approval.

Go agentNode.js backendPostgreSQLRedisNext.js

Overview

CloudTwin implements a full closed loop — not just monitoring and alerting. Every metric that arrives from the agent flows through a five-stage pipeline that ends with an audited, human-approved action on the live container.

Monitor

Go agent polls Docker Engine API every 10s and POSTs snapshots to your backend over HTTPS.

Predict

Rules engine evaluates CPU and memory thresholds. What-if simulator extrapolates trends at any load multiplier.

Decide

Decision engine proposes a restart with a 10-minute cooldown. You approve or reject it on the dashboard.

Act

Approved actions go into a Redis queue. The agent pops, executes via Docker API, and acks the result.

Quick start

Sign up at cloud-twin.vercel.app/signup, copy your API key and workspace ID from the onboarding page, then run the agent on any machine with Docker.

1

Start infrastructure

bash
cd infra && docker compose up -d

Starts PostgreSQL (5432), Redis (6379), and Mosquitto (1883).

2

Apply schema and seed rules

bash
cd backend && npm install && node src/db/init.js && npm start

Creates all tables and inserts 3 default alert rules.

3

Start the frontend

bash
cd frontend && npm install && npm run dev

Runs on http://localhost:3000. Sign up to get your API key.

4

Run the agent

bash
export CLOUDTWIN_BACKEND_URL=https://cloudtwin.onrender.com
export CLOUDTWIN_API_KEY=ct_your_key_here
export CLOUDTWIN_WORKSPACE_ID=your_workspace_id

cd agent && go run main.go

Your containers appear in the dashboard within 10 seconds.

Architecture

Three independent planes communicate over HTTPS and a Redis queue. The backend never connects directly to the monitored host — the agent polls for approved commands and executes them locally via the Docker socket.

text
Agent (Go binary on host)
  ├── Metrics loop     POST /api/snapshots every 10s
  └── Command loop     GET /commands/claim every 5s → docker restart

Backend (Node.js / Express)
  ├── Ingestor         receives snapshots, upserts twin state
  ├── Rules engine     evaluates CPU/memory thresholds
  ├── Decision engine  proposes restart (10-min cooldown)
  ├── Redis queue      approved actions queued per twin
  └── REST API         /twins /alerts /actions /simulate /settings

Frontend (Next.js)
  ├── Dashboard        live twin grid with health gauge + sparklines
  ├── Actions panel    approve / reject with full audit trail
  └── Simulation       what-if load projection

Agent

A compiled Go binary that runs on any machine with Docker. It reads the Docker Engine API via Unix socket — no Docker account, no OAuth, no cloud dependency. Two goroutines run independently: one publishes metrics, one polls for commands.

Install

bash
git clone https://github.com/AaryanBairagi/CloudTwin.git
cd CloudTwin/agent
go build -o cloudtwin-agent .
./cloudtwin-agent

How it works

The metrics loop calls GET /containers/json and GET /containers/:id/stats on the Docker socket, builds a snapshot, and POSTs it to the backend with an API key header. The command loop polls /commands/claim and executes any approved action via POST /containers/:id/restart.

Digital twins

Each container gets one row in the twins table. It is upserted on every snapshot so the dashboard always reflects live state. The health score (0–100) is a weighted penalty model:

text
score = 100 - (cpu_percent × 0.5) - (memory_percent × 0.3)

healthy   score ≥ 80
warning   score ≥ 50
critical  score < 50

The state column stores the full latest snapshot as JSONB. Metric history is stored separately in metric_snapshots and used by sparklines and the what-if simulator.

Alert rules

Rules are stored per workspace in Postgres and evaluated on every incoming snapshot. Three default rules are created at signup. You can toggle them from the Settings page.

criticalCPU > 90%→ creates alert + proposes restart
warningCPU > 75%→ creates alert + proposes restart
warningMEMORY > 85%→ creates alert + proposes restart

Self-healing actions

When a rule trips, the decision engine checks for a recent action (10-minute cooldown) then inserts a pending action. A human approves or rejects it on the dashboard. Approved actions are pushed to a Redis queue keyed by workspace and twin ID.

text
pending   → human approves
approved  → pushed to Redis queue
queued    → agent pops command
executing → docker restart running
completed → agent acked success
failed    → agent acked failure
rejected  → human rejected (terminal)

Every transition is timestamped in Postgres. Nothing is silent — rejected and failed actions are permanently recorded.

What-if simulation

The simulator fits a linear slope to the last 10 minutes of metric history and extrapolates 5 minutes forward at a chosen load multiplier. No ML — just linear regression against real metric_snapshots data.

text
slope        = (latest - oldest) / elapsed_seconds
projected    = current × multiplier + slope × 300s
time_to_90   = (90 - current × multiplier) / (slope × multiplier)

recommendation:
  ADD_REPLICA     projected CPU ≥ 90% or projected memory ≥ 88%
  MONITOR_CLOSELY projected CPU ≥ 70% or projected memory ≥ 75%
  SAFE            otherwise

API Reference — Auth

Dashboard routes require Authorization: Bearer <jwt>. Agent routes require X-API-Key: <key>. JWTs expire after 7 days.

POST/auth/signupJWT

Create account — returns token + agentApiKey (shown once)

POST/auth/loginJWT

Sign in — returns JWT token

GET/auth/meJWT

Current user info and workspace ID

GET/auth/api-keyJWT

List API keys (hashed, not raw)

POST/auth/api-key/rotateJWT

Rotate agent API key — new key returned once

API Reference — Twins

GET/twinsJWT

All twins in workspace, ordered by last update

GET/twins/:idJWT

Single twin with full state snapshot

GET/twins/:id/historyJWT

Metric snapshot history, last 100 entries

API Reference — Alerts

GET/alertsJWT

All alerts, most recent first, limit 100

API Reference — Actions

GET/actionsJWT

All actions in workspace, ordered by created_at

POST/actions/:id/approveJWT

Approve pending action — pushes to Redis queue

POST/actions/:id/rejectJWT

Reject pending action — terminal state

API Reference — Agent endpoints

These routes use API key auth, not JWT. The agent sends these automatically.

POST/api/snapshotsAPI Key

Ingest container metrics snapshot

GET/commands/claim?twinId=API Key

Pop next approved command for a twin

POST/commands/:id/ackAPI Key

Report execution result — closes the loop

API Reference — Simulation

GET/simulate/:twinId?loadMultiplier=2JWT

Project metrics at given load multiple

GET/settings/rulesJWT

List workspace alert rules

PATCH/settings/rules/:idJWT

Toggle rule enabled/disabled