Tanvrit Compute

Security

Tanvrit Compute runs untrusted code on behalf of authenticated users across a fleet of agents that may be on hardware you don't fully control. The security model is built around three principles:

  1. Every actor has its own credential, scoped as narrowly as possible.
  2. Sensitive material never leaks into job environments.
  3. Every action is auditable, and nothing is hard-deleted for at least 30
  4. days.

This page documents the model in full.

Credentials

There are three distinct credential types in the system, and they live on different code paths.

Agent API keys

Every tanvrit-agent worker has its own API key, issued at registration time by the control plane. Properties:

  • Node-scoped. An agent key can only act on behalf of the node it was
  • issued for. It cannot submit jobs as a user, list other nodes, or read user data.

  • BCrypt-hashed at rest. The plaintext is shown to the operator exactly
  • once, at creation time. The control plane stores only the hash.

  • Rotatable from the portal. Click Rotate key on a node, get a new
  • plaintext, paste it into /etc/tanvrit/agent.json, restart the agent. The old hash is invalidated immediately.

  • Sent in X-API-Key on every agent → control-plane request.
  • Rotated on every registration, and persisted by the agent to
  • $TANVRIT_STATE_DIR/agent-identity.properties (mode 0600) before its first authenticated call. Only the bcrypt hash is stored server-side, so there is no plaintext to hand back — a re-attach therefore issues a new key rather than reusing the old one. See self-hosting.md → Node identity.

  • Valid regardless of node status. An OFFLINE, PAUSED or DRAINING node still
  • authenticates; only a deleted node does not. Auth used to be scoped to ONLINE nodes, which meant the 60-second stale sweep silently invalidated the key of any machine that went to sleep — its next heartbeat 401'd and the agent registered itself all over again as a second node.

Registering a node

POST /api/compute/nodes/register is reachable without credentials, and stays that way: no node in the live fleet has a provisioned tnv_… user key, so anything gated on one would not run where duplicate rows are actually created.

Continuity therefore rests on X-Agent-Id, and the owner is enforced as a binding rather than as part of the lookup key:

  • The identity mapping is keyed on sha256(agentId) alone. A returning machine
  • re-attaches to the node row it already owns whether or not it has a key.

  • The agent id is a credential. It is 256 bits of SecureRandom, generated
  • once, kept 0600, and never logged — the same strength as the node key it stands in for. Anyone holding it can re-attach as that node and be issued a fresh node key, exactly as the real agent can.

  • An identity first claimed with a resolvable bootstrap key is bound to that
  • owner (derived from the credential, never from the request body). From then on only that owner can re-attach it; a caller who merely knows the agent id is refused and gets an unrelated new row. Enrolling a key only ever tightens this.

  • An identity claimed without a key stays unowned and re-attaches on the
  • agent id. To stop a guessable operator-supplied TANVRIT_AGENT_ID from standing in for a secret, an unowned claim requires 32 characters or more; a shorter one is refused and the machine gets a new row each time. Generate it with openssl rand -hex 32.

  • The agent id is hashed (sha256) before it is stored, so a database dump does
  • not hand out re-registration tokens.

  • Deleting a node through the platform-admin route revokes its identity mapping,
  • so re-registration cannot be used to undo the delete.

  • Set COMPUTE_REQUIRE_REGISTRATION_KEY=true to require a resolvable bootstrap
  • key once every agent has one. Nothing validated TANVRIT_API_KEY server-side before this change, so check the [compute-register] logs first — an agent running with a garbage or revoked key works today and will start failing.

Rate limiting on this endpoint uses two buckets: one keyed on the resolved client IP (fair per agent behind a proxy) and one on the socket peer, which no header can move. CF-Connecting-IP and X-Forwarded-For are client-writable whenever the origin is reachable without the proxy in front, so the first bucket alone would let a caller pick a fresh allowance per request on an endpoint that writes rows.

Node key format

Keys issued from the identity work onward are cnk_<nodeId>.<secret>. The node id half is not a secret — it appears in every API response — and it exists so validation is one indexed lookup plus one bcrypt verify. Before that, every call to /heartbeat, /poll, /result and /logs — none of which sit behind a login — bcrypt-verified against every node in the fleet, which is a CPU amplifier an unauthenticated caller could aim. Keys minted earlier are bare UUIDs and still authenticate by scan; a caller that keeps failing is throttled per socket peer, and legacy keys drain as nodes re-attach or keys are rotated.

User API keys

For programmatic access from the SDK or CLI without an interactive login. Properties:

  • Format: tnv_<prefix> where <prefix> is the first 8 characters of the
  • key id, used as a fast lookup index. The full secret is stored as a hash.

  • User-scoped. All actions are attributed to the owning user.
  • Per-key quotas. Each key carries:
  • - maxJobsPerHour — sliding window enforced by the control plane. - maxConcurrent — number of in-flight jobs the key may have at any time. - allowedJobTypes — subset of COMMAND, DOCKER, GITHUB_ACTION, AI_TOOL. A pytest-only key, for example, can be restricted to DOCKER jobs.

  • Sent in X-API-Key for SDK requests, or as `Authorization: Bearer
  • tnv_…` for the CLI.

JWT (interactive)

The portal and the marketing site's /login page issue a short-lived JWT on successful login. JWTs:

  • Are bound to a single user account.
  • Carry the user's role claims (user, admin, owner).
  • Are sent in Authorization: Bearer … for user-facing endpoints.
  • Are not accepted on agent-facing endpoints. The agent surface only
  • trusts X-API-Key.

Webhook signing

Outbound webhooks (job-state notifications, schedule fires, etc.) are signed with HMAC-SHA256 using the per-webhook secret you configured. The signature is delivered in:

X-Tanvrit-Signature: sha256=<hex>

<hex> is the HMAC of the exact request body bytes. To verify in your receiver:

const expected = "sha256=" + crypto
  .createHmac("sha256", WEBHOOK_SECRET)
  .update(rawBody)        // raw bytes, NOT a JSON re-serialisation
  .digest("hex");

if (!timingSafeEqual(expected, headers["x-tanvrit-signature"])) {
  return reject(401);
}

Use a constant-time comparison. Webhooks that don't verify are dropped on the floor.

Sensitive environment variable stripping

When the control plane forwards a job spec to an agent, the following keys are stripped from env regardless of who set them:

  • TANVRIT_API_KEY
  • AWS_SECRET_ACCESS_KEY
  • GITHUB_TOKEN

(plus any key matching *_SECRET, *_TOKEN, or *_PASSWORD by suffix).

This prevents a user from accidentally — or intentionally — leaking their own control-plane key into a job environment where another job on the same agent could read it through /proc. If you genuinely need to pass a credential to a job, use the dedicated secrets API: secrets are mounted at job start and never persisted to disk on the agent.

Audit log

Every API key use is recorded in an append-only audit log with:

  • key_id (never the plaintext)
  • actor — user id or node id
  • endpoint
  • ip_address
  • user_agent
  • result — ok, unauthorized, forbidden, rate_limited
  • created_at

The audit log is queryable from the portal's Audit tab and via:

GET /api/compute/audit?keyId=…&since=24h

A failed-auth burst on a single key id will trigger a Slack alert if you have the alerts integration enabled.

Soft delete & forensic trail

Nothing is hard-deleted in the MVP. When you delete a node, a job, an API key, or a schedule, the row is marked deleted_at = now() and hidden from normal queries. The data remains in the database for 30 days, after which a daily housekeeping job purges it.

This is deliberate: in a security incident you almost always need access to the state of the system before the actor noticed they had been caught and started cleaning up. The 30-day window gives the team time to investigate, extract evidence, and rotate credentials before any rows physically leave the database.

If you need to recover a soft-deleted row inside the window, contact support or use the admin endpoint:

POST /api/admin/restore
{ "table": "jobs", "id": "job_…" }

Hard deletion on demand (for GDPR right-to-erasure requests) is available through the same admin path with ?purge=true.